Langfuse ObservabilitySAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: langfuse-observability description: 'Set up comprehensive observability for Langfuse with metrics, dashboards, and alerts. Use when implementing monitoring for LLM operations, setting up dashboards, or configuring alerting for Langfuse integration health. Trigger with phrases like "langfuse monitoring", "langfuse metrics", "langfuse observability", "monitor langfuse", "langfuse alerts", "langfuse dashboard". ' allowed-tools: Read, Write, Edit version: 1.17.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - langfuse - monitoring - observability - llm compatibility: Designed for Claude Code --- # Langfuse Observability ## Overview Set up monitoring for your Langfuse integration: Prometheus metrics for trace/generation throughput, Grafana dashboards, alert rules, and integration with Langfuse's built-in analytics dashboards and Metrics API. ## Prerequisites - Langfuse SDK integrated and producing traces - For custom metrics: Prometheus + Grafana (or compatible stack) - For Langfuse analytics: access to the Langfuse UI dashboard ## Instructions ### Step 1: Langfuse Built-In Dashboards Langfuse provides pre-built dashboards in the UI at `https://cloud.langfuse.com` (or your self-hosted URL): - **Overview**: Total traces, generations, scores, and errors - **Cost Dashboard**: Token usage and costs over time, broken down by model, user, session - **Latency Dashboard**: Response times across models and user segments - **Custom Dashboards**: Build your own with the query engine (multi-level aggregations, filters by user/model/tag) **Accessing via Metrics API:** ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Fetch aggregated metrics programmatically const traces = await langfuse.api.traces.list({ fromTimestamp: new Date(Date.now() - 3600000).toISOString(), // Last hour limit: 100, }); console.log(`Traces in last hour: ${traces.data.length}`); // Get observations with cost data const observations = await langfuse.api.observations.list({ type: "GENERATION", fromTimestamp: new Date(Date.now() - 86400000).toISOString(), limit: 500, }); const totalCost = observations.data.reduce( (sum, obs) => sum + (obs.calculatedTotalCost || 0), 0 ); console.log(`Total cost (24h): $${totalCost.toFixed(4)}`); ``` ### Step 2: Prometheus Metrics for Your App Track the health of your Langfuse integration with custom Prometheus metrics: ```typescript // src/lib/langfuse-metrics.ts import { Counter, Histogram, Gauge, Registry } from "prom-client"; const registry = new Registry(); export const metrics = { tracesCreated: new Counter({ name: "langfuse_traces_created_total", help: "Total traces created", labelNames: ["status"], registers: [registry], }), generationDuration: new Histogram({ name: "langfuse_generation_duration_seconds", help: "LLM generation latency", labelNames: ["model"], buckets: [0.1, 0.5, 1, 2, 5, 10, 30], registers: [registry], }), tokensUsed: new Counter({ name: "langfuse_tokens_total", help: "Total tokens used", labelNames: ["model", "type"], registers: [registry], }), costUsd: new Counter({ name: "langfuse_cost_usd_total", help: "Total LLM cost in USD", labelNames: ["model"], registers: [registry], }), flushErrors: new Counter({ name: "langfuse_flush_errors_total", help: "Total flush/export errors", registers: [registry], }), }; export { registry }; ``` ```typescript // src/lib/traced-llm.ts -- Instrumented LLM wrapper import { observe, updateActiveObservation } from "@langfuse/tracing"; import { metrics } from "./langfuse-metrics"; import OpenAI from "openai"; const openai = new OpenAI(); export const tracedLLM = observe( { name: "llm-call", asType: "generation" }, async (model: string, messages: OpenAI.ChatCompletionMessageParam[]) => { const start = Date.now(); updateActiveObservation({ model, input: messages }); try { const response = await openai.chat.completions.create({ model, messages }); const duration = (Date.now() - start) / 1000; metrics.generationDuration.observe({ model }, duration); metrics.tracesCreated.inc({ status: "success" }); if (response.usage) { metrics.tokensUsed.inc({ model, type: "prompt" }, response.usage.prompt_tokens); metrics.tokensUsed.inc({ model, type: "completion" }, response.usage.completion_tokens); } updateActiveObservation({ output: response.choices[0].message.content, usage: { promptTokens: response.usage?.prompt_tokens, completionTokens: response.usage?.completion_tokens, }, }); return response.choices[0].message.content; } catch (error) { metrics.tracesCreated.inc({ status: "error" }); throw error; } } ); ``` ### Step 3: Expose Metrics Endpoint ```typescript // src/routes/metrics.ts import { registry } from "../lib/langfuse-metrics"; app.get("/metrics", async (req, res) => { res.set("Content-Type", registry.contentType); res.end(await registry.metrics()); }); ``` ### Step 4: Prometheus Scrape Config ```yaml # prometheus.yml scrape_configs: - job_name: "llm-app" scrape_interval: 15s static_configs: - targets: ["llm-app:3000"] ``` ### Step 5: Grafana Dashboard ```json { "panels": [ { "title": "LLM Requests/min", "type": "graph", "targets": [{ "expr": "rate(langfuse_traces_created_total[5m]) * 60" }] }, { "title": "Generation Latency P95", "type": "graph", "targets": [{ "expr": "histogram_quantile(0.95, rate(langfuse_generation_duration_seconds_bucket[5m]))" }] }, { "title": "Cost/Hour", "type": "stat", "targets": [{ "expr": "rate(langfuse_cost_usd_total[1h]) * 3600" }] }, { "title": "Error Rate", "type": "graph", "targets": [{ "exp
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__langfuse-observability.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Langfuse Observability skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Langfuse Observability safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Langfuse Observability access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Langfuse Observability work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.