Langfuse Cost TuningSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: langfuse-cost-tuning description: 'Monitor and optimize LLM costs using Langfuse analytics and dashboards. Use when tracking LLM spending, identifying cost anomalies, or implementing cost controls for AI applications. Trigger with phrases like "langfuse costs", "LLM spending", "track AI costs", "langfuse token usage", "optimize LLM budget". ' allowed-tools: Read, Write, Edit version: 1.17.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - langfuse - monitoring - llm - analytics compatibility: Designed for Claude Code --- # Langfuse Cost Tuning ## Overview Track, analyze, and optimize LLM costs using Langfuse's built-in token/cost tracking, the Metrics API for programmatic cost analysis, model routing for cost reduction, and automated budget alerts. ## Prerequisites - Langfuse tracing with token usage captured (via `observeOpenAI` or manual `usage` fields) - For Metrics API: `@langfuse/client` installed - Understanding of LLM pricing models ## How Langfuse Tracks Costs Langfuse automatically calculates costs for supported models (OpenAI, Anthropic, Google) when token usage is captured. For custom models, you can configure pricing in the Langfuse UI under **Settings > Model Definitions**. Cost tracking works on observations of type **generation** and **embedding**. The `observeOpenAI` wrapper captures usage automatically; for manual tracing, include `usage` in your observation updates. ## Instructions ### Step 1: Ensure Token Usage is Captured ```typescript // Automatic: observeOpenAI captures everything import { observeOpenAI } from "@langfuse/openai"; const openai = observeOpenAI(new OpenAI()); // Tokens, model, latency, and cost are all auto-tracked // Manual: include usage in generation observations import { startActiveObservation, updateActiveObservation } from "@langfuse/tracing"; await startActiveObservation( { name: "llm-call", asType: "generation" }, async () => { updateActiveObservation({ model: "gpt-4o" }); // Model required for cost calc const response = await openai.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: prompt }], }); updateActiveObservation({ output: response.choices[0].message.content, usage: { promptTokens: response.usage?.prompt_tokens, completionTokens: response.usage?.completion_tokens, totalTokens: response.usage?.total_tokens, }, // Optional: override inferred cost (in USD) // costInUsd: 0.0015, }); } ); ``` ### Step 2: Query Costs via Metrics API ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Fetch aggregated cost metrics async function getCostReport(days: number) { const fromTimestamp = new Date(Date.now() - days * 86400000).toISOString(); // Use the API to list traces with cost data const traces = await langfuse.api.traces.list({ fromTimestamp, limit: 1000, orderBy: "timestamp", }); const costByModel = new Map<string, { cost: number; tokens: number; count: number }>(); for (const trace of traces.data) { const observations = await langfuse.api.observations.list({ traceId: trace.id, type: "GENERATION", }); for (const obs of observations.data) { const model = obs.model || "unknown"; const existing = costByModel.get(model) || { cost: 0, tokens: 0, count: 0 }; existing.cost += obs.calculatedTotalCost || 0; existing.tokens += obs.totalTokens || 0; existing.count += 1; costByModel.set(model, existing); } } console.log("\n=== LLM Cost Report ==="); console.log(`Period: Last ${days} days\n`); let totalCost = 0; for (const [model, data] of costByModel.entries()) { console.log(`${model}:`); console.log(` Calls: ${data.count}`); console.log(` Tokens: ${data.tokens.toLocaleString()}`); console.log(` Cost: $${data.cost.toFixed(4)}`); totalCost += data.cost; } console.log(`\nTotal: $${totalCost.toFixed(4)}`); } getCostReport(7); ``` ### Step 3: Implement Smart Model Routing Route requests to cheaper models when appropriate: ```typescript import { observe, updateActiveObservation } from "@langfuse/tracing"; interface ModelConfig { model: string; costPer1MInput: number; costPer1MOutput: number; maxComplexity: "simple" | "moderate" | "complex"; } const MODELS: ModelConfig[] = [ { model: "gpt-4o-mini", costPer1MInput: 0.15, costPer1MOutput: 0.60, maxComplexity: "simple" }, { model: "gpt-4o", costPer1MInput: 2.50, costPer1MOutput: 10.00, maxComplexity: "moderate" }, { model: "claude-sonnet-4-20250514", costPer1MInput: 3.00, costPer1MOutput: 15.00, maxComplexity: "complex" }, ]; function selectModel(task: string, inputLength: number): ModelConfig { const simpleTasks = ["classify", "extract", "summarize-short", "translate"]; const isSimple = simpleTasks.some((t) => task.includes(t)); const isShort = inputLength < 500; if (isSimple && isShort) return MODELS[0]; // gpt-4o-mini if (isSimple || inputLength < 2000) return MODELS[1]; // gpt-4o return MODELS[2]; // claude-sonnet-4 } const costOptimizedLLM = observe( { name: "cost-optimized-llm", asType: "generation" }, async (task: string, input: string) => { const config = selectModel(task, input.length); updateActiveObservation({ model: config.model, metadata: { task, selectedReason: `${config.maxComplexity} tier`, estimatedCostPer1M: config.costPer1MInput, }, }); const response = await callModel(config.model, input); updateActiveObservation({ output: response.content, usage: response.usage, }); return response; } ); ``` ### Step 4: Budget Alerts ```typescript // scripts/cost-alert.ts -- run as cron job import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); const ALERT_THRESHOLDS = { dai
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__langfuse-cost-tuning.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Langfuse Cost Tuning skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Langfuse Cost Tuning safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Langfuse Cost Tuning access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Langfuse Cost Tuning work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.