Langfuse Core Workflow BSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: langfuse-core-workflow-b description: 'Execute Langfuse secondary workflow: Evaluation, scoring, and datasets. Use when implementing LLM evaluation, adding user feedback, or setting up automated quality scoring and experiment datasets. Trigger with phrases like "langfuse evaluation", "langfuse scoring", "rate llm outputs", "langfuse feedback", "langfuse datasets", "langfuse experiments". ' allowed-tools: Read, Write, Edit, Bash(npm:*), Grep version: 1.17.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - langfuse - llm - workflow - evaluation compatibility: Designed for Claude Code --- # Langfuse Core Workflow B: Evaluation, Scoring & Datasets ## Overview Implement LLM output evaluation using Langfuse scores (numeric, categorical, boolean), the experiment runner SDK for dataset-driven benchmarks, prompt management with versioned prompts, and LLM-as-a-Judge evaluation patterns. ## Prerequisites - Langfuse SDK configured with API keys - Traces already being collected (see `langfuse-core-workflow-a`) - For v4+: `@langfuse/client` installed ## Instructions ### Step 1: Score Traces via SDK Langfuse supports three score data types: **Numeric**, **Categorical**, and **Boolean**. ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Numeric score (e.g., 0-1 quality rating) await langfuse.score.create({ traceId: "trace-abc-123", name: "relevance", value: 0.92, dataType: "NUMERIC", comment: "Highly relevant answer with good context usage", }); // Categorical score (e.g., pass/fail classification) await langfuse.score.create({ traceId: "trace-abc-123", observationId: "gen-xyz-456", // Optional: score a specific generation name: "quality-tier", value: "excellent", dataType: "CATEGORICAL", }); // Boolean score (e.g., thumbs up/down) await langfuse.score.create({ traceId: "trace-abc-123", name: "user-approved", value: 1, // 1 = true, 0 = false dataType: "BOOLEAN", comment: "User clicked thumbs up", }); ``` ### Step 2: User Feedback Collection ```typescript // API endpoint for frontend feedback widget app.post("/api/feedback", async (req, res) => { const { traceId, rating, comment } = req.body; // Thumbs up/down await langfuse.score.create({ traceId, name: "user-feedback", value: rating === "positive" ? 1 : 0, dataType: "BOOLEAN", comment, }); // Granular star rating (1-5) if (req.body.stars) { await langfuse.score.create({ traceId, name: "star-rating", value: req.body.stars, dataType: "NUMERIC", comment: `${req.body.stars}/5 stars`, }); } res.json({ success: true }); }); ``` ### Step 3: Prompt Management ```typescript // Fetch a versioned prompt from Langfuse const textPrompt = await langfuse.prompt.get("summarize-article", { type: "text", label: "production", // or "latest", "staging" }); // Compile with variables -- replaces {{variable}} placeholders const compiled = textPrompt.compile({ maxLength: "100 words", tone: "professional", }); // Chat prompts return message arrays const chatPrompt = await langfuse.prompt.get("customer-support", { type: "chat", }); const messages = chatPrompt.compile({ customerName: "Alice", issue: "billing question", }); // messages = [{ role: "system", content: "..." }, { role: "user", content: "..." }] ``` ### Step 4: Create and Populate Datasets ```typescript // Create a dataset for evaluation await langfuse.api.datasets.create({ name: "customer-support-v1", description: "Test cases for customer support chatbot", metadata: { version: "1.0", domain: "support" }, }); // Add test items const testCases = [ { input: { query: "How do I cancel my subscription?" }, expectedOutput: { intent: "cancellation", sentiment: "neutral" }, metadata: { category: "billing" }, }, { input: { query: "Your product is amazing!" }, expectedOutput: { intent: "feedback", sentiment: "positive" }, metadata: { category: "feedback" }, }, ]; for (const testCase of testCases) { await langfuse.api.datasetItems.create({ datasetName: "customer-support-v1", input: testCase.input, expectedOutput: testCase.expectedOutput, metadata: testCase.metadata, }); } ``` ### Step 5: Run Experiments with the Experiment Runner ```typescript import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Define the task function -- your LLM application logic async function classifyIntent(input: { query: string }): Promise<string> { const response = await openai.chat.completions.create({ model: "gpt-4o-mini", messages: [ { role: "system", content: "Classify the user intent. Return one word." }, { role: "user", content: input.query }, ], temperature: 0, }); return response.choices[0].message.content?.trim() || ""; } // Define evaluator functions function exactMatch({ output, expectedOutput }: { output: string; expectedOutput: { intent: string }; }) { return { name: "exact-match", value: output.toLowerCase() === expectedOutput.intent.toLowerCase() ? 1 : 0, dataType: "BOOLEAN" as const, }; } // Run the experiment const result = await langfuse.runExperiment({ datasetName: "customer-support-v1", runName: "gpt-4o-mini-classifier-v1", runDescription: "Testing intent classification with gpt-4o-mini", task: classifyIntent, evaluators: [exactMatch], }); console.log(`Experiment complete. ${result.runs.length} items evaluated.`); // View results in Langfuse UI: Datasets > customer-support-v1 > Runs ``` ### Step 6: LLM-as-a-Judge Evaluation ```typescript async function llmJudge({ output, input, expectedOutput }: { output: string; input: { query: string }; expectedOutput: { intent: string; sentiment: string }; }) { const judgment = await openai.chat.completions.create({ model: "gpt-4o", temperature: 0, messages:
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__langfuse-core-workflow-b.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Langfuse Core Workflow B skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Langfuse Core Workflow B safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Langfuse Core Workflow B access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Langfuse Core Workflow B work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.