Langfuse Ci IntegrationSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: langfuse-ci-integration description: 'Configure Langfuse CI/CD integration with GitHub Actions and automated testing. Use when setting up automated testing, configuring CI pipelines, or integrating Langfuse tests into your build process. Trigger with phrases like "langfuse CI", "langfuse GitHub Actions", "langfuse automated tests", "CI langfuse", "langfuse pipeline". ' allowed-tools: Read, Write, Edit, Bash(gh:*) version: 1.17.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - langfuse - testing - ci-cd compatibility: Designed for Claude Code --- # Langfuse CI Integration ## Overview Integrate Langfuse into CI/CD pipelines: trace validation tests, prompt regression testing, experiment-driven quality gates, automated prompt deployment from version control, and score monitoring. ## Prerequisites - Langfuse API keys stored as GitHub secrets (`LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY`) - Test framework (Vitest or Jest) - OpenAI API key for LLM tests ## Instructions ### Step 1: GitHub Actions Workflow for AI Quality Tests ```yaml # .github/workflows/langfuse-tests.yml name: AI Quality Tests on: pull_request: paths: ["src/ai/**", "src/prompts/**", "tests/ai/**"] jobs: ai-quality: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: { node-version: "20", cache: "npm" } - run: npm ci - name: Run AI quality tests with tracing env: LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }} LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }} LANGFUSE_BASE_URL: ${{ vars.LANGFUSE_BASE_URL || 'https://cloud.langfuse.com' }} OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }} run: npx vitest run tests/ai/ --reporter=verbose - name: Langfuse connectivity check env: LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }} LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }} run: | node -e " const { LangfuseClient } = require('@langfuse/client'); const lf = new LangfuseClient(); lf.prompt.get('__ci-health__').catch(() => {}); console.log('Langfuse SDK initialized OK'); " ``` ### Step 2: Prompt Regression Tests ```typescript // tests/ai/prompt-quality.test.ts import { describe, it, expect, afterAll } from "vitest"; import { LangfuseClient } from "@langfuse/client"; import { startActiveObservation, updateActiveObservation } from "@langfuse/tracing"; import OpenAI from "openai"; const langfuse = new LangfuseClient(); const openai = new OpenAI(); describe("Prompt Quality Regression", () => { it("summarization prompt produces valid output", async () => { const prompt = await langfuse.prompt.get("summarize-article", { type: "text" }); const compiled = prompt.compile({ maxLength: "100 words" }); const result = await startActiveObservation( { name: "ci-test-summarize", asType: "generation" }, async () => { updateActiveObservation({ model: "gpt-4o-mini", input: compiled }); const response = await openai.chat.completions.create({ model: "gpt-4o-mini", messages: [{ role: "user", content: compiled }], temperature: 0, }); const output = response.choices[0].message.content || ""; updateActiveObservation({ output, usage: { promptTokens: response.usage?.prompt_tokens, completionTokens: response.usage?.completion_tokens, }, }); return output; } ); expect(result.length).toBeGreaterThan(20); expect(result.length).toBeLessThan(600); }); it("classification prompt returns valid intent", async () => { const prompt = await langfuse.prompt.get("classify-intent", { type: "text" }); const compiled = prompt.compile({ userMessage: "I want to cancel my subscription" }); const response = await openai.chat.completions.create({ model: "gpt-4o-mini", messages: [{ role: "user", content: compiled }], temperature: 0, }); const intent = response.choices[0].message.content?.trim().toLowerCase() || ""; const validIntents = ["billing", "cancellation", "support", "feedback"]; expect(validIntents).toContain(intent); }); }); ``` ### Step 3: Experiment-Driven Quality Gates ```typescript // tests/ai/experiment-gate.test.ts import { describe, it, expect } from "vitest"; import { LangfuseClient } from "@langfuse/client"; import OpenAI from "openai"; const langfuse = new LangfuseClient(); const openai = new OpenAI(); describe("Quality Gate: Intent Classification", () => { it("scores above 80% accuracy on test dataset", async () => { async function classifyIntent(input: { query: string }) { const response = await openai.chat.completions.create({ model: "gpt-4o-mini", messages: [ { role: "system", content: "Classify intent. Return one word." }, { role: "user", content: input.query }, ], temperature: 0, }); return response.choices[0].message.content?.trim() || ""; } const result = await langfuse.runExperiment({ datasetName: "intent-classification-test", runName: `ci-${process.env.GITHUB_SHA?.slice(0, 7) || "local"}`, task: classifyIntent, evaluators: [ ({ output, expectedOutput }) => ({ name: "exact-match", value: output.toLowerCase() === expectedOutput.intent.toLowerCase() ? 1 : 0, dataType: "BOOLEAN" as const, }), ], }); // Calculate accuracy const scores = result.runs.flatMap((r) => r.scores || []); const accuracy = scores.filter((s) => s.value === 1).length / scores.length; console.log(`Accuracy: ${(accuracy * 100).toFixed(1)}%`); expect(accuracy).toBeGreaterThanOrEqual(0.8); }); }); ``` ### Step 4: Au
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__langfuse-ci-integration.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Langfuse Ci Integration skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Langfuse Ci Integration safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Langfuse Ci Integration access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Langfuse Ci Integration work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.