Langfuse Rate LimitsSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-09Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: langfuse-rate-limits description: 'Implement Langfuse rate limiting, batching, and backoff patterns. Use when handling rate limit errors, optimizing trace ingestion, or managing high-volume LLM observability workloads. Trigger with phrases like "langfuse rate limit", "langfuse throttling", "langfuse 429", "langfuse batching", "langfuse high volume". ' allowed-tools: Read, Write, Edit version: 1.17.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - langfuse - observability - llm compatibility: Designed for Claude Code --- # Langfuse Rate Limits ## Overview Handle Langfuse API rate limits with optimized SDK batching, exponential backoff with jitter, concurrent request limiting, and configurable sampling for ultra-high-volume workloads. ## Prerequisites - Langfuse SDK installed and configured - High-volume trace workload (1,000+ events/minute) ## Instructions ### Step 1: Optimize SDK Batching Configuration The Langfuse SDK batches events internally before sending. Tuning batch settings is the first defense against rate limits. ```typescript // v3 Legacy: Direct configuration import { Langfuse } from "langfuse"; const langfuse = new Langfuse({ flushAt: 50, // Events per batch (default: 15, max ~200) flushInterval: 10000, // Milliseconds between flushes (default: 10000) requestTimeout: 30000, // Timeout per batch request }); // v4+: Configure via OTel span processor import { LangfuseSpanProcessor } from "@langfuse/otel"; import { NodeSDK } from "@opentelemetry/sdk-node"; const processor = new LangfuseSpanProcessor({ exportIntervalMillis: 10000, // Flush interval maxExportBatchSize: 50, // Events per batch }); const sdk = new NodeSDK({ spanProcessors: [processor] }); sdk.start(); ``` ### Step 2: Implement Retry with Exponential Backoff For custom API calls (scores, datasets, prompts) that hit rate limits: ```typescript async function withRetry<T>( fn: () => Promise<T>, options: { maxRetries?: number; baseDelayMs?: number; maxDelayMs?: number } = {} ): Promise<T> { const { maxRetries = 5, baseDelayMs = 1000, maxDelayMs = 30000 } = options; for (let attempt = 0; attempt <= maxRetries; attempt++) { try { return await fn(); } catch (error: any) { const status = error?.status || error?.response?.status; // Only retry on rate limits (429) and server errors (5xx) if (attempt === maxRetries || (status && status < 429)) { throw error; } // Honor Retry-After header if present const retryAfter = error?.response?.headers?.["retry-after"]; let delay: number; if (retryAfter) { delay = parseInt(retryAfter, 10) * 1000; } else { // Exponential backoff with jitter delay = Math.min(baseDelayMs * Math.pow(2, attempt), maxDelayMs); delay += Math.random() * 500; // Jitter } console.warn(`Rate limited. Retry ${attempt + 1}/${maxRetries} in ${Math.round(delay)}ms`); await new Promise((r) => setTimeout(r, delay)); } } throw new Error("Unreachable"); } // Usage with Langfuse client operations const langfuse = new LangfuseClient(); await withRetry(() => langfuse.score.create({ traceId: "trace-123", name: "quality", value: 0.95, dataType: "NUMERIC", }) ); ``` ### Step 3: Queue-Based Concurrency Limiting Use `p-queue` to cap concurrent Langfuse API calls: ```typescript import PQueue from "p-queue"; import { LangfuseClient } from "@langfuse/client"; const langfuse = new LangfuseClient(); // Max 10 concurrent API calls, 50 per second const queue = new PQueue({ concurrency: 10, interval: 1000, intervalCap: 50, }); // Queue score submissions async function queueScore(params: { traceId: string; name: string; value: number; }) { return queue.add(() => langfuse.score.create({ ...params, dataType: "NUMERIC", }) ); } // Queue dataset item creation async function queueDatasetItem(datasetName: string, item: any) { return queue.add(() => langfuse.api.datasetItems.create({ datasetName, input: item.input, expectedOutput: item.expectedOutput, }) ); } // Monitor queue health setInterval(() => { console.log(`Queue: ${queue.pending} pending, ${queue.size} queued`); }, 10000); ``` ### Step 4: Configurable Sampling for Ultra-High Volume When tracing volume exceeds rate limits, sample traces instead of dropping them: ```typescript import { observe, updateActiveObservation, startActiveObservation } from "@langfuse/tracing"; class TraceSampler { private rate: number; private windowCounts: number[] = []; private windowMs = 60000; // 1 minute window private maxPerWindow: number; constructor(sampleRate: number, maxPerMinute: number) { this.rate = sampleRate; this.maxPerWindow = maxPerMinute; } shouldSample(tags?: string[]): boolean { // Always sample errors if (tags?.includes("error") || tags?.includes("critical")) { return true; } // Check window limit const now = Date.now(); this.windowCounts = this.windowCounts.filter((t) => t > now - this.windowMs); if (this.windowCounts.length >= this.maxPerWindow) { return false; } // Probabilistic sampling if (Math.random() > this.rate) { return false; } this.windowCounts.push(now); return true; } } // 10% sampling, max 1000 traces/minute const sampler = new TraceSampler(0.1, 1000); async function sampledOperation(name: string, fn: () => Promise<any>) { if (!sampler.shouldSample()) { return fn(); // Run without tracing } return startActiveObservation(name, async () => { updateActiveObservation({ metadata: { sampled: true } }); return fn(); }); } ``` ## Rate Limit Reference | Tier | Traces/min | Batch Size | Strategy | |------|------------|------------|----------| | Hobby | ~500 | 15 | Default settings | | Pro | ~5,000 | 50 | Increase `flushAt` |
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__langfuse-rate-limits.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Langfuse Rate Limits skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Langfuse Rate Limits safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Langfuse Rate Limits access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Langfuse Rate Limits work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.