Groq Core Workflow BSAFE
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Overview
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
4f83675ca38aOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: groq-core-workflow-b description: | Use when you need Groq's non-chat endpoints — transcribing or translating audio with Whisper, understanding images with Llama 4 vision, generating speech (TTS), or benchmarking models for speed vs quality. Trigger with phrases like "groq whisper", "groq transcription", "groq audio", "groq vision", "groq TTS", "groq speech". allowed-tools: Read, Bash(npm:*) version: 1.11.0 license: MIT author: Jeremy Longshore <[email protected]> tags: - saas - groq - workflow - audio - vision compatibility: Designed for Claude Code --- # Groq Core Workflow B: Audio, Vision & Speech ## Overview Beyond chat completions, Groq offers ultra-fast Whisper transcription (216x real-time), Llama 4 vision, and text-to-speech — all on the same `groq-sdk` client. This skill covers transcription/translation, vision, TTS, and model benchmarking, with full runnable code in [references/implementation.md](references/implementation.md) and worked scripts in [references/examples.md](references/examples.md). ## Prerequisites - `groq-sdk` installed, `GROQ_API_KEY` set (the SDK reads it from the environment automatically) - For audio: audio files in a supported format - For vision: image URLs or base64-encoded images ## Audio Models | Model ID | Languages | Speed | Best For | |----------|-----------|-------|----------| | `whisper-large-v3` | 100+ | 164x real-time | Best accuracy, multilingual | | `whisper-large-v3-turbo` | 100+ | 216x real-time | Best speed/accuracy balance | **Supported audio formats**: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm ## Instructions Each workflow is a single SDK call on the shared `groq` client. Pick the endpoint for your task, then follow the full walkthrough in [references/implementation.md](references/implementation.md) for the complete, copy-pasteable version of each. 1. **Transcription** — `groq.audio.transcriptions.create({ file, model: "whisper-large-v3-turbo", response_format })`. Use `response_format: "verbose_json"` with `timestamp_granularities: ["segment"]` to get per-segment start/end times. 2. **Translation** — `groq.audio.translations.create({ file, model: "whisper-large-v3" })` transcribes any-language audio directly to English text. 3. **Vision** — a normal `groq.chat.completions.create` call where `content` is an array mixing `{ type: "text" }` and `{ type: "image_url" }` parts. Accepts up to 5 images (URL or `data:` base64) with `meta-llama/llama-4-scout-17b-16e-instruct`. 4. **Text-to-Speech** — `groq.audio.speech.create({ model: "playai-tts", input, voice, response_format })`, then write `Buffer.from(await response.arrayBuffer())` to a file. 5. **Benchmarking** — loop a prompt across several chat models and time each call to compare latency and tokens/sec (see [references/examples.md](references/examples.md)). Minimal transcription skeleton: ```typescript import Groq from "groq-sdk"; import fs from "fs"; const groq = new Groq(); async function transcribe(filePath: string): Promise<string> { const transcription = await groq.audio.transcriptions.create({ file: fs.createReadStream(filePath), model: "whisper-large-v3-turbo", response_format: "json", }); return transcription.text; } ``` ## Output - **Transcription/translation**: a `transcription.text` string. With `verbose_json`, a `segments[]` array where each segment has `start`, `end`, and `text`. - **Vision**: the assistant reply at `completion.choices[0].message.content` (a natural-language answer about the image(s)). - **Text-to-Speech**: an audio response you convert to a `Buffer` and write to disk (`wav`, `mp3`, `flac`, `opus`, or `aac`). - **Benchmarking**: one console line per model — latency in ms, throughput in tok/s, and total tokens. ## Vision Model Limits - Maximum 5 images per request - Supported formats: JPEG, PNG, GIF, WebP - Images fetched from URL or embedded as base64 - Vision models also support tool use, JSON mode, and streaming ## Error Handling | Error | Cause | Solution | |-------|-------|----------| | `Invalid file format` | Unsupported audio type | Convert to mp3/wav/flac first | | `File too large` | Audio exceeds 25MB | Split into smaller chunks | | `model_not_found` | Vision model ID wrong | Use full path: `meta-llama/llama-4-scout-17b-16e-instruct` | | `max_images_exceeded` | >5 images in request | Reduce to 5 or fewer images | | `429` on Whisper | Audio RPM limit hit | Queue transcription requests | ## Examples Complete, runnable scripts live in [references/examples.md](references/examples.md): - **Python transcription with timestamps** — transcribe a local MP3 and print each segment with its start/end time. - **Model benchmarking** — run one prompt across `llama-3.1-8b-instant`, `llama-3.3-70b-versatile`, and `llama-3.3-70b-specdec` and print latency + throughput per model. Quick vision example (analyze one image by URL): ```typescript const completion = await groq.chat.completions.create({ model: "meta-llama/llama-4-scout-17b-16e-instruct", messages: [{ role: "user", content: [ { type: "text", text: "What is in this image?" }, { type: "image_url", image_url: { url: imageUrl } }, ], }], max_tokens: 1024, }); console.log(completion.choices[0].message.content); ``` ## Resources - [Groq Speech-to-Text](https://console.groq.com/docs/speech-to-text) - [Groq Text-to-Speech](https://console.groq.com/docs/text-to-speech) - [Groq Vision](https://console.groq.com/docs/vision) - [Groq Models](https://console.groq.com/docs/models) ## Next Steps For common errors and troubleshooting patterns across all Groq workflows, see the `groq-common-errors` skill. For chat completions, streaming, tool use, and JSON mode, see `groq-core-workflow-a`.
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4f83675ca38afull audit observations/trust-audit/skill/jeremylongshore__groq-core-workflow-b.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 4f83675ca38a | SAFE | B | 89 | first audit |
Questions
What does the Groq Core Workflow B skill do?
Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com.
Is Groq Core Workflow B safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Groq Core Workflow B access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Groq Core Workflow B work with?
Its documentation mentions claude-code. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (4f83675ca38a), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.