WhisperSAFE
An MCP Server for audio transcription using OpenAI
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
A Model Context Protocol (MCP) server for advanced audio transcription and processing using OpenAI's Whisper and GPT-4o models.
[](https://pypi.org/project/mcp-server-whisper/) [](https://opensource.org/licenses/MIT) [](https://www.python.org/downloads/) [](https://github.com/astral-sh/uv)
[!WARNING] This project has moved. Active development has migrated to [TJC-LP/sanzaru](https://github.com/TJC-LP/sanzaru). This repository is no longer maintained and will be archived. Please update your dependencies and issues to the new repo.
Overview
MCP Server Whisper provides a standardized way to process audio files through OpenAI's latest transcription and speech services. By implementing the Model Context Protocol, it enables AI assistants like Claude to seamlessly interact with audio processing capabilities.
Key features:
- 🔍 Advanced file searching with regex patterns, file metadata filtering, and sorting capabilities
- ⚡ MCP-native parallel processing - call multiple tools simultaneously
- 🔄 Format conversion between supported audio types
- 📦 Automatic compression for oversized files
- 🎯 Multi-model transcription with support for all OpenAI audio models
- 🗣️ Interactive audio chat with GPT-4o audio models
- ✏️ Enhanced transcription with specialized prompts and timestamp support
- 🎙️ Text-to-speech generation with customizable voices, instructions, and speed
- 📊 Comprehensive metadata including duration, file size, and form
5a2467f78ed7OBSERVED · 2026-10-08Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control.
claude mcp add mcp-server-whisper -- uvx mcp-server-whisper
{
"mcpServers": {
"mcp-server-whisper": {
"command": "uvx",
"args": [
"mcp-server-whisper"
]
}
}
}Exposed tools (5)
4 read · 1 write · 0 destructive.
| Tool | Risk | Description |
|---|---|---|
compress_audio | read | Compress audio file if it |
convert_audio | read | Convert audio file to supported format (mp3 or wav). |
create_audio | write | Generate text-to-speech audio from text prompts with customizable voices. |
get_latest_audio | read | Get the most recently modified audio file and returns its path with model support info. |
transcribe_with_enhancement | read | Transcribe audio with GPT-4 using specific enhancement prompts. |
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | WARN |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- found
Findings (11)
return OpenAIClientWrapper(api_key="sk-test-fake-key-for-testing")
.pre-commit-config.yaml
resolver.resolve_input("../../../etc/passwd")A few notes on efficiency: if performance is a concern or if you plan to transcribe very large files, you could enhance this function by splitting audio into chunks (using `AudioSegment[:segment_durat
- **Transcription Result State:** The transcript is the key piece of state that must carry forward. Once the tool returns the text, the **MCP client (AI)** will include that text in the conversation.
- **Direct GPT-4o Audio Input:** If using GPT-4o (the “omni” model that can handle audio natively ([Hello GPT-4o - OpenAI](https://openai.com/index/hello-gpt-4o/#:~:text=Hello%20GPT,and%20text%20in%20
**Synchronizing State**: If you have multiple components – say one component recording audio, another transcribing, another handling the chat logic – you need to sync their operation. A common approac
In any multi-step pipeline, errors can occur at various points: microphone failure, file read error, transcription misrecognition, API error from GPT, etc. **Error recovery** means handling the error
For more control, you can use the low-level server implementation directly. This gives you full access to the protocol and allows you to customize every aspect of your server, including lifecycle mana
- Load environment variables from your `.env` file
- **Environment and Config:** Use `.env` or environment variables to load API keys in your app. In tests, you can either set dummy keys (since you won’t actually call the API if mocking) or ensure the
Gates applied: no_behavioural_pass.
5a2467f78ed7full audit observations/trust-audit/mcp-server/arcaputo3__whisper.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 5a2467f78ed7 | SAFE | B | 89 | first audit |
Questions
What is the Whisper MCP server?
An MCP Server for audio transcription using OpenAI
What tools does Whisper expose?
5 in total: 4 read-only, 1 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.
Is Whisper safe to connect to an agent?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.
What credentials does Whisper need?
No credential environment variables were found in its source, so it appears to need none.
How current is this page?
The grade is for one exact copy of the source (5a2467f78ed7), read on 2026-10-08. The repository is watched and re-audited when it changes.