Atlas / MCP servers / arcaputo3 / Whisper

WhisperSAFE

mcp/arcaputo3/whisper

An MCP Server for audio transcription using OpenAI

Verdict
SAFE
Grade
B
Trust score
89 /100
Exposed tools
5 4r · 1w · 0d
Transport
—
License
MIT
Stars
60
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

A Model Context Protocol (MCP) server for advanced audio transcription and processing using OpenAI's Whisper and GPT-4o models.

[](https://pypi.org/project/mcp-server-whisper/) [](https://opensource.org/licenses/MIT) [](https://www.python.org/downloads/) [](https://github.com/astral-sh/uv)

[!WARNING] This project has moved. Active development has migrated to [TJC-LP/sanzaru](https://github.com/TJC-LP/sanzaru). This repository is no longer maintained and will be archived. Please update your dependencies and issues to the new repo.

Overview

MCP Server Whisper provides a standardized way to process audio files through OpenAI's latest transcription and speech services. By implementing the Model Context Protocol, it enables AI assistants like Claude to seamlessly interact with audio processing capabilities.

Key features:

  • 🔍 Advanced file searching with regex patterns, file metadata filtering, and sorting capabilities
  • ⚡ MCP-native parallel processing - call multiple tools simultaneously
  • 🔄 Format conversion between supported audio types
  • 📦 Automatic compression for oversized files
  • 🎯 Multi-model transcription with support for all OpenAI audio models
  • 🗣️ Interactive audio chat with GPT-4o audio models
  • ✏️ Enhanced transcription with specialized prompts and timestamp support
  • 🎙️ Text-to-speech generation with customizable voices, instructions, and speed
  • 📊 Comprehensive metadata including duration, file size, and form
Read from source at commit 5a2467f78ed7OBSERVED · 2026-10-08
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control.

claude-code
claude mcp add mcp-server-whisper -- uvx mcp-server-whisper
claude-desktop
{
  "mcpServers": {
    "mcp-server-whisper": {
      "command": "uvx",
      "args": [
        "mcp-server-whisper"
      ]
    }
  }
}
03

Exposed tools (5)

4 read · 1 write · 0 destructive.

ToolRiskDescription
compress_audioreadCompress audio file if it
convert_audioreadConvert audio file to supported format (mp3 or wav).
create_audiowriteGenerate text-to-speech audio from text prompts with customizable voices.
get_latest_audioreadGet the most recently modified audio file and returns its path with model support info.
transcribe_with_enhancementreadTranscribe audio with GPT-4 using specific enhancement prompts.
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeWARN
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
found

Findings (11)

MEDIUMHard-coded secrets · secret.generic · CWE-798, CWE-321
tests/test_openai_client.py:27
return OpenAIClientWrapper(api_key="sk-test-fake-key-for-testing")
LOWInventory / provenance · inv.hidden_file · CWE-1104
.pre-commit-config.yaml
.pre-commit-config.yaml
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
tests/test_path_resolver.py:66
resolver.resolve_input("../../../etc/passwd")
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/mcp-overview.md:140
A few notes on efficiency: if performance is a concern or if you plan to transcribe very large files, you could enhance this function by splitting audio into chunks (using `AudioSegment[:segment_durat
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/mcp-overview.md:158
- **Transcription Result State:** The transcript is the key piece of state that must carry forward. Once the tool returns the text, the **MCP client (AI)** will include that text in the conversation.
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/mcp-overview.md:168
- **Direct GPT-4o Audio Input:** If using GPT-4o (the “omni” model that can handle audio natively ([Hello GPT-4o - OpenAI](https://openai.com/index/hello-gpt-4o/#:~:text=Hello%20GPT,and%20text%20in%20
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/mcp-overview.md:182
**Synchronizing State**: If you have multiple components – say one component recording audio, another transcribing, another handling the chat logic – you need to sync their operation. A common approac
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/mcp-overview.md:187
In any multi-step pipeline, errors can occur at various points: microphone failure, file read error, transcription misrecognition, API error from GPT, etc. **Error recovery** means handling the error
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine
INFOPrompt injection · prompt.authority_framing · CWE-94, CWE-1427
docs/mcp-readme.md:365
For more control, you can use the low-level server implementation directly. This gives you full access to the protocol and allows you to customize every aspect of your server, including lifecycle mana
INFOPrompt injection · prompt.credential_read · CWE-94, CWE-1427
README.md:82
- Load environment variables from your `.env` file
Why it matters. asks the agent to read credentials
INFOPrompt injection · prompt.credential_read · CWE-94, CWE-1427
docs/mcp-overview.md:293
- **Environment and Config:** Use `.env` or environment variables to load API keys in your app. In tests, you can either set dummy keys (since you won’t actually call the API if mocking) or ensure the
Why it matters. asks the agent to read credentials

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 5a2467f78ed7full audit observations/trust-audit/mcp-server/arcaputo3__whisper.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-085a2467f78ed7SAFEB89first audit
06

Questions

What is the Whisper MCP server?

An MCP Server for audio transcription using OpenAI

What tools does Whisper expose?

5 in total: 4 read-only, 1 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.

Is Whisper safe to connect to an agent?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.

What credentials does Whisper need?

No credential environment variables were found in its source, so it appears to need none.

How current is this page?

The grade is for one exact copy of the source (5a2467f78ed7), read on 2026-10-08. The repository is watched and re-audited when it changes.

Advertisement