SpeechCAUTION
Speech MCP: A Goose MCP extension for voice interaction with audio visualization
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
A Goose MCP extension for voice interaction with modern audio visualization.
https://github.com/user-attachments/assets/f10f29d9-8444-43fb-a919-c80b9e0a12c8
Overview
Speech MCP provides a voice interface for Goose, allowing users to interact through speech rather than text. It includes:
- Real-time audio processing for speech recognition
- Local speech-to-text using faster-whisper (a faster implementation of OpenAI's Whisper model)
- High-quality text-to-speech with multiple voice options
- Modern PyQt-based UI with audio visualization
- Simple command-line interface for voice interaction
Features
- Modern UI: Sleek PyQt-based interface with audio visualization and dark theme
- Voice Input: Capture and transcribe user speech using faster-whisper
- Voice Output: Convert agent responses to speech with 54+ voice options
- Multi-Speaker Narration: Generate audio files with multiple voices for stories and dialogues
- Single-Voice Narration: Convert any text to speech with your preferred voice
- Audio/Video Transcription: Transcribe speech from various media formats with optional timestamps and speaker detection
- Voice Persistence: Remembers your preferred voice between sessions
- Continuous Conversation: Automatically listen for user input after agent responses
- Silence Detection: Automatically stops recording when the user stops speaking
- Robust Error Handling: Graceful recovery from common failure modes with helpful voice suggestions
Installation
Important Note: After installation, the first time you use the speech interface, it may take several minutes to download the Kokoro voice models (approximately 523 KB per voice). During this initial setup period, the system will use a more robotic-sounding fallback voice. Once the Kokoro voices are downloaded, the high-quality voices will be used automatically.
⚠️ IMPORTANT PREREQUISITES ⚠️
Before installing
dce2018e7b2bOBSERVED · 2026-10-07Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control.
claude mcp add speech-mcp -- uvx speech-mcp
{
"mcpServers": {
"speech-mcp": {
"command": "uvx",
"args": [
"speech-mcp"
]
}
}
}Exposed tools (7)
6 read · 1 write · 0 destructive.
| Tool | Risk | Description |
|---|---|---|
close_ui | read | |
launch_ui | read | |
narrate | read | |
narrate_conversation | read | |
reply | read | |
start_conversation | write | |
transcribe | read |
Trust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (5 observation(s))
- Network
- none-observed
- Shell
- declared (3 observation(s))
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (3)
start_listening.wav
stop_listening.wav
.goosehints
Gates applied: no_behavioural_pass.
dce2018e7b2bfull audit observations/trust-audit/mcp-server/kvadratni__speech.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | dce2018e7b2b | CAUTION | B | 89 | first audit |
Questions
What is the Speech MCP server?
Speech MCP: A Goose MCP extension for voice interaction with audio visualization
What tools does Speech expose?
7 in total: 6 read-only, 1 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.
Is Speech safe to connect to an agent?
With care. The audit graded it B (89/100) and found 3 things worth knowing before you trust this server, listed below with the exact line each was found on.
What credentials does Speech need?
No credential environment variables were found in its source, so it appears to need none.
How current is this page?
The grade is for one exact copy of the source (dce2018e7b2b), read on 2026-10-07. The repository is watched and re-audited when it changes.