Rubber DuckSAFE
An MCP server that acts as a bridge to query multiple OpenAI-compatible LLMs with MCP tool access. Just like rubber duck debugging, explain your problems to various AI "ducks" who can actually research and get different perspectives!
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
An MCP (Model Context Protocol) server that acts as a bridge to query multiple LLMs -- both OpenAI-compatible HTTP APIs and CLI coding agents. Just like rubber duck debugging, explain your problems to various AI "ducks" and get different perspectives!
[](https://www.npmjs.com/package/mcp-rubber-duck) [](https://github.com/nesquikm/mcp-rubber-duck/pkgs/container/mcp-rubber-duck) [](https://registry.modelcontextprotocol.io)
Why direct provider integration? MCP's sampling primitive -- a server borrowing the host's model -- was deprecated in the 2026-07-28 spec RC in favor of servers integrating directly with LLM provider APIs. Rubber Duck has always worked this way (it brings its own ducks), so it's aligned with where the protocol is heading -- no migration required.Features
- Universal OpenAI Compatibility -- Works with any OpenAI-compatible API endpoint
- CLI Agent Support -- Use CLI coding agents (Claude Code, Codex, Gemini CLI, Grok, Aider) as ducks
- Multiple Ducks -- Configure and query multiple LLM providers simultaneously
- Conversation Management -- Maintain context across multiple messages
- Duck Council -- Get responses from all your configured LLMs at once
- Consensus Voting -- Multi-duck voting with reasoning and confidence scores
- LLM-as-Judge -- Have ducks evaluate and rank each other's responses
- Iterative Refinement -- Two ducks collaboratively improve responses
- Structured Debates -- Oxford, Socratic, and adversarial debate formats
- MCP Prompts -- 8 reusable prompt templates for mul
95dec8c5fee4OBSERVED · 2026-10-06Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add mcp-rubber-duck --env OPENAI_API_KEY=${OPENAI_API_KEY} --env GEMINI_API_KEY=${GEMINI_API_KEY} --env GROQ_API_KEY=${GROQ_API_KEY} -- npx -y [email protected]Exposed tools (45)
43 read · 1 write · 1 destructive. Blast radius: 1 tool can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
approve_mcp_request | read | |
architecture | read | Structured architecture or design review from multiple engineering perspectives. Each reviewer focuses on different cross-cutting concerns. |
ask_duck | read | |
assumptions | read | Surface and challenge hidden assumptions in a plan, design, or idea. Identifies implicit premises that could be risky if wrong. |
blindspots | read | Hunt for missing considerations, overlooked risks, and gaps in a proposal. Acts as a panel of critical reviewers looking for what might be underweighted. |
challenge | read | The problem or challenge to solve |
chat_with_duck | read | |
clear_conversations | destructive | |
concerns | read | Areas where you feel uncertain or worried |
constraints | read | Known hard constraints that are definitely true |
context | read | Additional background or constraints |
convergence_criteria | read | What makes a solution |
covered | read | What you think you have already addressed or considered |
criteria | read | Evaluation criteria (comma-separated, e.g., |
design | read | Description of the architecture, system design, or technical approach |
dimensions | read | Risk dimensions to focus on (e.g., |
diverge_converge | read | Structure divergent thinking (explore many options) followed by convergence (evaluate and select). Maximizes creative exploration before narrowing down. |
duck_council | read | |
duck_judge | read | |
fetch_url | read | Fetch URL content |
get_pending_approvals | read | |
list_ducks | read | |
list_models | read | |
mcp__fs__read | read | [fs] Read file |
mcp__fs__search | read | [fs] Search files |
mcp_status | read | |
options | read | The options to compare (comma-separated list or detailed descriptions) |
perspectives | read | Analyze a problem from multiple perspectives. Each LLM adopts a different analytical lens (e.g., security, performance, UX) for comprehensive multi-angle analysis. |
plan | read | The plan, design, or idea to analyze for hidden assumptions |
priorities | read | Non-functional priorities (e.g., |
problem | read | The problem, design, or code to analyze |
proposal | read | The plan, code, design, or proposal to review for blindspots |
read_file | read | Read a file |
red_team | read | Conduct attack surface analysis from multiple angles. Each reviewer focuses on different risk dimensions (security, privacy, abuse, compliance). |
reframe | read | Reframe a problem from multiple angles and abstraction levels. Helps break out of mental ruts by viewing the problem differently. |
risk_tolerance | read | Your risk tolerance level: low, medium, or high |
stuck_on | read | What specifically you are stuck on or frustrated by |
target | read | The system, feature, code, or plan to red-team |
test_tool | read | Test tool |
threat_model | read | Known threat actors, attack scenarios, or security context |
tradeoffs | read | Compare options with explicit criteria and trade-off analysis. Provides structured evaluation to help make informed decisions. |
uncertainties | read | Areas where you feel most unsure or want extra scrutiny |
weights | write | Which criteria matter most (priority order or weights) |
width | read | Exploration width: |
workloads | read | Key use cases, workloads, or scenarios the design must handle |
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (8 observation(s))
- Network
- declared (3 observation(s))
- Shell
- none-observed
- Dependencies
- not all pinned
- Secrets in source
- none-found
Findings (15)
clear_conversations
.env.desktop.example
.env.pi.example
.env.template
.releaserc.json
.trivyignore
import { PatternBlockerConfig, ContentPart } from '../../config/types.js';import { GuardrailPhase, GuardrailContext, GuardrailResult } from '../../types.js';import { PIIRedactorConfig, ContentPart } from '../../../config/types.js';import { logger } from '../../../utils/logger.js';import { RateLimiterConfig } from '../../config/types.js';@modelcontextprotocol/ext-apps, @modelcontextprotocol/sdk, ajv, dotenv, openai, winston, zod, @eslint/js
Then add to your `~/.zshrc` or `~/.bashrc`:
curl -fsSL https://ollama.ai/install.sh | sh
4. Post-response: If `restore_on_response=true`, `[EMAIL_1]` -> `[email protected]`
Gates applied: no_behavioural_pass.
95dec8c5fee4full audit observations/trust-audit/mcp-server/nesquikm__rubber-duck.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-06 | 95dec8c5fee4 | SAFE | B | 89 | first audit |
Questions
What is the Rubber Duck MCP server?
An MCP server that acts as a bridge to query multiple OpenAI-compatible LLMs with MCP tool access. Just like rubber duck debugging, explain your problems to various AI "ducks" who can actually research and get different perspectives!
What tools does Rubber Duck expose?
45 in total: 43 read-only, 1 that write, and 1 that can delete or overwrite (clear_conversations). Every one is listed on this page with its risk.
Is Rubber Duck safe to connect to an agent?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B. Separately from the audit: 1 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does Rubber Duck need?
It reads CUSTOM_API1_API_KEY, CUSTOM_API2_API_KEY, CUSTOM_EMPTYMODELS_API_KEY, CUSTOM_INCOMPLETE1_API_KEY, CUSTOM_INTEGRATION_API_KEY, CUSTOM_LMSTUDIO_API_KEY, CUSTOM_MINIMAL_API_KEY, CUSTOM_MYAPI_API_KEY, CUSTOM_MYUPPERAPI_API_KEY, CUSTOM_NICKNAMED_API_KEY, CUSTOM_OLLAMA_API_KEY and CUSTOM_TESTMODELS_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Rubber Duck run?
It speaks stdio and streamable-http, so it runs as a local process your client starts. It is published on npm as mcp-rubber-duck at 1.20.9.
How current is this page?
The grade is for one exact copy of the source (95dec8c5fee4), read on 2026-10-06. The repository is watched and re-audited when it changes.