Atlas / MCP servers / nesquikm / Rubber Duck

Rubber DuckSAFE

mcp/nesquikm/rubber-duck

An MCP server that acts as a bridge to query multiple OpenAI-compatible LLMs with MCP tool access. Just like rubber duck debugging, explain your problems to various AI "ducks" who can actually research and get different perspectives!

Verdict
SAFE
Grade
B
Trust score
89 /100
Exposed tools
45 43r · 1w · 1d
Transport
stdio · streamable-http
License
MIT
Stars
178
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

An MCP (Model Context Protocol) server that acts as a bridge to query multiple LLMs -- both OpenAI-compatible HTTP APIs and CLI coding agents. Just like rubber duck debugging, explain your problems to various AI "ducks" and get different perspectives!

[](https://www.npmjs.com/package/mcp-rubber-duck) [](https://github.com/nesquikm/mcp-rubber-duck/pkgs/container/mcp-rubber-duck) [](https://registry.modelcontextprotocol.io)

Why direct provider integration? MCP's sampling primitive -- a server borrowing the host's model -- was deprecated in the 2026-07-28 spec RC in favor of servers integrating directly with LLM provider APIs. Rubber Duck has always worked this way (it brings its own ducks), so it's aligned with where the protocol is heading -- no migration required.

Features

  • Universal OpenAI Compatibility -- Works with any OpenAI-compatible API endpoint
  • CLI Agent Support -- Use CLI coding agents (Claude Code, Codex, Gemini CLI, Grok, Aider) as ducks
  • Multiple Ducks -- Configure and query multiple LLM providers simultaneously
  • Conversation Management -- Maintain context across multiple messages
  • Duck Council -- Get responses from all your configured LLMs at once
  • Consensus Voting -- Multi-duck voting with reasoning and confidence scores
  • LLM-as-Judge -- Have ducks evaluate and rank each other's responses
  • Iterative Refinement -- Two ducks collaboratively improve responses
  • Structured Debates -- Oxford, Socratic, and adversarial debate formats
  • MCP Prompts -- 8 reusable prompt templates for mul
Read from source at commit 95dec8c5fee4OBSERVED · 2026-10-06
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code (npm)
claude mcp add mcp-rubber-duck --env OPENAI_API_KEY=${OPENAI_API_KEY} --env GEMINI_API_KEY=${GEMINI_API_KEY} --env GROQ_API_KEY=${GROQ_API_KEY} -- npx -y [email protected]
03

Exposed tools (45)

43 read · 1 write · 1 destructive. Blast radius: 1 tool can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.

ToolRiskDescription
approve_mcp_requestread
architecturereadStructured architecture or design review from multiple engineering perspectives. Each reviewer focuses on different cross-cutting concerns.
ask_duckread
assumptionsreadSurface and challenge hidden assumptions in a plan, design, or idea. Identifies implicit premises that could be risky if wrong.
blindspotsreadHunt for missing considerations, overlooked risks, and gaps in a proposal. Acts as a panel of critical reviewers looking for what might be underweighted.
challengereadThe problem or challenge to solve
chat_with_duckread
clear_conversationsdestructive
concernsreadAreas where you feel uncertain or worried
constraintsreadKnown hard constraints that are definitely true
contextreadAdditional background or constraints
convergence_criteriareadWhat makes a solution
coveredreadWhat you think you have already addressed or considered
criteriareadEvaluation criteria (comma-separated, e.g.,
designreadDescription of the architecture, system design, or technical approach
dimensionsreadRisk dimensions to focus on (e.g.,
diverge_convergereadStructure divergent thinking (explore many options) followed by convergence (evaluate and select). Maximizes creative exploration before narrowing down.
duck_councilread
duck_judgeread
fetch_urlreadFetch URL content
get_pending_approvalsread
list_ducksread
list_modelsread
mcp__fs__readread[fs] Read file
mcp__fs__searchread[fs] Search files
mcp_statusread
optionsreadThe options to compare (comma-separated list or detailed descriptions)
perspectivesreadAnalyze a problem from multiple perspectives. Each LLM adopts a different analytical lens (e.g., security, performance, UX) for comprehensive multi-angle analysis.
planreadThe plan, design, or idea to analyze for hidden assumptions
prioritiesreadNon-functional priorities (e.g.,
problemreadThe problem, design, or code to analyze
proposalreadThe plan, code, design, or proposal to review for blindspots
read_filereadRead a file
red_teamreadConduct attack surface analysis from multiple angles. Each reviewer focuses on different risk dimensions (security, privacy, abuse, compliance).
reframereadReframe a problem from multiple angles and abstraction levels. Helps break out of mental ruts by viewing the problem differently.
risk_tolerancereadYour risk tolerance level: low, medium, or high
stuck_onreadWhat specifically you are stuck on or frustrated by
targetreadThe system, feature, code, or plan to red-team
test_toolreadTest tool
threat_modelreadKnown threat actors, attack scenarios, or security context
tradeoffsreadCompare options with explicit criteria and trade-off analysis. Provides structured evaluation to help make informed decisions.
uncertaintiesreadAreas where you feel most unsure or want extra scrutiny
weightswriteWhich criteria matter most (priority order or weights)
widthreadExploration width:
workloadsreadKey use cases, workloads, or scenarios the design must handle
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfaceWARN
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
declared (8 observation(s))
Network
declared (3 observation(s))
Shell
none-observed
Dependencies
not all pinned
Secrets in source
none-found

Findings (15)

MEDIUMFilesystem / path · mcp.destructive_tools · CWE-22, CWE-59
clear_conversations
Why it matters. 1 tool(s) can delete or overwrite
Fix. prefer a read-only mode or scoped tokens; the page states the blast radius
LOWInventory / provenance · inv.hidden_file · CWE-1104
.env.desktop.example
.env.desktop.example
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWInventory / provenance · inv.hidden_file · CWE-1104
.env.pi.example
.env.pi.example
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWInventory / provenance · inv.hidden_file · CWE-1104
.env.template
.env.template
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWInventory / provenance · inv.hidden_file · CWE-1104
.releaserc.json
.releaserc.json
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWInventory / provenance · inv.hidden_file · CWE-1104
.trivyignore
.trivyignore
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
src/guardrails/plugins/pattern-blocker.ts:3
import { PatternBlockerConfig, ContentPart } from '../../config/types.js';
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
src/guardrails/plugins/pii-redactor/index.ts:2
import { GuardrailPhase, GuardrailContext, GuardrailResult } from '../../types.js';
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
src/guardrails/plugins/pii-redactor/index.ts:3
import { PIIRedactorConfig, ContentPart } from '../../../config/types.js';
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
src/guardrails/plugins/pii-redactor/index.ts:6
import { logger } from '../../../utils/logger.js';
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
src/guardrails/plugins/rate-limiter.ts:3
import { RateLimiterConfig } from '../../config/types.js';
LOWSupply chain · supply.unpinned · CWE-829, CWE-1357
package.json
@modelcontextprotocol/ext-apps, @modelcontextprotocol/sdk, ajv, dotenv, openai, winston, zod, @eslint/js
Why it matters. 24 dependency range(s) float
Fix. pin exact versions or ship a lockfile
LOWPrompt injection · prompt.persistence · CWE-94, CWE-1427
docs/setup.md:143
Then add to your `~/.zshrc` or `~/.bashrc`:
Why it matters. instructs the agent to persist itself in the user's environment
LOWSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
docs/provider-setup.md:10
curl -fsSL https://ollama.ai/install.sh | sh
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
docs/guardrails.md:119
4. Post-response: If `restore_on_response=true`, `[EMAIL_1]` -> `[email protected]`
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine

Gates applied: no_behavioural_pass.

Audited 2026-10-06 · audit v0.4.1 · source sha 95dec8c5fee4full audit observations/trust-audit/mcp-server/nesquikm__rubber-duck.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-0695dec8c5fee4SAFEB89first audit
06

Questions

What is the Rubber Duck MCP server?

An MCP server that acts as a bridge to query multiple OpenAI-compatible LLMs with MCP tool access. Just like rubber duck debugging, explain your problems to various AI "ducks" who can actually research and get different perspectives!

What tools does Rubber Duck expose?

45 in total: 43 read-only, 1 that write, and 1 that can delete or overwrite (clear_conversations). Every one is listed on this page with its risk.

Is Rubber Duck safe to connect to an agent?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B. Separately from the audit: 1 of its tools can destroy data, so scope the token you give it to what you actually need.

What credentials does Rubber Duck need?

It reads CUSTOM_API1_API_KEY, CUSTOM_API2_API_KEY, CUSTOM_EMPTYMODELS_API_KEY, CUSTOM_INCOMPLETE1_API_KEY, CUSTOM_INTEGRATION_API_KEY, CUSTOM_LMSTUDIO_API_KEY, CUSTOM_MINIMAL_API_KEY, CUSTOM_MYAPI_API_KEY, CUSTOM_MYUPPERAPI_API_KEY, CUSTOM_NICKNAMED_API_KEY, CUSTOM_OLLAMA_API_KEY and CUSTOM_TESTMODELS_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does Rubber Duck run?

It speaks stdio and streamable-http, so it runs as a local process your client starts. It is published on npm as mcp-rubber-duck at 1.20.9.

How current is this page?

The grade is for one exact copy of the source (95dec8c5fee4), read on 2026-10-06. The repository is watched and re-audited when it changes.

Advertisement