LLM CouncilCAUTION
Multi-LLM deliberation council — MCP server + Claude Code skill. GPT-5, Gemini 2.5, Claude as peers. Vote, debate, synthesize, critique, MAV protocols.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
Multi LLM deliberation council: query frontier models in parallel, then combine their responses through structured protocols (voting, debate, synthesis, critique, red teaming, verification). Ships as an MCP server and Claude Code skill.
Architecture
flowchart TD
A[User / Claude Code] -->|MCP stdio| B[LLM Council Server]
B --> C{Protocol Router}
C -->|vote| D[Parallel Query + Peer Review]
C -->|debate| E[Multi Round Debate + KS Stopping]
C -->|synthesize| F[Fan Out + Chairman Synthesis]
C -->|critique| G[Peer Critique]
C -->|redteam| H[Adversarial Red Team]
C -->|mav| I[Multi Agent Verification]
D --> J[OpenAI GPT 5.4]
D --> K[Gemini 2.5 Pro]
D --> L[Claude Sonnet 4.6]
E --> J
E --> K
E --> L
F --> J
F --> K
F --> L
F -->|synthesis| M[Chairman Model]
subgraph Providers
J
K
L
end
subgraph Broker [Peer Discovery :7899]
N[Register] --> O[Peers]
P[Messages] --> O
end
B -.->|optional| BrokerFeatures
- Vote. Each model answers independently, then all models anonymously rank each other. Winner selected by first place votes with confidence score.
- Debate. Multi round argumentation with optional KS statistic adaptive stopping. Chairman synthesizes the final round into a consensus answer.
- Synthesize. Fan out to all models, then a chairman model produces an authoritative synthesis combining the best elements.
- Critique. Models answer, then peer review each other's responses with structured feedback.
- Red Team. Adversarial variant of critique that actively probes for flaws, hallucinations, and failure modes.
- MAV (Model as Verifier). Cross checks a candidate answer using multiple models. Each model scores the candidate, highest scored response becomes verified output.
MCP Tools
94862e4939faOBSERVED · 2026-10-09Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add llmcouncil --env ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} --env GEMINI_API_KEY=${GEMINI_API_KEY} --env OPENAI_API_KEY=${OPENAI_API_KEY} -- npx -y [email protected]{
"mcpServers": {
"llmcouncil": {
"command": "npx",
"args": [
"-y",
"[email protected]"
],
"env": {
"ANTHROPIC_API_KEY": "${ANTHROPIC_API_KEY}",
"GEMINI_API_KEY": "${GEMINI_API_KEY}",
"OPENAI_API_KEY": "${OPENAI_API_KEY}"
}
}
}
}Exposed tools (8)
5 read · 3 write · 0 destructive.
| Tool | Risk | Description |
|---|---|---|
council_configure | write | Update the default council configuration for this session. Changes persist until the server restarts. Use this to set preferred models, chairman, or default protocol without passing them on every call. |
council_critique | read | Peer critique: models answer, then critique each other |
council_debate | read | Multi round debate: models argue in rounds, optionally stopping early when responses converge (KS statistic). Chairman synthesizes the final round into a consensus answer. |
council_deliberate | write | Run a multi LLM council deliberation. Sends the same question to multiple frontier models and combines their responses using a chosen protocol (synthesize, vote, debate, critique, redteam, or MAV verification). Returns responses, synthesis/consensus, cost breakdown, and latency. |
council_estimate_cost | write | Estimate the USD cost of a council run before executing it. Returns per model and total cost estimates based on token counts and current pricing. |
council_status | read | Check which LLM providers have API keys configured and list available models with their pricing. Use this to verify the council is ready before running a deliberation. |
council_verify | read | MAV (Model as Verifier): cross checks an answer using multiple models. Each model scores the candidate answer, and the highest scored response becomes the verified output. Use this to fact check or validate an existing answer. |
council_vote | read | Quick voting: each model answers, then all models rank each other |
Trust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | WARN |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- declared (3 observation(s))
- Shell
- none-observed
- Dependencies
- not all pinned
- Secrets in source
- none-found
Findings (4)
const BROKER_URL = "http://127.0.0.1:7899";
console.log(`Listens on ${c.yellow}http://127.0.0.1:7899${c.reset}`);@anthropic-ai/sdk, @google/genai, @modelcontextprotocol/sdk, openai, @types/node, typescript
Gates applied: no_behavioural_pass, no_license.
94862e4939fafull audit observations/trust-audit/mcp-server/rachittshah__llm-council-2.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 94862e4939fa | CAUTION | B | 89 | first audit |
Questions
What is the LLM Council MCP server?
Multi-LLM deliberation council — MCP server + Claude Code skill. GPT-5, Gemini 2.5, Claude as peers. Vote, debate, synthesize, critique, MAV protocols.
What tools does LLM Council expose?
8 in total: 5 read-only, 3 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.
Is LLM Council safe to connect to an agent?
With care. The audit graded it B (89/100) and found 4 things worth knowing before you trust this server, listed below with the exact line each was found on.
What credentials does LLM Council need?
It reads ANTHROPIC_API_KEY, GEMINI_API_KEY and OPENAI_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does LLM Council run?
It speaks stdio, so it runs as a local process your client starts. It is published on npm as llmcouncil at 1.0.0.
How current is this page?
The grade is for one exact copy of the source (94862e4939fa), read on 2026-10-09. The repository is watched and re-audited when it changes.