ThoughtboxBLOCK
Thoughtbox is an intention ledger for agents. Evaluate AI's decisions against its decision-making.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
Multi-agent collaborative reasoning that's auditable. Thoughtbox is an MCP server for structured, multi-agent reasoning, with a companion web app for workspace and inspection flows. Every step is recorded as a structured thought in a persistent reasoning ledger that can be visualized, exported, and analyzed.
Runtime modes: Local development can use filesystem or in-memory storage. Deployed mode uses Supabase-backed storage, and the current production MCP server runs on Cloud Run.
Observatory UI showing a reasoning session with 14 thoughts and a branch exploration (purple nodes 13-14) forking from thought 5.
Code Mode
Thoughtbox exposes exactly two MCP tools using the Code Mode pattern:
- `thoughtbox_search` — Write JavaScript to query the operation/prompt/resource catalog. The LLM has full programmatic filtering power over the catalog.
- `thoughtbox_execute` — Write JavaScript using the
tbSDK to chain operations. Access thoughts, sessions, knowledge, notebooks, hub, observability, and protocol tools through a unified namespace.
Workflow: search to discover available operations, then execute code against them. Use console.log() for debugging — output is captured in response logs.
This replaces per-operation tool registration with a two-tool surface that scales without context window bloat.
Multi-Agent Collaboration
The Hub is the coordination layer. Agents register with role-specific profiles, join shared workspaces, and work through a structured problem-solving workflow — all via thoughtbox_execute.
The workflow: register → create workspace → create problem → claim → work → propose solution → peer review → merge → consensus
Workspace primitives:
- Problem — A unit of work with dependencies, sub-problems, and status tracking (open → in-progress → resolved → closed)
- Proposal — A proposed solution with a source branch reference
437eeaaba9a8OBSERVED · 2026-10-07Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add thoughtbox -- npx -y @kastalien-research/[email protected]
Exposed tools (164)
118 read · 42 write · 4 destructive. Blast radius: 4 tools can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
ARCHITECT | read | Structural design specialist |
DEBUGGER | read | Root cause analysis specialist |
M10-test | read | Testing per-session isolation |
MANAGER | read | Delegation and team coordination specialist |
RESEARCHER | read | Parallel hypothesis investigation specialist |
REVIEWER | read | Code and proposal review specialist |
SECURITY | read | Risk and vulnerability detection specialist |
add_dependency | write | Add a dependency between problems. The problem cannot be claimed until its dependency is resolved. |
affected | read | Transitive dependents of a claim: claims reachable by reverse depends_on edges. Cycle-safe and depth-capped (default 10, max 20); each result carries its edge distance. |
arch-review | read | Architecture review workspace |
assert | write | Create a new typed claim in a hub workspace with status asserted. Returns the claim with its generated id. |
bad | read | ... |
bind_final_validator | read | Pin a notebook code cell as the predicate that gates terminalState= |
blocked_problems | read | List problems that are blocked by unresolved dependencies. |
branch | read | Toolhost for managing reasoning branches. |
branch-management | write | Spawn and merge reasoning branches |
branch-retrieval | read | List and inspect branches |
branch_get | read | Retrieve a specific branch with its metadata and all thoughts. |
branch_list | read | List all branches for a session with thought counts and status. |
branch_merge | write | Merge branch results back into the main track. |
branch_read | read | Read thoughts from this branch, ordered by thought number. |
branch_spawn | write | Create a new reasoning branch from a main-track thought. |
branch_status | read | Get the status of the current branch (thought count, context). |
branch_thought | read | Record a thought on this branch. Thoughts are numbered sequentially within the branch. |
caching-decision | read | Redis selected |
cell-operations | write | Add, update, list, and retrieve cells |
changed_since | read | Digest query: claims whose status changed strictly after a timestamp, oldest transition first. For session-start recall and natural-boundary syncs (spec §11.1). Optionally narrowed to one hub workspace. Read-only. |
checkpoint | write | Submit a diff for Cassandra adversarial audit. The audit result (approved/rejected) and optional feedback are recorded. |
claim_diff | read | Claim-level diff between two branches in a workspace: added, removed, shared, superseded claims and crossing contradicts edges (SPEC-MERGE-EVIDENCE diffBranchClaims; the same computation feeding merge-evidence notebooks). A claim belongs to a branch when an evidenceRef is |
claim_problem | read | Claim a problem to work on. Auto-generates a branch name if not provided. Sets status to in-progress. |
clear_folder | destructive | Whether to clear the .interleaved-thinking folder after completion, keeping only final-answer.md (default: false) |
collab-ws | read | Collaboration workspace |
collection-ds | read | Collection |
complete | read | End the Theseus session with a terminal state. Bridges a knowledge entry if a summary is provided. |
create_problem | write | Define a new problem to be solved within a workspace. |
create_proposal | write | Propose a solution to a problem. References a thought branch containing the work. |
create_sub_problem | write | Create a sub-problem under an existing parent problem. Sub-problems inherit workspace scope. |
create_workspace | write | Create a new collaboration workspace. The creating agent becomes the coordinator. |
decision | read | ... |
decision-1 | read | ... |
decision-2 | read | ... |
delete | destructive | Remove a stored variable. Returns { deleted: boolean } (false if the name was not set). |
deployment-ds | read | Deployment |
endorse_consensus | read | Endorse an existing consensus marker to show agreement. |
entity-management | write | Create, read, and observe entities in the knowledge graph |
evidence-engine | write | Persist, run, inspect, cancel, and retrieve artifacts for Notebook Evidence Engine modes |
execution | write | Run code cells and install dependencies |
extract_claims | read | Extract atomic claims from a text artifact |
get | write | Read back a value stored earlier in this MCP session via tb.vars.set. Throws a clear error naming the variable if it was never set (or the session restarted) — use tb.vars.list() to see what exists. |
get_profile_prompt | read | Get the behavioral prompt for a specific profile role. Includes domain-specific mental models and guidelines. |
get_session | read | Read a reasoning session through the broker proxy |
graph-structure | write | Create relations and traverse the graph |
health | read | Check health of Thoughtbox infrastructure services (Thoughtbox server, Supabase OTEL store). Optionally filter to specific services. |
init | write | Start a Theseus refactoring session with a declared file scope. All subsequent operations are scoped to these files until a visa expands the boundary. |
input | read | Your task/question. May include |
integration-test | read | Integration test workspace |
interleaved-thinking | read | Use this Thoughtbox server as a reasoning workspace to alternate between internal reasoning steps and external tool/action invocation. Enables structured multi-phase execution with tooling inventory, sufficiency assessment, strategy development, and execution. |
invalidate | read | Mark a claim invalidated. Append-history-style: the claim row is preserved (no hard delete). Idempotent on already-invalidated claims; rejected for superseded claims. |
join_workspace | read | Join an existing workspace. Returns current workspace state including problems and proposals. |
knowledge_add_observation | write | Add an observation to an existing entity. Observations are timestamped notes that track how an entity evolves over time. |
knowledge_create_entity | write | Create a new entity in the knowledge graph. Entities represent insights, concepts, workflows, decisions, or agents. |
knowledge_create_relation | write | Create a directed relation between two entities. Relations form the edges of the knowledge graph. |
knowledge_get_entity | read | Retrieve full details of a specific entity by ID, including all properties and metadata. |
knowledge_list_entities | read | List entities with optional filtering by type, visibility, name pattern, or date range. |
knowledge_query_graph | read | Traverse the knowledge graph starting from an entity. Follows relations up to a specified depth, with optional filtering by relation type. |
knowledge_stats | read | Get aggregate statistics for the knowledge graph: entity counts by type, relation counts, and general health metrics. |
link | write | Create a typed edge between two claims in the same workspace. depends_on edges feed the affected traversal. Idempotent per (from, to, kind). |
list | read | List the names (and serialised sizes in bytes) of all variables stored in this MCP session. |
list_consensus | read | List all consensus markers in a workspace with endorsement counts. |
list_mcp_assets | read | Overview of all MCP capabilities, tools, resources, and quickstart guide |
list_problems | read | List all problems in a workspace with their status and assignments. |
list_proposals | read | List all proposals in a workspace with their review status. |
list_workspaces | read | List all available workspaces. Does not require registration. |
mark_consensus | read | Record a consensus decision. Links to a thought reference for traceability. |
merge_proposal | write | Merge an approved proposal. Requires at least one approval review. |
never | read | never |
notebook-management | write | Create, list, load, and export notebooks |
notebook_add_cell | write | Add a cell to a notebook (title, markdown, or executable code) |
notebook_cancel_run | write | Cancel a running notebook evidence run. |
notebook_create | write | Create a new headless notebook for literate programming or Notebook Evidence Engine workflows with JavaScript/TypeScript support. See thoughtbox://notebook/capabilities for evidence modes and templates. |
notebook_export | read | Export a notebook to .src.md format. Always returns the notebook content as a string. Optionally writes to a filesystem path if provided. - STDIO mode: Provide |
notebook_fitness | read | Read the fitness ledger for a runbook template (SPEC-AGX-SUBSTRATE §7). |
notebook_get_artifact | read | Retrieve a Notebook Evidence Engine artifact by artifactId. |
notebook_get_cell | read | Get complete details of a specific cell including content and execution results |
notebook_get_run | write | Retrieve a run record, including its typed status, outputs, and artifact references. |
notebook_install_deps | write | Install pnpm dependencies defined in the notebook |
notebook_instantiate | read | Reconstruct a live notebook from a persisted, versioned runbook template and |
notebook_list | read | List all active notebooks with their metadata |
notebook_list_cells | read | List all cells in a notebook with their metadata |
notebook_list_runs | read | List notebook evidence runs, optionally filtered by notebookId. |
notebook_load | read | Load a notebook from .src.md format. Accepts exactly one source: - |
notebook_persist | read | Persist the current notebook document. Always stores an in-process artifact for replay/export; with a durable backend configured (FileSystem locally, Supabase deployed) the document also persists durably — upsert by notebookId, latest wins — and the response |
notebook_run_cell | write | Execute a code cell and capture output (stdout, stderr, exit code). |
notebook_start_run | write | Execute a notebook |
notebook_update_cell | write | Update the content of an existing cell |
notebook_validate | write | Run a code cell as a deterministic predicate over JSON-serialisable observed data. The cell receives the observed value as a JSON file pointed to by process.env.TB_OBSERVED_PATH and must write its verdict to process.env.TB_VERDICT_PATH as { |
outcome | read | Record whether tests passed or failed after a modification. Tracks the B (brittleness) counter. |
parallel-verification | write | Execute parallel hypothesis exploration using Thoughtbox. Pass your task/question - optional max_branches param can be included or will default to 3. |
persistent | read | ... |
plan | write | Record a primary action step with a pre-committed recovery step. Each move must be paired with a notebook code cell that will deterministically decide that move |
post_message | write | Post a message to a problem |
post_system_message | write | Post a system message to a problem |
probe | read | Issue a single outbound broker proxy call |
query | read | List claims in a workspace filtered by type, status, creating agent, and/or statement substring. |
quick_join | read | Register and join a workspace in a single call. Combines register + join_workspace for efficient onboarding. |
read_channel | read | Read messages from a problem |
ready_problems | read | List problems that are ready to claim (no unresolved dependencies, status is open). |
reflect | read | Form a falsifiable hypothesis when the state step tracker reaches S=2 (both primary and backup moves failed). Must include explicit disproof criteria. Required before the protocol allows further action steps. Failed moves are forbidden; generate new primary + backup. |
register | read | Register as an agent in the hub. Required before any other hub operation. Returns a unique agentId. |
remove_dependency | destructive | Remove a dependency between problems. |
research | read | Research workspace |
review | read | ... |
review-ws | read | Review workspace |
review_proposal | read | Review a proposal with approve/request-changes/reject verdict. |
runbook_add_await_cell | write | Author an await cell (SPEC-AGX-SUBSTRATE B6): a claim subscription plus a |
runbook_advance | read | Pull-based advancement of a durable runbook instance (tb.runbook.advance — |
runbook_status | read | Read-only snapshot of a durable runbook instance: derived status |
session | read | Toolhost for managing Thoughtbox reasoning sessions. List, search, retrieve, resume, export, and analyze sessions. |
session-analysis | read | Analyze session structure and quality metrics |
session-analysis-guide | read | Process guide for qualitative analysis of reasoning sessions. Teaches how to identify key moments and record learnings to the knowledge graph. |
session-management | read | Resume and export sessions |
session-retrieval | read | List, search, and retrieve session details |
session_analyze | read | Analyze the structure and quality metrics of a reasoning session. Returns objective metrics (linearity, revision rate, branch depth, convergence) - qualitative analysis is done client-side using the session-analysis-guide resource. |
session_export | read | Export a session to markdown format, optionally using cipher notation for compression. Useful for injecting session context into hooks or sharing reasoning chains. |
session_get | read | Retrieve full details of a specific reasoning session, including all thoughts, branches, and metadata. |
session_info | read | Retrieve detailed information about a specific reasoning session including thought count, duration, and metadata. |
session_list | read | List previous reasoning sessions for the current user. Returns session metadata including title, tags, thought count, and timestamps. |
session_query_thoughts | read | Structured queries over a session |
session_resume | read | Load a previous session into the active ThoughtHandler, allowing continuation of reasoning from where it left off. After resuming, subsequent thoughtbox calls will append to this session. |
session_resume_latest | read | Resume the most recently updated session in the current workspace without knowing its ID. Convenience wrapper around session_list + session_resume; optionally filter by tags. Returns the same payload as session_resume, or { success: false } when the workspace has no sessions. |
session_search | read | Search for reasoning sessions by title or tags using a keyword query. |
session_timeline | read | Retrieve the chronological timeline of OTEL events (tool calls, API requests, errors) for a Claude Code session. Events are ordered by timestamp. |
sessions | read | List reasoning sessions with optional filtering by status (active, idle, or all) and a result limit. |
sleep_forever | read | Sleeps past every budget to exercise timeout enforcement |
status | write | Fetch a merge commit by id: status, evidence notebook id + hash, verdict, attribution, and decision timestamps. Read-only. |
subscribe | read | Register a subscriber (agent or runbook cell ref) on a claim. Defaults to the acting agent. Idempotent per (claim, subscriber). |
supersede | read | Replace a claim with a new one: asserts the replacement and marks the old claim superseded with a superseded_by pointer. The old claim is preserved. |
support | read | Attach evidence to a claim and mark it supported. Rejected for invalidated or superseded claims. |
task | read | The task to complete using interleaved thinking approach |
test | read | ... |
test-ws | read | ... |
thoughtbox-interleaved-guides | read | IRCoT-style interleaved reasoning guides centered on Thoughtbox as the canonical reasoning workspace. Mode parameter: research, analysis, or development. |
thoughtbox_core_outcomes | read | Causal-lift core outcomes: end-user tasks where only the final result matters |
thoughtbox_execute | write | Run JavaScript using the \ |
thoughtbox_negative_controls | read | Causal-lift negative controls: simple, direct tasks where Thoughtbox should stay out |
thoughtbox_notebook | write | Notebook toolhost for literate programming with JavaScript/TypeScript. Create, manage, and execute interactive notebooks with markdown documentation and executable code cells. |
thoughtbox_search | read | Discover Thoughtbox operations, prompts, resources, and public tool surfaces by querying this catalog with JavaScript. |
thoughtbox_session | read | Toolhost for managing Thoughtbox reasoning sessions. List, search, retrieve, resume, export, and analyze sessions. |
thoughtbox_theseus | read | Theseus Protocol: friction-gated refactoring for autonomous agents. Prevents scope drift via boundary locking, test-write locks, epistemic visas, and adversarial auditing (Cassandra). Operations: - init: Start a refactoring session with declared file scope (args: { scope: [ |
thoughtbox_thought | write | Advanced reasoning tracking tool. Submit thoughts, track state changes, audit decisions, and build branches or revisions. |
thoughts_limit | read | Maximum number of thoughts to use (default: 100) |
unsubscribe | destructive | Remove a subscriber from a claim. No-op when the subscription does not exist. |
update_problem | write | Update problem status or resolution. Status transitions: open → in-progress → resolved → closed. |
v1-decision | read | Agreed on Redis |
visa | read | Request an epistemic visa to touch an out-of-scope file. Requires justification and acknowledgment of the scope-creep anti-pattern. |
whoami | read | Get current agent identity, role, and workspace memberships. |
workspace_digest | read | Get a comprehensive digest of workspace state: agents, problems, proposals, and consensus. |
workspace_status | read | Get current status of a workspace including agent activity and problem summary. |
ws | read | test |
ws-1 | read | ... |
ws-2 | read | ... |
ws-a | read | A |
ws-b | read | B |
x | read | x |
Trust audit
BLOCKgrade F · trust 30/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | FAIL |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (8 observation(s))
- Network
- declared (10 observation(s))
- Shell
- declared (5 observation(s))
- Dependencies
- not all pinned
- Secrets in source
- found
Findings (25)
await exec(`git merge-base --is-ancestor ${sha} HEAD`, { cwd: ROOT });- **Local DB URL**: `postgresql://postgres:[email protected]:54322/postgres`
- **Local DB URL**: `postgresql://postgres:[email protected]:54322/postgres`
> Operating in the current Afghan media environment presents numerous challenges, including[...] Despite these challenges, Amu TV has managed to continue to provide a vital service to the Afghan popul
timeline.ts
drift-http.ts
claim-diff.ts
apps/web/.augment/skills/nextjs-supabase-auth
apps/web/.augment/skills/vercel-react-best-practices
apps/web/.claude/skills/nextjs-supabase-auth
apps/web/.claude/skills/vercel-react-best-practices
apps/web/.roo/skills/nextjs-supabase-auth
`Δcost=$${costDelta?.toFixed(4) ?? "?"} Δlatency=${(latencyDeltaMs / 1000).toFixed(1)}s ` +`Δtokens=${tokenDelta ?? "?"} vs scratchpad baseline`,- DATABASE_URL=postgres://plausible:password123@plausible-db:5432/plausible
'postgresql://postgres:[email protected]:54322/postgres';
clear_folder, delete, remove_dependency, unsubscribe
.gcloudignore
.gitleaks.toml
.eslintignore
.terraform.lock.hcl
import { resendWelcomeEmailAction } from '../../actions'} from "../../merge-db";
entry = `${entry} ([${shortHash}](../../commit/${commit.hash}))`;import type { ThoughtData } from '../../persistence/types.js';Gates applied: no_behavioural_pass.
437eeaaba9a8full audit observations/trust-audit/mcp-server/kastalien-research__thoughtbox.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | 437eeaaba9a8 | BLOCK | F | 30 | first audit |
Questions
What is the Thoughtbox MCP server?
Thoughtbox is an intention ledger for agents. Evaluate AI's decisions against its decision-making.
What tools does Thoughtbox expose?
164 in total: 118 read-only, 42 that write, and 4 that can delete or overwrite (clear_folder, delete, remove_dependency, unsubscribe). Every one is listed on this page with its risk.
Is Thoughtbox safe to connect to an agent?
No — not without reading the findings first. The audit graded it F (30/100) and found 4 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 4 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does Thoughtbox need?
It reads ANTHROPIC_API_KEY, LANGSMITH_API_KEY, NEXT_PUBLIC_SUPABASE_ANON_KEY, OAUTH_JWT_SECRET, STRIPE_SECRET_KEY, STRIPE_WEBHOOK_SECRET, SUPABASE_ACCESS_TOKEN, SUPABASE_SERVICE_ROLE_KEY, SUPABASE_TEST_ANON_KEY, SUPABASE_TEST_JWT_SECRET, SUPABASE_TEST_SERVICE_ROLE_KEY and TB_BRANCH_SIGNING_SECRET from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Thoughtbox run?
It speaks streamable-http and stdio, so it runs as a local process your client starts. It is published on npm as thoughtbox-claude-code at 0.1.8.
How current is this page?
The grade is for one exact copy of the source (437eeaaba9a8), read on 2026-10-07. The repository is watched and re-audited when it changes.