MatryoshkaBLOCK
MCP server for token-efficient large document analysis via the use of REPL state
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
[](https://github.com/yogthos/Matryoshka/actions/workflows/test.yml) [](https://safeskill.dev/scan/yogthos-matryoshka)
Process documents 100x larger than your LLM's context window—without vector databases or chunking heuristics.
The Problem
LLMs have fixed context windows. Traditional solutions (RAG, chunking) lose information or miss connections across chunks. RLM takes a different approach: the model reasons about your query and outputs symbolic commands that a logic engine executes against the document.
Based on the Recursive Language Models paper.
How It Works
Unlike traditional approaches where an LLM writes arbitrary code, RLM uses [Nucleus](https://github.com/michaelwhitford/nucleus)—a constrained symbolic language based on S-expressions. The LLM outputs Nucleus commands, which are parsed, type-checked, and executed by Lattice, our logic engine.
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ │ User Query │────▶│ LLM Reasons │────▶│ Nucleus Command │ │ "total sales?" │ │ about intent │ │ (sum RESULTS) │ └─────────────────┘ └─────────────────┘ └────────┬────────┘ │ ┌─────────────────┐ ┌─────────────────┐ ┌────────▼────────┐ │ Final Answer │◀────│ Lattice Engine │◀────│ Parser │ │ 13,000,000 │ │ Executes │ │ Validates │ └─────────────────┘ └─────────────────┘ └─────────────────┘
Why this works better than code generation:
- Reduced entropy - Nucleus has a rigid grammar with fewer valid outputs than JavaScript
- Fail-fast validation - Parser rejects malformed commands before execution
- Safe execution - Lattice only executes k
3a2549633d52OBSERVED · 2026-10-07Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add matryoshka-rlm --env API_KEY=${API_KEY} --env MISSING_KEY=${MISSING_KEY} --env TEST_API_KEY=${TEST_API_KEY} --env TEST_EMBEDDED_TOKEN=${TEST_EMBEDDED_TOKEN} -- npx -y [email protected]{
"mcpServers": {
"matryoshka-rlm": {
"command": "npx",
"args": [
"-y",
"[email protected]"
],
"env": {
"API_KEY": "${API_KEY}",
"MISSING_KEY": "${MISSING_KEY}",
"TEST_API_KEY": "${TEST_API_KEY}",
"TEST_EMBEDDED_TOKEN": "${TEST_EMBEDDED_TOKEN}"
}
}
}
}Exposed tools (78)
71 read · 4 write · 3 destructive. Blast radius: 3 tools can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
analyze_document | read | Analyze a document using the Recursive Language Model (RLM). |
auto_synthesized | read | Synthesized from ${inputs.length} examples |
bad | read | |
batch_llm_query | write | Execute multiple LLM queries in parallel. Much faster than sequential llm_query calls when you need to process multiple chunks or sections. All prompts are sent simultaneously. |
bracket_content | read | Extract content from brackets |
bracket_extract | read | Extract content from square brackets |
child1 | read | |
child2 | read | |
context.match | read | Search context with a regular expression. Returns all matches. |
context.slice | read | Get a portion of the context by character indices. Use sparingly - prefer fuzzy_search to find relevant sections first. |
count_tokens | read | Estimate the token count of text. Useful for checking if a section will fit in an LLM call. Uses word-based heuristics (most words = 1 token). |
currency | read | Currency pattern |
currencyExtractor | read | Extracts currency as number |
currency_decimal | read | Extract decimal from dollar currency |
currency_extractor | read | |
currency_integer | read | Extract integer from dollar currency |
date | read | Date pattern |
derived | read | |
digits | read | |
empty | read | |
extractor1 | read | |
fuzzy_search | read | Find approximate keyword matches using fuzzy string matching. Returns matching lines with line numbers and match scores (lower score = better match). NOTE: Takes a single search term, NOT regex. For regex patterns use grep() instead. |
good | read | |
grep | read | Fast regex search with line numbers. More efficient than context.match for finding specific patterns. Returns matches with line numbers and character indices. |
high | read | |
high-usage | read | high usage component |
imported | read | |
integer_commas | read | Parse integer with comma separators |
integer_plain | read | Parse plain integer |
key_equals_value_extract | read | Extract value from key=value pattern |
key_value_extract_key | read | Extract key from key: value pattern |
key_value_extract_value | read | Extract value from key: value pattern |
lattice_bindings | read | Show current handle bindings. Returns all active handles with their stubs: $grep_error: Array(500) [preview of first item...] $filter: Array(50) [preview...] RESULTS: -> $filter Use this to see what data you have available before deciding what to expand. |
lattice_close | read | Close the current session and free memory. |
lattice_expand | read | Get full data from a handle when you need to inspect actual results. USE THIS WHEN: - You need to see actual content to make decisions - You want to verify what |
lattice_help | read | Get complete Nucleus command reference documentation. |
lattice_llm_respond | read | Resolve a pending (llm_query ...) suspension with a single string response. When lattice_query hits (llm_query ...) and the client doesn |
lattice_load | read | Load a document for analysis. Call this first before querying. |
lattice_memo_delete | destructive | Delete a memo that is no longer needed to free memory. Use this when context becomes stale or irrelevant (e.g., after finishing work on a module). Check lattice_bindings to see current memos and their handles. |
lattice_query | write | Execute a Nucleus query on the loaded document. RETURNS HANDLE STUBS (not full data): - Array results return a handle like |
lattice_reset | destructive | Clear all handles and bindings but keep the document loaded. Use this to start fresh analysis. |
lattice_stats | read | Get statistics about the currently loaded document. |
lattice_status | read | Get current session status including document info, active handles, and timeout. |
llm_query | read | Query a sub-LLM to process a chunk of text. Expensive operation - batch related information when possible to minimize calls. Use format: |
locate_line | read | Extract lines by line number (1-based). More precise than slice() for line-based navigation. Supports negative indices to count from end. |
log_level_extract | read | Extract log level from log line |
lonely | read | |
low | read | |
minikanren_synthesized | read | Synthesized via miniKanren relational search |
new-component | read | |
nucleus_commands | read | Get reference documentation for available Nucleus commands. |
nucleus_execute | write | Execute Nucleus commands directly on a document without LLM orchestration. |
number | read | Number pattern |
original | read | test description |
paren_content | read | Extract content from parentheses |
parent | read | |
percentage_to_decimal | read | Convert percentage to decimal |
prefix_suffix_strip | destructive | Remove prefix ${JSON.stringify(inputPrefix.slice(0, 50))} and suffix ${JSON.stringify(inputSuffix.slice(0, 50))} |
regex | read | |
regex1 | read | |
regex2 | read | |
removable | read | |
split_comma | read | Split by comma |
split_pipe | read | Split by pipe |
structured_currency_extract | read | Extract currency value from structured text |
structured_number_extract | read | Extract number from structured text |
template_synthesized | read | Template synthesized from ${inputs.length} examples |
test | read | |
test1 | read | |
test2 | read | |
test_high | read | |
test_low | read | |
test_regex | read | test |
test_transform | read | test |
text_stats | read | Get document metadata WITHOUT reading tokens: length, line count, and 5-line samples from start/middle/end. Use this first to understand document structure. |
toNumber | read | |
trigger | write | |
zero-usage | read | never used component |
Trust audit
BLOCKgrade F · trust 59/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (3 observation(s))
- Network
- declared (10 observation(s))
- Shell
- declared (5 observation(s))
- Dependencies
- not all pinned
- Secrets in source
- none-found
Findings (20)
const fn = new Function("result", `"use strict"; return (${invariant});`);const fn = new Function("input", `return ${generatedCode}`) as (return new Function("item", `"use strict"; return (${code});`);exec(sql: string, params?: SqlValue[]): SqlJsQueryExecResult[];
exec(sql: string): void {const url = new URL(req.url || "/", `http://127.0.0.1:${localPort}`);lattice_memo_delete, lattice_reset, prefix_suffix_strip
console.log(`Token savings: ${savings}%`);const result = await tool.executeAsync({ type: "load", filePath: "/etc/passwd" });import { runRLM } from "../../src/rlm.js";import { createNucleusAdapter } from "../../src/adapters/nucleus.js";import { runRLMFromContent } from "../../src/rlm.js";import { createNucleusAdapter } from "../../src/adapters/nucleus.js";import { runRLMFromContent } from "../../src/rlm.js";const response = await fetch(`http://127.0.0.1:${port}/load`, {await fetch(`http://127.0.0.1:${port}/load`, {const response = await fetch(`http://127.0.0.1:${port}/reset`, {const response = await fetch(`http://127.0.0.1:${port}/health`);expect(() => safeEvalSynthesized(`(x) => atob(x)`)).toThrow();
@modelcontextprotocol/sdk, @tree-sitter-grammars/tree-sitter-markdown, @types/sql.js, @vscode/tree-sitter-wasm, @yogthos/tree-sitter-clojure, graphology, graphology-communities-louvain, graphology-dag
Gates applied: no_behavioural_pass.
3a2549633d52full audit observations/trust-audit/mcp-server/yogthos__matryoshka.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | 3a2549633d52 | BLOCK | F | 59 | first audit |
Questions
What is the Matryoshka MCP server?
MCP server for token-efficient large document analysis via the use of REPL state
What tools does Matryoshka expose?
78 in total: 71 read-only, 4 that write, and 3 that can delete or overwrite (lattice_memo_delete, lattice_reset, prefix_suffix_strip). Every one is listed on this page with its risk.
Is Matryoshka safe to connect to an agent?
No — not without reading the findings first. The audit graded it F (59/100) and found 5 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 3 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does Matryoshka need?
It reads API_KEY, MISSING_KEY, TEST_API_KEY and TEST_EMBEDDED_TOKEN from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Matryoshka run?
It speaks stdio, so it runs as a local process your client starts. It is published on npm as matryoshka-rlm at 0.2.39.
How current is this page?
The grade is for one exact copy of the source (3a2549633d52), read on 2026-10-07. The repository is watched and re-audited when it changes.