Atlas / MCP servers / yogthos / Matryoshka

MatryoshkaBLOCK

mcp/yogthos/matryoshka

MCP server for token-efficient large document analysis via the use of REPL state

Verdict
BLOCK
Grade
F
Trust score
59 /100
Exposed tools
78 71r · 4w · 3d
Transport
stdio
License
Apache-2.0
Stars
149
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

[](https://github.com/yogthos/Matryoshka/actions/workflows/test.yml) [](https://safeskill.dev/scan/yogthos-matryoshka)

Process documents 100x larger than your LLM's context window—without vector databases or chunking heuristics.

The Problem

LLMs have fixed context windows. Traditional solutions (RAG, chunking) lose information or miss connections across chunks. RLM takes a different approach: the model reasons about your query and outputs symbolic commands that a logic engine executes against the document.

Based on the Recursive Language Models paper.

How It Works

Unlike traditional approaches where an LLM writes arbitrary code, RLM uses [Nucleus](https://github.com/michaelwhitford/nucleus)—a constrained symbolic language based on S-expressions. The LLM outputs Nucleus commands, which are parsed, type-checked, and executed by Lattice, our logic engine.

┌─────────────────┐     ┌─────────────────┐     ┌─────────────────┐
│   User Query    │────▶│   LLM Reasons   │────▶│ Nucleus Command │
│ "total sales?"  │     │  about intent   │     │  (sum RESULTS)  │
└─────────────────┘     └─────────────────┘     └────────┬────────┘
│
┌─────────────────┐     ┌─────────────────┐     ┌────────▼────────┐
│  Final Answer   │◀────│ Lattice Engine  │◀────│     Parser      │
│   13,000,000    │     │    Executes     │     │    Validates    │
└─────────────────┘     └─────────────────┘     └─────────────────┘

Why this works better than code generation:

  1. Reduced entropy - Nucleus has a rigid grammar with fewer valid outputs than JavaScript
  2. Fail-fast validation - Parser rejects malformed commands before execution
  3. Safe execution - Lattice only executes k
Read from source at commit 3a2549633d52OBSERVED · 2026-10-07
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code
claude mcp add matryoshka-rlm --env API_KEY=${API_KEY} --env MISSING_KEY=${MISSING_KEY} --env TEST_API_KEY=${TEST_API_KEY} --env TEST_EMBEDDED_TOKEN=${TEST_EMBEDDED_TOKEN} -- npx -y [email protected]
claude-desktop
{
  "mcpServers": {
    "matryoshka-rlm": {
      "command": "npx",
      "args": [
        "-y",
        "[email protected]"
      ],
      "env": {
        "API_KEY": "${API_KEY}",
        "MISSING_KEY": "${MISSING_KEY}",
        "TEST_API_KEY": "${TEST_API_KEY}",
        "TEST_EMBEDDED_TOKEN": "${TEST_EMBEDDED_TOKEN}"
      }
    }
  }
}
03

Exposed tools (78)

71 read · 4 write · 3 destructive. Blast radius: 3 tools can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.

ToolRiskDescription
analyze_documentreadAnalyze a document using the Recursive Language Model (RLM).
auto_synthesizedreadSynthesized from ${inputs.length} examples
badread
batch_llm_querywriteExecute multiple LLM queries in parallel. Much faster than sequential llm_query calls when you need to process multiple chunks or sections. All prompts are sent simultaneously.
bracket_contentreadExtract content from brackets
bracket_extractreadExtract content from square brackets
child1read
child2read
context.matchreadSearch context with a regular expression. Returns all matches.
context.slicereadGet a portion of the context by character indices. Use sparingly - prefer fuzzy_search to find relevant sections first.
count_tokensreadEstimate the token count of text. Useful for checking if a section will fit in an LLM call. Uses word-based heuristics (most words = 1 token).
currencyreadCurrency pattern
currencyExtractorreadExtracts currency as number
currency_decimalreadExtract decimal from dollar currency
currency_extractorread
currency_integerreadExtract integer from dollar currency
datereadDate pattern
derivedread
digitsread
emptyread
extractor1read
fuzzy_searchreadFind approximate keyword matches using fuzzy string matching. Returns matching lines with line numbers and match scores (lower score = better match). NOTE: Takes a single search term, NOT regex. For regex patterns use grep() instead.
goodread
grepreadFast regex search with line numbers. More efficient than context.match for finding specific patterns. Returns matches with line numbers and character indices.
highread
high-usagereadhigh usage component
importedread
integer_commasreadParse integer with comma separators
integer_plainreadParse plain integer
key_equals_value_extractreadExtract value from key=value pattern
key_value_extract_keyreadExtract key from key: value pattern
key_value_extract_valuereadExtract value from key: value pattern
lattice_bindingsreadShow current handle bindings. Returns all active handles with their stubs: $grep_error: Array(500) [preview of first item...] $filter: Array(50) [preview...] RESULTS: -> $filter Use this to see what data you have available before deciding what to expand.
lattice_closereadClose the current session and free memory.
lattice_expandreadGet full data from a handle when you need to inspect actual results. USE THIS WHEN: - You need to see actual content to make decisions - You want to verify what
lattice_helpreadGet complete Nucleus command reference documentation.
lattice_llm_respondreadResolve a pending (llm_query ...) suspension with a single string response. When lattice_query hits (llm_query ...) and the client doesn
lattice_loadreadLoad a document for analysis. Call this first before querying.
lattice_memo_deletedestructiveDelete a memo that is no longer needed to free memory. Use this when context becomes stale or irrelevant (e.g., after finishing work on a module). Check lattice_bindings to see current memos and their handles.
lattice_querywriteExecute a Nucleus query on the loaded document. RETURNS HANDLE STUBS (not full data): - Array results return a handle like
lattice_resetdestructiveClear all handles and bindings but keep the document loaded. Use this to start fresh analysis.
lattice_statsreadGet statistics about the currently loaded document.
lattice_statusreadGet current session status including document info, active handles, and timeout.
llm_queryreadQuery a sub-LLM to process a chunk of text. Expensive operation - batch related information when possible to minimize calls. Use format:
locate_linereadExtract lines by line number (1-based). More precise than slice() for line-based navigation. Supports negative indices to count from end.
log_level_extractreadExtract log level from log line
lonelyread
lowread
minikanren_synthesizedreadSynthesized via miniKanren relational search
new-componentread
nucleus_commandsreadGet reference documentation for available Nucleus commands.
nucleus_executewriteExecute Nucleus commands directly on a document without LLM orchestration.
numberreadNumber pattern
originalreadtest description
paren_contentreadExtract content from parentheses
parentread
percentage_to_decimalreadConvert percentage to decimal
prefix_suffix_stripdestructiveRemove prefix ${JSON.stringify(inputPrefix.slice(0, 50))} and suffix ${JSON.stringify(inputSuffix.slice(0, 50))}
regexread
regex1read
regex2read
removableread
split_commareadSplit by comma
split_pipereadSplit by pipe
structured_currency_extractreadExtract currency value from structured text
structured_number_extractreadExtract number from structured text
template_synthesizedreadTemplate synthesized from ${inputs.length} examples
testread
test1read
test2read
test_highread
test_lowread
test_regexreadtest
test_transformreadtest
text_statsreadGet document metadata WITHOUT reading tokens: length, line count, and 5-line samples from start/middle/end. Use this first to understand document structure.
toNumberread
triggerwrite
zero-usagereadnever used component
04

Trust audit

BLOCKgrade F · trust 59/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeFAIL
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfaceWARN
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
declared (3 observation(s))
Network
declared (10 observation(s))
Shell
declared (5 observation(s))
Dependencies
not all pinned
Secrets in source
none-found

Findings (20)

HIGHCode injection · code.eval_exec · CWE-78, CWE-94, CWE-95
src/constraints/verifier.ts:359
const fn = new Function("result", `"use strict"; return (${invariant});`);
Why it matters. evaluates text as code
Fix. remove; use a parser or a dispatch table
HIGHCode injection · code.eval_exec · CWE-78, CWE-94, CWE-95
src/logic/synthesis-integrator.ts:814
const fn = new Function("input", `return ${generatedCode}`) as (
Why it matters. evaluates text as code
Fix. remove; use a parser or a dispatch table
HIGHCode injection · code.eval_exec · CWE-78, CWE-94, CWE-95
src/persistence/predicate-compiler.ts:191
return new Function("item", `"use strict"; return (${code});`);
Why it matters. evaluates text as code
Fix. remove; use a parser or a dispatch table
HIGHCode injection · code.eval_exec · CWE-78, CWE-94, CWE-95
src/persistence/session-db.ts:31
exec(sql: string, params?: SqlValue[]): SqlJsQueryExecResult[];
Why it matters. evaluates text as code
Fix. remove; use a parser or a dispatch table
HIGHCode injection · code.eval_exec · CWE-78, CWE-94, CWE-95
src/persistence/session-db.ts:113
exec(sql: string): void {
Why it matters. evaluates text as code
Fix. remove; use a parser or a dispatch table
MEDIUMNetwork egress · net.raw_ip · CWE-200, CWE-319
src/tool/adapters/http.ts:222
const url = new URL(req.url || "/", `http://127.0.0.1:${localPort}`);
MEDIUMFilesystem / path · mcp.destructive_tools · CWE-22, CWE-59
lattice_memo_delete, lattice_reset, prefix_suffix_strip
Why it matters. 3 tool(s) can delete or overwrite
Fix. prefer a read-only mode or scoped tokens; the page states the blast radius
LOWInformation disclosure · disclose.log_secret · CWE-209, CWE-532
tests/persistence/comparison.test.ts:326
console.log(`Token savings: ${savings}%`);
LOWFilesystem / path · fs.system_paths · CWE-22, CWE-59
tests/tool/lattice-tool.test.ts:292
const result = await tool.executeAsync({ type: "load", filePath: "/etc/passwd" });
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
demos/phase1-rlm-query/harness.ts:13
import { runRLM } from "../../src/rlm.js";
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
demos/phase1-rlm-query/harness.ts:14
import { createNucleusAdapter } from "../../src/adapters/nucleus.js";
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
demos/phase3-multi-context/after.test.ts:22
import { runRLMFromContent } from "../../src/rlm.js";
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
demos/phase3-multi-context/after.test.ts:23
import { createNucleusAdapter } from "../../src/adapters/nucleus.js";
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
demos/phase3-multi-context/before.test.ts:20
import { runRLMFromContent } from "../../src/rlm.js";
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
tests/tool/http-adapter.test.ts:77
const response = await fetch(`http://127.0.0.1:${port}/load`, {
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
tests/tool/http-adapter.test.ts:94
await fetch(`http://127.0.0.1:${port}/load`, {
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
tests/tool/http-adapter.test.ts:100
const response = await fetch(`http://127.0.0.1:${port}/reset`, {
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
tests/tool/http-adapter.test.ts:115
const response = await fetch(`http://127.0.0.1:${port}/health`);
LOWObfuscation / stealth · obf.decode_call · CWE-506, CWE-94
tests/synthesis/coordinator.test.ts:387
expect(() => safeEvalSynthesized(`(x) => atob(x)`)).toThrow();
LOWSupply chain · supply.unpinned · CWE-829, CWE-1357
package.json
@modelcontextprotocol/sdk, @tree-sitter-grammars/tree-sitter-markdown, @types/sql.js, @vscode/tree-sitter-wasm, @yogthos/tree-sitter-clojure, graphology, graphology-communities-louvain, graphology-dag
Why it matters. 30 dependency range(s) float
Fix. pin exact versions or ship a lockfile

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha 3a2549633d52full audit observations/trust-audit/mcp-server/yogthos__matryoshka.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-073a2549633d52BLOCKF59first audit
06

Questions

What is the Matryoshka MCP server?

MCP server for token-efficient large document analysis via the use of REPL state

What tools does Matryoshka expose?

78 in total: 71 read-only, 4 that write, and 3 that can delete or overwrite (lattice_memo_delete, lattice_reset, prefix_suffix_strip). Every one is listed on this page with its risk.

Is Matryoshka safe to connect to an agent?

No — not without reading the findings first. The audit graded it F (59/100) and found 5 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 3 of its tools can destroy data, so scope the token you give it to what you actually need.

What credentials does Matryoshka need?

It reads API_KEY, MISSING_KEY, TEST_API_KEY and TEST_EMBEDDED_TOKEN from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does Matryoshka run?

It speaks stdio, so it runs as a local process your client starts. It is published on npm as matryoshka-rlm at 0.2.39.

How current is this page?

The grade is for one exact copy of the source (3a2549633d52), read on 2026-10-07. The repository is watched and re-audited when it changes.

Advertisement