jDocMunchBLOCK
The leading, most token-efficient MCP server for documentation exploration and retrieval via structured section indexing
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
jDocMunch is an MCP server for coding agents that retrieves the exact documentation section a task needs, without loading whole files into the context window.
Index a documentation set once by heading hierarchy, then fetch a single section, a heading subtree, or a ranked search result — extracted byte-precisely from the original file.
Install · Quickstart · Benchmarks · Commercial licensing
[](https://pypi.org/project/jdocmunch-mcp/) [](https://pypi.org/project/jdocmunch-mcp/) [](https://doi.org/10.5281/zenodo.20102349)
Free for personal use. Commercial use requires a paid license — terms below.
Why jDocMunch?
The problem. An agent asked "how do I configure authentication?" opens a documentation file, skims hundreds of paragraphs it does not need, opens another, and repeats. Large context windows do not fix this. They just make the waste affordable enough to ignore until the bill arrives, and they crowd out the context the model actually needed.
The mechanism. jDocMunch parses a documentation set into a section tree keyed by heading hierarchy, stores each section's byte offsets into the original file, and exposes retrieval over MCP. Sections keep durable identities across re-indexing as long as path, heading text, and heading level are unchanged.
The outcome. The unit of access changes from file to section. An agent retrieves the installation s
d388761fe3faOBSERVED · 2026-10-06Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add jdocmunch-mcp -- None jdocmunch-mcp==1.145.1
Exposed tools (23)
20 read · 2 write · 1 destructive. Blast radius: 1 tool can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
count_sections | write | v1.59+ — count sections matching the same filter set as search_sections (path_glob, role/roles/exclude_roles, tags/exclude_tags, min/max_level, min/max_byte_length) but skip ranking. Use for UI counters or |
delete_index | destructive | Remove a repo index and its cached raw files. Deletes the index and its cached files, never your source documents. There is no undo; re-index to restore. |
describe_section | read | v1.54+ — consolidated handle bundle: full metadata + ancestor breadcrumb + prev/next/parent/first_child neighbors for one section in a single call. Saves three round-trips vs calling get_section_summary + get_section_path + section_neighbors separately. No content reads. |
doc_index_repo | read | Index a GitHub repository |
get_all_roles | read | v1.50+ — list every distinct role classification across the repo with per-role section counts and id samples. Companion to v1.46 get_all_tags. Sections without metadata.role are bucketed under |
get_doc | read | v1.58+ — single-doc detail view. Pairs with list_docs (cross-doc inventory). Returns section list (handles), role_distribution, tag_distribution, byte_size, format, indexed_at for one doc. No content reads. |
get_document_outline | read | Get the section hierarchy for a single document file, without content. Headings only, no content. Read a section with get_section. |
get_index_overview | read | v1.56+ — single-call repo snapshot: doc_count, section_count, total_byte_size, format_breakdown, top_tags, top_roles, indexed_at. Composition of v1.46/v1.50/v1.55 aggregations. Use for |
get_recent_changes | read | v1.47+ — list sections that have drifted from index state (edited_uncommitted or stale_index buckets via the v1.16 FreshnessProbe). By default compares the index against the cached raw-content mirror, NOT live workspace files; pass live_source=true to read the live files under the index |
get_section | read | Retrieve the full content of a specific section using byte-range reads. Use after identifying section IDs via search_sections or get_toc. Returns this section |
get_section_context | read | Retrieve a section with its full hierarchy context: ancestor headings (root → parent) for orientation, the target section |
get_section_descendants | read | v1.43+ — return every descendant of a section (BFS over parent_id) in document order with depth offset. Pairs with get_section_path (ancestors). Optional max_depth caps the walk; max_depth=1 returns immediate children only. Handles only — no content. |
get_section_excerpt | read | v1.41+ — return a short content preview (default 500 bytes) for one section. Trimmed to last newline before the cap so it ends on a paragraph boundary. Use to peek at content before paying for a full get_section read. _meta.tokens_saved reports the byte-savings vs full content. |
get_section_excerpts | read | v1.49+ — batch counterpart to get_section_excerpt. Resolves N previews in one call against a single index load. Per-id errors reported in-line. _meta.tokens_saved aggregates byte savings across the batch. Previews are truncated by design; read full content with get_sections. |
get_section_path | read | v1.40+ — return the breadcrumb chain (root → ... → target) for a section_id. Walks parent_id upward; cycle-protected. Handles only ({id, title, level, doc_path}) per step plus depth. |
get_section_summaries | read | v1.48+ — batch version of get_section_summary. Resolve metadata for many ids in one call against a single index load. Per-id errors are reported in-line on the corresponding result entry rather than aborting the batch. |
get_section_summary | read | v1.38+ — return full indexed metadata (title, summary, role, tags, metadata, parent_id, children, content_hash, byte_start/end, byte_length) for one section without fetching content. Use to inspect role/tags before deciding whether to read the content via get_section. |
get_sections | read | Batch content retrieval for multiple sections in one call. Content only for the ids you pass; unknown ids come back as per-id errors, not a failed call. |
get_toc | read | Get a flat table of contents for all sections in a repo, sorted by document order. Content is excluded — use get_section to retrieve content. Scope with path_glob; for a SINGLE document use get_document_outline (this tool has no doc_path parameter). |
get_toc_tree | read | Get a nested table of contents tree per document. Shows parent/child heading relationships. Content is excluded. |
search_sections | read | Search sections by relevance. Hybrid (BM25 lexical + semantic embedding) fusion when the index was built with use_embeddings=true; falls back to lexical-only otherwise. Returns summaries only — use get_section for full content. |
search_titles | read | v1.57+ — fast title-only token-overlap match. Different from search_sections (full hybrid retrieval). Use for navigation: |
section_neighbors | write | v1.37+ — return prev/next siblings (in document order), parent, and first child for a section. Handles only (id, title, level, doc_path) — no content. Use for fast sequential navigation without re-querying search_sections. |
Trust audit
BLOCKgrade D · trust 69/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (7 observation(s))
- Network
- declared (5 observation(s))
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (25)
"id_rsa",
"id_rsa.*",
"id_ed25519",
"id_ed25519.*",
"id_ecdsa",
delete_index
exec(compile(stripped, vi.__file__ + " (pre-fix)", "exec"), mod.__dict__)
key = hashlib.sha1(resolved_path.encode("utf-8")).hexdigest()return hashlib.sha1(payload).hexdigest()
digest = hashlib.sha1(json.dumps(entries).encode("utf-8")).hexdigest()[:16]shape = hashlib.sha1(
return hashlib.sha1(common_dir_norm.encode("utf-8", errors="replace")).hexdigest()[:16]["local/../../escape", "./local", "local/.", "noslash", ""],
assert store.raw_doc_bytes(owner, name, ["", None, "nope.md", "../../outside.md"]) == 0
@pytest.mark.parametrize("evil", ["../outside.md", "../../outside.md"])raw_files["../../etc/passwd"] = "evil"
"../../etc/passwd",
required: [callback_url, events]
callback_url:
_configure_local(monkeypatch, "http://127.0.0.1:8000/v1", "llama-3.3-70b")
assert str(built._client.base_url).rstrip("/") == "http://127.0.0.1:8000/v1"repo=repo, section_id=sid, verify=False, compress_code=True, storage_path=storage,
flat.frombytes(base64.b64decode(payload))
jDocMunch contributes an anonymous token savings delta to a live global counter with each tool call, POSTed to `https://j.gravelle.us/APIs/savings/post.php` and displayed at [jcodemunch.com](https://j
curl -X POST https://api.example.com/auth/token \
Gates applied: no_behavioural_pass.
d388761fe3fafull audit observations/trust-audit/mcp-server/jgravelle__jdocmunch.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-06 | d388761fe3fa | BLOCK | D | 69 | first audit |
Questions
What is the jDocMunch MCP server?
The leading, most token-efficient MCP server for documentation exploration and retrieval via structured section indexing
What tools does jDocMunch expose?
23 in total: 20 read-only, 2 that write, and 1 that can delete or overwrite (delete_index). Every one is listed on this page with its risk.
Is jDocMunch safe to connect to an agent?
No — not without reading the findings first. The audit graded it D (69/100) and found 5 critical or high issues in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 1 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does jDocMunch need?
It reads ANTHROPIC_API_KEY, GITHUB_TOKEN, GOOGLE_API_KEY, JDOCMUNCH_OPENAI_COMPAT_API_KEY, JDOCMUNCH_SESSION_TOKEN_BUDGET, JDOCMUNCH_SUMMARIZER_API_KEY and OPENAI_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does jDocMunch run?
It speaks stdio and streamable-http, so it runs as a local process your client starts. It is published on PyPI as jdocmunch-mcp.
How current is this page?
The grade is for one exact copy of the source (d388761fe3fa), read on 2026-10-06. The repository is watched and re-audited when it changes.