AnansiBLOCK
A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when needed, and a coherent identity layer (Chrome TLS fingerprints, matched personas, vendor-aware Cloudflare/Akamai/DataDome handling) works to slip past bot detection. Ships with an MCP server
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
The spider that learns.
[](LICENSE) [](https://python.org) [](docs/mcp.md)
A self-healing web scraper for hostile sites — and it's driveable by any LLM.
Every scraper starts working. The question is how long before it breaks. Anansi is built on a different assumption: the web is adversarial and unstable, and your scraper should handle that without your involvement.
When a site changes its layout, Anansi finds the data anyway and remembers the fix. When a page needs a browser to render, it switches to one silently. When bot detection gets in the way, it presents a coherent identity — matched TLS fingerprint, persona, and headers — that works to slip past detection instead of tripping it. And when you re-crawl, unchanged pages are skipped before a request is even made. The result is a crawler that survives redesigns, handles hostile sites, and gets better the longer it runs.
Ships with an MCP server so any LLM can drive a full crawl through conversation.
Highlights
- Selectors that repair themselves — CSS selectors carry confidence scores; when one breaks, four healing strategies compete and the winner is persisted. How it works →
- A browser only when you need one — every response is checked for JS shells and silently retried in a stealth Playwright browser, cached per domain. How it works →
- A coherent anti-bot identity — matched TLS/HTTP-2 fingerprint, persona, and headers, with vendor-aware handling of Cloudflare, Akamai, and DataDome, plus an opt-in Patchright engine for CDP-level st
fe09918c9226OBSERVED · 2026-10-07Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control.
claude mcp add anansi-scraper -- uvx anansi-scraper
{
"mcpServers": {
"anansi-scraper": {
"command": "uvx",
"args": [
"anansi-scraper"
]
}
}
}Exposed tools (20)
19 read · 0 write · 1 destructive. Blast radius: 1 tool can delete or overwrite — an agent that can be talked into calling a tool can be talked into calling this one.
| Tool | Risk | Description |
|---|---|---|
cancel_crawl | read | |
clear_cache | destructive | |
crawl_metrics | read | |
crawl_site | read | |
export_crawl | read | |
extract | read | |
fetch_and_extract | read | |
fetch_url | read | |
fetch_urls | read | |
get_crawl_items | read | |
internal-nacl-plugin | read | |
internal-pdf-viewer | read | Portable Document Format |
list_crawls | read | |
mhjfbmdgcfjbbpaeojofohoefgiehjai | read | |
pause_crawl | read | |
resume_crawl | read | |
screenshot_url | read | |
selector_health | read | |
train_selector | read | |
validate_selector | read |
Trust audit
BLOCKgrade D · trust 69/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | WARN |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (1 observation(s))
- Network
- declared (6 observation(s))
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (9)
or str(ip) in {"169.254.169.254", "fd00:ec2::254"}clear_cache
sse
page_hash = hashlib.md5(result.html.encode()).hexdigest()
result = await srv.screenshot_url("http://127.0.0.1/")srv._validate_url("http://127.0.0.1/")srv._validate_url("http://127.0.0.1/") # must not raiseheaders={"location": "http://127.0.0.1/"},decoded = base64.b64decode(result["data_b64"])
Gates applied: no_behavioural_pass.
fe09918c9226full audit observations/trust-audit/mcp-server/mdowis__anansi.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | fe09918c9226 | BLOCK | D | 69 | first audit |
Questions
What is the Anansi MCP server?
A self-healing scraper for hostile sites: broken selectors repair themselves, browser rendering kicks in when needed, and a coherent identity layer (Chrome TLS fingerprints, matched personas, vendor-aware Cloudflare/Akamai/DataDome handling) works to slip past bot detection. Ships with an MCP server
What tools does Anansi expose?
20 in total: 19 read-only, 0 that write, and 1 that can delete or overwrite (clear_cache). Every one is listed on this page with its risk.
Is Anansi safe to connect to an agent?
No — not without reading the findings first. The audit graded it D (69/100) and found 1 critical or high issue in the source. Each one is listed on this page with the file and line it is on. Separately from the audit: 1 of its tools can destroy data, so scope the token you give it to what you actually need.
What credentials does Anansi need?
No credential environment variables were found in its source, so it appears to need none.
How does Anansi run?
It speaks sse, so it runs as a service you connect to over the network. It is published on PyPI as anansi-scraper.
How current is this page?
The grade is for one exact copy of the source (fe09918c9226), read on 2026-10-07. The repository is watched and re-audited when it changes.