WebclawBLOCK
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
English | 简体中文
webclaw
Turn websites into clean markdown, JSON, and LLM-ready context. CLI, MCP server, REST API, and SDKs for AI agents and RAG pipelines.
ab608ea8739bOBSERVED · 2026-09-23Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.
claude mcp add mcp -- npx -y @webclaw/[email protected]
Trust audit
BLOCKgrade D · trust 69/100 Do not install this without reading the findings. The audit found something that could harm you or your machine.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | FAIL |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (5 observation(s))
- Network
- declared (13 observation(s))
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (16)
/.env' ... curl
/.env" ... curl
.fetch_smart_with_headers("http://127.0.0.1/", &[("Cookie", "fixture=1")])std::env::set_var("WEBCLAW_API_URL", "http://127.0.0.1:8099/v1/");assert_eq!(client.base_url(), "http://127.0.0.1:8099/v1");
assert_eq!(result.as_deref(), Some("http://10.0.0.1:3128"));validate_public_http_url("https://93.184.216.34/")"https://example.com/../../etc/passwd",
assert!(safe_relative_filename("a/../../b.md").is_err());let err = write_to_file(&dir, "../../tmp/webclaw_pwned.md", "x").unwrap_err();
"../../targets_1000.txt",
"../../../targets_1000.txt",
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
- Reddit threads extract reliably again. The old anonymous JSON endpoint is no longer available, so webclaw now reads old.reddit.com directly without an API key or JavaScript. You get the post plus th
assets/sponsors/coldproxy-banner.png
Gates applied: no_behavioural_pass.
ab608ea8739bfull audit observations/trust-audit/mcp-server/0xmassi__webclaw.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-23 | ab608ea8739b | BLOCK | D | 69 | source changed, verdict held |
Questions
What is the Webclaw MCP server?
Fast, local-first web content extraction for LLMs. Scrape, crawl, extract structured data — all from Rust. CLI, REST API, and MCP server.
Is Webclaw safe to connect to an agent?
No — not without reading the findings first. The audit graded it D (69/100) and found 2 critical or high issues in the source. Each one is listed on this page with the file and line it is on.
What credentials does Webclaw need?
It reads FIRECRAWL_API_KEY and GITHUB_TOKEN from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.
How does Webclaw run?
It speaks stdio, so it runs as a local process your client starts. It is published on npm as @webclaw/mcp at 0.6.24.
How current is this page?
The grade is for one exact copy of the source (ab608ea8739b), read on 2026-09-23. The repository is watched and re-audited when it changes.