Atlas / MCP servers / webscraping-ai / WebScraping.AI

WebScraping.AISAFE

mcp/webscraping-ai/webscraping-ai

A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.

Verdict
SAFE
Grade
B
Trust score
89 /100
Exposed tools
9 9r · 0w · 0d
Transport
stdio · streamable-http
License
—
Stars
44
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

[](https://www.npmjs.com/package/webscraping-ai-mcp) [](https://github.com/webscraping-ai/webscraping-ai-mcp-server/actions/workflows/ci.yml)

Prefer zero setup? Use the hosted remote MCP server: add https://mcp.webscraping.ai/mcp to your MCP client and sign in with your WebScraping.AI account — OAuth handles auth, no API key or local install needed. This repo is the open-source stdio version for self-hosting and customization.

A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities — Chromium JavaScript rendering, rotating datacenter/residential/stealth proxies, and AI-powered question answering and structured field extraction on any page.

Sign up to get an API key — a free trial, no credit card required. See the API documentation for the full parameter reference.

Features

  • Question answering about web page content
  • Structured data extraction from web pages
  • HTML content retrieval with JavaScript rendering
  • Plain text extraction from web pages
  • CSS selector-based content extraction
  • Google search results (SERP) as parsed JSON
  • Structured JSON for pages on supported sites (e.g. YouTube, TikTok, X, LinkedIn, Instagram, Reddit) from their normal URL
  • Multiple proxy types (datacenter, residential, stealth) with country selection
  • JavaScript rendering using headless Chrome/Chromium
  • Concurrent request management with rate limiting
  • Custom JavaScript execution on target pages
  • Device emulation (desktop, mobile, tablet)
  • Account usage monitoring
  • Content sandboxing option - Wraps scraped content with security boundaries to help protect ag
Read from source at commit 6aad507dd68aOBSERVED · 2026-10-08
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code (npm)
claude mcp add webscraping-ai-mcp --env WEBSCRAPING_AI_API_KEY=${WEBSCRAPING_AI_API_KEY} -- npx -y [email protected]
03

Exposed tools (9)

9 read · 0 write · 0 destructive.

ToolRiskDescription
webscraping_ai_accountread
webscraping_ai_dataread
webscraping_ai_fieldsread
webscraping_ai_htmlread
webscraping_ai_questionread
webscraping_ai_selectedread
webscraping_ai_selected_multipleread
webscraping_ai_serpread
webscraping_ai_textread
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
declared (1 observation(s))
Shell
none-observed
Dependencies
not all pinned
Secrets in source
none-found

Findings (5)

LOWInventory / provenance · inv.hidden_file · CWE-1104
.eslintignore
.eslintignore
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWInventory / provenance · inv.no_license · CWE-1104
Why it matters. no LICENSE file and no repo licence
Fix. add a licence
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
src/stdio.test.js:86
WEBSCRAPING_AI_API_URL: `http://127.0.0.1:${stub.address().port}`,
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
src/stdio.test.js:425
WEBSCRAPING_AI_API_URL: `http://127.0.0.1:${stub.address().port}`,
LOWSupply chain · supply.unpinned · CWE-829, CWE-1357
package.json
@modelcontextprotocol/sdk, axios, dotenv, p-queue, zod, @jest/globals, eslint, eslint-config-prettier
Why it matters. 11 dependency range(s) float
Fix. pin exact versions or ship a lockfile

Gates applied: no_behavioural_pass, no_license.

Audited 2026-10-08 · audit v0.4.1 · source sha 6aad507dd68afull audit observations/trust-audit/mcp-server/webscraping-ai__webscraping-ai.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-086aad507dd68aSAFEB89first audit
06

Questions

What is the WebScraping.AI MCP server?

A Model Context Protocol (MCP) server implementation that integrates with WebScraping.AI for web data extraction capabilities.

What tools does WebScraping.AI expose?

9 in total: 9 read-only, 0 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.

Is WebScraping.AI safe to connect to an agent?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.

What credentials does WebScraping.AI need?

It reads WEBSCRAPING_AI_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does WebScraping.AI run?

It speaks stdio and streamable-http, so it runs as a local process your client starts. It is published on npm as webscraping-ai-mcp at 1.2.2.

How current is this page?

The grade is for one exact copy of the source (6aad507dd68a), read on 2026-10-08. The repository is watched and re-audited when it changes.

Advertisement