Atlas / MCP servers / mldsveda / PyScrappy

PyScrappySAFE

mcp/mldsveda/pyscrappy

Adaptive Python web scraping toolkit + MCP server for AI agents. Self-healing selectors that survive site changes, TLS-fingerprint stealth to bypass anti-bot filters, CSS/XPath parsing, and 24 built-in scrapers, clean, structured, LLM-ready data from any URL.

Verdict
SAFE
Grade
B
Trust score
89 /100
Exposed tools
24 23r · 1w · 0d
Transport
sse · stdio · streamable-http
License
MIT
Stars
260
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

Adaptive Python web scraping toolkit (self-healing, stealth)+ MCP server for AI agents

[](https://www.python.org/downloads/) [](https://pypi.org/project/PyScrappy/) [](https://github.com/mldsveda/PyScrappy/blob/main/LICENSE) [](https://pepy.tech/project/pyscrappy) [](https://glama.ai/mcp/servers/mldsveda/PyScrappy) [](https://pyscrappy.vercel.app) [](https://mcptoplist.com/server/io.github.mldsveda%2Fpyscrappy)

PyScrappy is an AI-native web scraping toolkit that turns websites into structured, LLM-ready data. Use it as a Python library or expose it as an MCP server for AI agents.

📖 Documentation: pyscrappy.vercel.app

Key features

  • Generic scraper — give it any URL, get back structured text, links, images, tables, and metadata
  • LLM-ready output — .to_markdown() turns any result into clean Markdown; also .to_json() and .to_dataframe()
  • MCP server — expose the scrapers as tools for AI agents (Claude, Cursor, local LLMs, ...)
  • JS rendering — optional Playwright backend for JavaScript-heavy sites
  • Custom selectors — pass CSS selectors to extract exactly what you need
  • Chainable `Selector` — navigate HTML directly with CSS/XPath, find_all, find_by_text, and find_similar (Scrapy/BeautifulSoup-styl
Read from source at commit d08a9dde9549OBSERVED · 2026-10-07
02

Connect

Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control. Replace the environment placeholders with a token scoped to the least it needs.

claude-code (pypi)
claude mcp add pyscrappy --env OMDB_API_KEY=${OMDB_API_KEY} -- uvx pyscrappy==1.6.4
03

Exposed tools (24)

23 read · 1 write · 0 destructive.

ToolRiskDescription
convert_currencyreadFetch live exchange rates and convert an amount from one currency to others.
define_wordreadLook up an English word and return its dictionary entry: definitions, part(s) of speech, and example sentences.
get_cryptoreadFetch live cryptocurrency market data and return a list of coin records, each with fields: id, symbol, name, current price (in vs_currency), market cap, and 24h price change (percent).
get_ubereats_menureadFetch an Uber Eats restaurant
get_weatherreadFetch the current weather conditions for a named place and return a dict with keys: temperature (number, degrees Celsius), humidity (number, percent), wind_speed (number, wind speed), condition (str, e.g.
list_available_scrapersreadList every scraper registered with this server and return their names for use with scrape_with.
lookup_moviereadLook up movie and TV data from IMDB via the OMDb API and return a JSON-serializable dict; a title search returns {
scrape_newsreadFetch news articles from an RSS/Atom feed, a news site (feed auto-discovered), or a single article, and return a list of article dicts (typically: title, url, published date, author, summary, and full text where available).
scrape_stockreadFetch stock market data from Yahoo Finance and return it as a dict.
scrape_urlreadScrape any HTTP(S) URL and return a ScrapeToolResult whose `data` holds one object per page containing extracted text (with word_count), links, images, tables, and page metadata.
scrape_wikipediareadFetch a Wikipedia article by title or search term and return its text content.
scrape_withwriteRun any registered scraper (built-in or plugin) by name and return that scraper
scrape_zomatoreadSearch Zomato for restaurants in a city and return a list of restaurant records.
search_amazonreadScrape Amazon search results for a query and return a list of matching products, each with its title, price, rating, and image URL.
search_booksreadSearch books by title, author, or free text via the Open Library search API and return a list of matching book records.
search_githubreadSearch GitHub for public repositories and return a list of repository records.
search_hackernewsreadSearch Hacker News stories and return a list of matching story dicts, each with title, url, points, author (username), and num_comments (comment count).
search_ikeareadSearch IKEA
search_imagesreadSearch the web for images and return a list of result objects with image URLs and metadata.
search_linkedin_jobsreadSearch LinkedIn public job postings and return a list of matched jobs.
search_neweggreadSearch Newegg for electronics and computer hardware, returning a list of product dicts each with title, price, product_url, image_url, rating, and item_number.
search_soundcloudreadSearch SoundCloud for tracks and return a list of track dicts, each with keys: title (str), artist (str), plays (int), likes (int), and url (str, the track page URL).
search_ubereatsreadSearch Uber Eats for restaurants delivering in a given city, returning a ScrapeToolResult envelope whose `data` is a list of restaurant objects (typically `name`, `eta`, delivery `fee`, and store `url`).
search_youtubereadSearch YouTube for videos matching a query and return a list of matching videos with their metadata.
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
declared (5 observation(s))
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (5)

LOWInventory / provenance · inv.hidden_file · CWE-1104
.pre-commit-config.yaml
.pre-commit-config.yaml
Why it matters. hidden member outside the usual dotfiles
Fix. review its purpose
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
tests/test_core/test_proxy.py:35
cfg = ScraperConfig(proxy="http://127.0.0.1:9")
LOWObfuscation / stealth · obf.decode_call · CWE-506, CWE-94
src/pyscrappy/core/http.py:178
content = base64.b64decode(data["body"])
LOWObfuscation / stealth · obf.decode_call · CWE-506, CWE-94
src/pyscrappy/scrapers/spotify.py:217
decoded = base64.b64decode(match.group(1))
LOWPrompt injection · prompt.transfer_instruction · CWE-94, CWE-1427
CHANGELOG.md:75
- **robots.txt server errors now fail closed.** A `5xx` (or a connection error / timeout) while fetching `robots.txt` is treated as disallow-all and is *not* cached, so a later request re-fetches — in
Why it matters. an instruction to move sensitive data to an outside destination
Fix. remove; a skill never needs the user's secrets off the machine

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha d08a9dde9549full audit observations/trust-audit/mcp-server/mldsveda__pyscrappy.json · Report an issue / request a re-scan
05

Audit history

Every audit this server has had. A grade with a past is a grade somebody is still checking.

DateSourceVerdictGradeScoreChange
2026-10-07d08a9dde9549SAFEB89first audit
06

Questions

What is the PyScrappy MCP server?

Adaptive Python web scraping toolkit + MCP server for AI agents. Self-healing selectors that survive site changes, TLS-fingerprint stealth to bypass anti-bot filters, CSS/XPath parsing, and 24 built-in scrapers, clean, structured, LLM-ready data from any URL.

What tools does PyScrappy expose?

24 in total: 23 read-only, 1 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.

Is PyScrappy safe to connect to an agent?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.

What credentials does PyScrappy need?

It reads OMDB_API_KEY from the environment. Give it a token scoped to the least it needs — an agent that can be talked into calling a tool can be talked into calling it with your credentials.

How does PyScrappy run?

It speaks sse, stdio and streamable-http, so it runs as a local process your client starts. It is published on PyPI as pyscrappy.

How current is this page?

The grade is for one exact copy of the source (d08a9dde9549), read on 2026-10-07. The repository is watched and re-audited when it changes.

Advertisement