Pdf ExtractionSAFE
MCP server to extract contents from a PDF file
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
MCP server to extract contents from PDF files, with fixes for Claude Code CLI installation.
This fork includes critical fixes for installing and running the server with Claude Code (the CLI version).
What's Different in This Fork
- Added `__main__.py` - Enables the package to be run as a module with
python -m pdf_extraction - Claude Code specific instructions - Clear installation steps that work with Claude Code CLI
- Tested installation process - Verified working with
claude mcp addcommand
Components
Tools
The server implements one tool:
- extract-pdf-contents: Extract contents from a local PDF file
- Takes
pdf_pathas a required string argument (local file path) - Takes
pagesas an optional string argument (comma-separated page numbers, supports negative indexing like-1for last page) - Supports both PDF text extraction and OCR for scanned documents
Installation for Claude Code CLI
Prerequisites
- Python 3.11 or higher
- pip or conda
- Claude Code CLI installed (
claudecommand)
Step 1: Clone and Install
# Clone this fork git clone https://github.com/lh/mcp-pdf-extraction-server.git cd mcp-pdf-extraction-server # Install in development mode pip install -e .
Step 2: Find the Installed Command
# Check where pdf-extraction was installed which pdf-extraction # Example output: /opt/homebrew/Caskroom/miniconda/base/bin/pdf-extraction
Step 3: Add to Claude Code
# Add the server using the full path from above claude mcp add pdf-extraction /opt/homebrew/Caskroom/miniconda/base/bin/pdf-extraction # Verify it was added claude mcp list
Step 4: Use in Claude
# Start a new Claude session claude # In Claude, type: /mcp # You should see: # MCP Server Status # • pdf-extraction: connected
Usage Example
Once connected, you can ask Claude to extract PDF contents:
"Can you extract the conte
12d07d102b3fOBSERVED · 2026-10-08Connect
Built from this server's own package name, version and transport as found in its source — not copied from anyone's documentation, so it cannot drift against a page we do not control.
claude mcp add pdf-extraction -- uvx pdf-extraction
{
"mcpServers": {
"pdf-extraction": {
"command": "uvx",
"args": [
"pdf-extraction"
]
}
}
}Exposed tools (1)
1 read · 0 write · 0 destructive.
| Tool | Risk | Description |
|---|---|---|
extract-pdf-contents | read | Extract contents from a local PDF file, given page numbers separated in comma. Negative page index number supported. |
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (1)
Gates applied: no_behavioural_pass, no_license.
12d07d102b3ffull audit observations/trust-audit/mcp-server/xraywu__pdf-extraction.json · Report an issue / request a re-scanAudit history
Every audit this server has had. A grade with a past is a grade somebody is still checking.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 12d07d102b3f | SAFE | B | 89 | first audit |
Questions
What is the Pdf Extraction MCP server?
MCP server to extract contents from a PDF file
What tools does Pdf Extraction expose?
1 in total: 1 read-only, 0 that write, and 0 that can delete or overwrite. Every one is listed on this page with its risk.
Is Pdf Extraction safe to connect to an agent?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean server reads B.
What credentials does Pdf Extraction need?
No credential environment variables were found in its source, so it appears to need none.
How does Pdf Extraction run?
It speaks stdio, so it runs as a local process your client starts. It is published on PyPI as pdf-extraction.
How current is this page?
The grade is for one exact copy of the source (12d07d102b3f), read on 2026-10-08. The repository is watched and re-audited when it changes.