Pdf Extraction GuideSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Install
Commands as the repository documents them. They are shown, not run.
pip install marker-pdf
Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: pdf-extraction-guide
description: "PDF parsing, text extraction, and document format conversion"
metadata:
openclaw:
emoji: "📄"
category: "tools"
subcategory: "document"
keywords: ["PDF parsing", "PDF extraction", "document chunking", "format conversion", "md2pdf"]
source: "wentor-research-plugins"
---
# PDF Extraction Guide
Extract text, tables, figures, and metadata from academic PDFs using Python libraries, with strategies for handling multi-column layouts, mathematical content, and scanned documents.
## PDF Extraction Tools Comparison
| Tool | Text | Tables | Figures | Layout | OCR | Speed |
|------|------|--------|---------|--------|-----|-------|
| PyMuPDF (fitz) | Excellent | Manual | Yes | Blocks | No (add with OCR engine) | Fast |
| pdfplumber | Good | Excellent | No | Tables focus | No | Medium |
| PyPDF2 / pypdf | Basic | No | No | No | No | Fast |
| Tabula-py | No | Excellent | No | No | No | Medium |
| GROBID | Structured | Yes | References | Academic layout | No | Slow (ML-based) |
| Nougat (Meta) | Excellent | Yes | Yes | Academic layout | Built-in | Slow (GPU) |
| Marker | Excellent | Yes | Yes | Multi-column | Built-in | Medium |
| pdf2image + Tesseract | Via OCR | Via OCR | Via OCR | No | Yes | Slow |
## PyMuPDF (fitz) — Fast Text Extraction
### Basic Text Extraction
```python
import fitz # pip install PyMuPDF
def extract_text(pdf_path):
"""Extract all text from a PDF with page numbers."""
doc = fitz.open(pdf_path)
full_text = []
for page_num, page in enumerate(doc, 1):
text = page.get_text("text")
full_text.append(f"--- Page {page_num} ---\n{text}")
doc.close()
return "\n".join(full_text)
# Usage
text = extract_text("paper.pdf")
print(text[:2000])
```
### Structured Block-Level Extraction
```python
def extract_structured(pdf_path):
"""Extract text with layout information (blocks, lines, spans)."""
doc = fitz.open(pdf_path)
pages = []
for page_num, page in enumerate(doc):
blocks = page.get_text("dict")["blocks"]
page_data = {"page": page_num + 1, "blocks": []}
for block in blocks:
if "lines" not in block:
continue # Skip image blocks
block_text = ""
max_font_size = 0
is_bold = False
for line in block["lines"]:
for span in line["spans"]:
block_text += span["text"]
max_font_size = max(max_font_size, span["size"])
if "Bold" in span.get("font", ""):
is_bold = True
block_text += "\n"
page_data["blocks"].append({
"text": block_text.strip(),
"font_size": max_font_size,
"is_bold": is_bold,
"bbox": block["bbox"] # (x0, y0, x1, y1)
})
pages.append(page_data)
doc.close()
return pages
# Identify section headings
pages = extract_structured("paper.pdf")
for page in pages:
for block in page["blocks"]:
if block["is_bold"] and block["font_size"] > 11:
print(f"[Heading] {block['text'][:80]}")
```
### Extract Images and Figures
```python
def extract_images(pdf_path, output_dir="./images"):
"""Extract all images from a PDF."""
import os
os.makedirs(output_dir, exist_ok=True)
doc = fitz.open(pdf_path)
img_count = 0
for page_num, page in enumerate(doc):
images = page.get_images(full=True)
for img_idx, img in enumerate(images):
xref = img[0]
pix = fitz.Pixmap(doc, xref)
if pix.n - pix.alpha > 3: # CMYK
pix = fitz.Pixmap(fitz.csRGB, pix)
filename = f"{output_dir}/page{page_num+1}_img{img_idx+1}.png"
pix.save(filename)
img_count += 1
doc.close()
print(f"Extracted {img_count} images to {output_dir}")
```
## pdfplumber — Table Extraction
```python
import pdfplumber
def extract_tables(pdf_path):
"""Extract all tables from a PDF."""
tables = []
with pdfplumber.open(pdf_path) as pdf:
for page_num, page in enumerate(pdf.pages):
page_tables = page.extract_tables()
for table_idx, table in enumerate(page_tables):
tables.append({
"page": page_num + 1,
"table_index": table_idx,
"data": table
})
return tables
# Convert extracted table to pandas DataFrame
import pandas as pd
tables = extract_tables("paper.pdf")
for t in tables:
if t["data"]:
df = pd.DataFrame(t["data"][1:], columns=t["data"][0])
print(f"\nTable on page {t['page']}:")
print(df.to_string())
```
## GROBID — Structured Academic Paper Parsing
GROBID uses machine learning to parse academic PDFs into structured TEI XML.
```python
import requests
def parse_with_grobid(pdf_path, grobid_url="http://localhost:8070"):
"""Parse a paper PDF using GROBID."""
with open(pdf_path, "rb") as f:
response = requests.post(
f"{grobid_url}/api/processFulltextDocument",
files={"input": f},
data={"consolidateHeader": 1, "consolidateCitations": 1}
)
if response.status_code == 200:
return response.text # TEI XML
else:
raise Exception(f"GROBID error: {response.status_code}")
# Parse the TEI XML
from lxml import etree
tei_xml = parse_with_grobid("paper.pdf")
root = etree.fromstring(tei_xml.encode())
ns = {"tei": "http://www.tei-c.org/ns/1.0"}
# Extract title
title = root.find(".//tei:titleStmt/tei:title", ns)
print(f"Title: {title.text if title is not None else 'N/A'}")
# Extract abstract
abstract = root.find(".//tei:profileDesc/tei:abstract", ns)
if abstract is not None:
print(f"Abstract: {abstract.text}")
# Extract references
refs = root.findall(".//tei:listBibl/tei:biblStruct", ns)
print(f"References found: {len(Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__pdf-extraction-guide.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Pdf Extraction Guide skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Pdf Extraction Guide safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Pdf Extraction Guide access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Pdf Extraction Guide work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.