Atlas / Skills / opensensenova / Pdf Analysis

Pdf AnalysisSAFE

skills/opensensenova/pdf-analysis

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
MIT
Stars
5,744
01

Overview

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Read from source at commit 657860e4d389OBSERVED · 2026-10-07
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: pdf-analysis
description: "PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。"
---

# PDF Analysis

## Step 0 — Detect PDF type (text vs scanned)

**Critical first step**: determine whether the PDF has extractable text or is a scanned image.
Never skip this — using the wrong parser wastes time and produces empty results.

```python
import fitz  # PyMuPDF

def detect_pdf_type(pdf_path, sample_pages=3):
    """
    Returns 'text' if PDF has extractable text, 'scanned' if image-based.
    Checks first N pages (or all if fewer).
    """
    doc = fitz.open(pdf_path)
    total_chars = 0
    pages_checked = min(sample_pages, len(doc))

    for i in range(pages_checked):
        page = doc[i]
        text = page.get_text("text")
        total_chars += len(text.strip())

    doc.close()
    avg_chars = total_chars / max(pages_checked, 1)
    pdf_type = 'text' if avg_chars > 50 else 'scanned'
    print(f"PDF type: {pdf_type} (avg {avg_chars:.0f} chars/page, checked {pages_checked} pages)")
    return pdf_type
```

---

## Core Method 1: Text PDF — Full Text Extraction (ALL pages)

```python
import fitz

def extract_text_pdf(pdf_path):
    """Extract text from all pages of a text-based PDF."""
    doc = fitz.open(pdf_path)
    total_pages = len(doc)
    print(f"Total pages: {total_pages}")

    all_text = []
    for i, page in enumerate(doc):
        text = page.get_text("text").strip()
        if text:
            all_text.append(f"=== Page {i+1} ===\n{text}")
        else:
            print(f"  Page {i+1}: no text (may be image — will caption later)")

    doc.close()
    return '\n\n'.join(all_text)

# ⚠️ MUST iterate ALL pages — never stop at page 1
full_text = extract_text_pdf(pdf_path)
print(f"Total text length: {len(full_text)} chars")
```

---

## Core Method 2: Text PDF — Table Extraction

For PDFs with tables, `pdfplumber` gives better table structure than `fitz`:

```python
import pdfplumber
import pandas as pd

def extract_tables_pdf(pdf_path):
    """Extract all tables from all pages as DataFrames."""
    all_tables = []
    with pdfplumber.open(pdf_path) as pdf:
        print(f"Total pages: {len(pdf.pages)}")
        for i, page in enumerate(pdf.pages):
            tables = page.extract_tables()
            for j, tbl in enumerate(tables):
                if not tbl:
                    continue
                # First row as header
                df = pd.DataFrame(tbl[1:], columns=tbl[0])
                # Clean: strip whitespace, replace None
                df = df.applymap(lambda x: x.strip() if isinstance(x, str) else x)
                df = df.dropna(how='all').reset_index(drop=True)
                all_tables.append({'page': i+1, 'table_idx': j, 'df': df})
                print(f"  Page {i+1}, Table {j}: {df.shape[0]}r × {df.shape[1]}c")
                print(df.head(3))
    return all_tables

# Verify table alignment after extraction:
# Print column headers and first 3 rows to confirm row/col mapping is correct
```

---

## Core Method 3: Scanned PDF — OCR via Caption

For scanned PDFs (image-based pages), render each page as PNG and caption:

```python
import fitz
import subprocess, json, os

CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"

def extract_scanned_pdf(pdf_path, prompt=None, dpi=150):
    """Render each page as image, then caption for text extraction."""
    doc = fitz.open(pdf_path)
    total_pages = len(doc)
    print(f"Scanned PDF: {total_pages} pages, captioning each...")

    all_text = []
    for i, page in enumerate(doc):
        # Render page to PNG
        mat = fitz.Matrix(dpi/72, dpi/72)
        pix = page.get_pixmap(matrix=mat)
        img_path = f"/tmp/pdf_page_{i+1}.png"
        pix.save(img_path)

        # Caption the page image
        cmd = ["python3", CAPTION, img_path, "--json"]
        if prompt:
            cmd += ["--prompt", prompt]
        else:
            cmd += ["--prompt", "提取页面中所有文字和表格内容,保持原始结构,Markdown格式输出。"]

        r = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
        if r.returncode == 0:
            desc = json.loads(r.stdout).get("description", "")
            all_text.append(f"=== Page {i+1} ===\n{desc}")
            print(f"  Page {i+1}: {len(desc)} chars extracted")
        else:
            print(f"  Page {i+1}: caption failed — {r.stderr[:100]}")

    doc.close()
    return '\n\n'.join(all_text)

# Usage for scanned invoice PDFs, bank statements, org charts, etc.
text = extract_scanned_pdf(pdf_path)
```

---

## Core Method 4: Hybrid PDF (mixed text + image pages)

```python
def extract_hybrid_pdf(pdf_path, text_prompt=None, image_prompt=None):
    """Handle PDFs where some pages have text, others are scanned."""
    doc_fitz = fitz.open(pdf_path)
    all_text = []

    for i, page in enumerate(doc_fitz):
        raw_text = page.get_text("text").strip()

        if len(raw_text) > 50:
            # Text page — use directly
            all_text.append(f"=== Page {i+1} (text) ===\n{raw_text}")
        else:
            # Image page — render and caption
            mat = fitz.Matrix(150/72, 150/72)
            pix = page.get_pixmap(matrix=mat)
            img_path = f"/tmp/hybrid_page_{i+1}.png"
            pix.save(img_path)

            cmd = ["python3", CAPTION, img_path, "--json"]
            prompt = image_prompt or "提取页面中所有文字和表格内容,Markdown格式输出。"
            cmd += ["--prompt", prompt]

            r = subprocess.run(cmd, capture_output=True, text=True, timeout=90)
            if r.returncode == 0:
                desc = json.loads(r.stdout).get("description", "")
                all_text.append(f"=== Page {i+1} (image→caption) ===\n{desc}")
            else:
                all_text.append(f"=== Page {i+1} (caption failed) ===")

    doc_fitz.close()
    return '\n\n'.join(all_text)
```

---

## Core Method 5: Extract Embedded Images / Charts from PDF

```python
import fitz

def extract_pdf_images(pdf_path, min_width=100, min_height=100):
    """Ex
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha 657860e4d389full audit observations/trust-audit/skill/opensensenova__pdf-analysis.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-07657860e4d389SAFEB89first audit
05

Questions

What does the Pdf Analysis skill do?

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Is Pdf Analysis safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Pdf Analysis access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (657860e4d389), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement