Sn Da Excel WorkflowSAFE
Modular SenseNova skills for building AI-powered office assistants and productivity workflows
Overview
Modular SenseNova skills for building AI-powered office assistants and productivity workflows
657860e4d389OBSERVED · 2026-10-07What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: sn-da-excel-workflow
description: "Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子 skill。**遇到以下任一情况就主动使用本 skill,不要自行写几行 pandas 就回答**:1用户出现触发词:Excel 分析 / 表格分析 / 数据分析 / 数据清洗 / 数据统计 / 数据筛选 / 数据可视化 / 数据导出 / 汇总统计 / 透视表 / 分组统计 / 交叉分析 / 趋势分析 / 对比分析 / 异常值检测 / 去重 / 缺失值处理 / Excel 报告 / 生成报表 / analyze Excel / data analysis / data cleaning / pivot table;2用户上传或指定了 .xlsx / .xls / .csv 文件并要求分析、清洗、统计或可视化;3任务涉及多 Sheet 读取、条件筛选、分类汇总、图表生成中的任意一项;4用户要求导出带格式的 Excel 报告或下载链接。仅不用于:不涉及表格数据的纯文本处理、图片分析(使用 sn-da-image-caption)、单个公式计算的简单问答。"
---
# Excel Data Analysis Workflow
End-to-end workflow for structured Excel analysis. Each step maps to a
capability sub-skill that can be loaded for detailed patterns.
## Workflow
### Step 1 — Count rows across all sheets (lightweight, no full load)
Count rows per sheet **without loading data into memory**. Use openpyxl
`read_only` mode — this works for any file size.
```python
import openpyxl, gc
wb = openpyxl.load_workbook(file_path, read_only=True, data_only=True)
total_rows = 0
sheet_info = {}
for name in wb.sheetnames:
ws = wb[name]
row_count = sum(1 for _ in ws.iter_rows(min_row=2, values_only=True))
total_rows += row_count
sheet_info[name] = row_count
print(f"Sheet '{name}': {row_count} rows")
wb.close()
print(f"总行数={total_rows}")
```
⚠️ **Do NOT use `pd.read_excel()` to count rows** — it loads all data into
memory, which will OOM on large files.
→ capability: `excel-reading/multi-sheet-reading`
### Step 2 — Large file gate (CRITICAL — choose strategy by row count)
| total_rows | Strategy | What to do |
|-----------|----------|------------|
| < 10k | Direct read | `df = pd.read_excel(file_path, sheet_name=target_sheet)` |
| 10k – 100k | Parquet cache | `pd.read_excel()` once → `df.to_parquet()` → all later reads from Parquet |
| **>= 100k** | **STOP. Load `sn-da-large-file-analysis` skill** | Read its SKILL.md, then follow its streaming read + Parquet pattern. **Do NOT use `pd.read_excel()` at all** — it will OOM or timeout on 100k+ rows. |
**For >= 100k rows:**
```
read_file(path="<skills_base>/sn-da-large-file-analysis/SKILL.md")
```
Then use `stream_excel_to_parquet()` from that skill — it reads via
openpyxl `iter_rows` in 50k-row chunks with constant memory.
**For 10k – 100k rows (only):**
```python
import pandas as pd
parquet_path = "/tmp/_auto_parquet.parquet"
df = pd.read_excel(file_path, sheet_name=target_sheet)
df.to_parquet(parquet_path, engine="pyarrow")
del df; gc.collect()
df = pd.read_parquet(parquet_path)
```
→ capability: `excel-reading/large-excel-reading`
### Step 3 — Inspect schema & data types
Preview target sheet structure. **For large files (>= 10k rows), only read
a small sample — never full load just to inspect.**
```python
# For any file size — read only first N rows for inspection
df_head = pd.read_excel(file_path, sheet_name=target_sheet, nrows=20)
print(f"Columns: {df_head.columns.tolist()}")
print(f"Dtypes:\n{df_head.dtypes}")
print(df_head.head(10))
```
→ capability: `excel-reading/range-reading`
### Step 4 — Data cleaning
Handle missing values, normalize text, clean invalid characters.
```python
# Missing values
null_count = df[col].isna().sum()
# Text cleaning: keep only Chinese characters
import re
def clean_text(val):
if pd.isna(val): return val
return "".join(re.findall(r"[\u4e00-\u9fff]", str(val))) or ""
df[col] = df[col].apply(clean_text)
```
⚠️ **Large file rule**: When `total_rows >= 100k`, do NOT use `df.apply(lambda...)`.
Use vectorized operations or `np.where()` instead. See `sn-da-large-file-analysis` skill
for the vectorized cheat sheet.
→ capabilities:
- `excel-data-cleaning/missing-value-handling`
- `excel-data-cleaning/invalid-data-cleaning`
- `excel-data-cleaning/text-normalization`
### Step 5 — Filter & extract
Apply condition or category filters, aggregate results.
```python
# Condition filter
mask = df[col].astype(str).str.strip() == target_value
filtered = df[mask]
# Category extraction (for headerless layouts)
df_raw = pd.read_excel(file_path, sheet_name=sheet, header=None)
# Walk rows to find category markers, collect items until next marker
```
→ capabilities:
- `excel-data-filtering/condition-filtering`
- `excel-data-filtering/category-filtering`
- `excel-data-filtering/threshold-filtering`
### Step 6 — Export results
Save filtered/cleaned data as Excel or CSV. Provide download link.
```python
output_path = "/mnt/data/result.xlsx"
result_df.to_excel(output_path, index=False)
print(f"[Download](sandbox:{output_path})")
```
→ capabilities:
- `excel-result-export/single-sheet-export`
- `excel-result-export/formatted-export`
## Key rules
- **Always count rows first** — gate large-file logic on the 10k threshold.
- **>= 100k rows → MUST load `sn-da-large-file-analysis` skill** — do not attempt to handle with `pd.read_excel()`.
- **Column names may contain spaces** (e.g. `'是否通 过'`) — use exact string indexing.
- **Headerless sheets** — use `header=None` and positional indexing.
- **Prohibited on large files (>= 100k rows)**:
- `pd.read_excel()` for full load (use streaming read → Parquet)
- `df.apply(lambda...)` or `df.iterrows()` (use vectorized ops or `itertuples()`)
- `fc-list`, `find ... fonts`, `subprocess` to search fonts, or `pip install` (use fixed font paths below)
- Printing all unique values or full DataFrames (use `.head()`, `.value_counts().head()`)
## CJK Font Setup (mandatory for charts)
When generating charts with matplotlib, **copy this block as-is**. Do NOT search for fonts.
```python
import os
import matplotlib
import matplotlib.pyplot as plt
import matplotlib.font_manager as fm
_FONT_PATHS = [
'/mnt/afs_agents/SimHei.ttf',
'/mnt/afs_agents/mnt/data/SimHei.ttf',
os.path.expanduser('~/.fonts/SimHei.ttf'),
'/usr/share/fonts/truetype/wqy/wqy-zenhei.Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
657860e4d389full audit observations/trust-audit/skill/opensensenova__sn-da-excel-workflow.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | 657860e4d389 | SAFE | B | 89 | first audit |
Questions
What does the Sn Da Excel Workflow skill do?
Modular SenseNova skills for building AI-powered office assistants and productivity workflows
Is Sn Da Excel Workflow safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Sn Da Excel Workflow access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (657860e4d389), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.