Atlas / Skills / opensensenova / Sn Da Excel Workflow

Sn Da Excel WorkflowSAFE

skills/opensensenova/sn-da-excel-workflow

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
MIT
Stars
5,744
01

Overview

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Read from source at commit 657860e4d389OBSERVED · 2026-10-07
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: sn-da-excel-workflow
description: "Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子 skill。**遇到以下任一情况就主动使用本 skill,不要自行写几行 pandas 就回答**:1用户出现触发词:Excel 分析 / 表格分析 / 数据分析 / 数据清洗 / 数据统计 / 数据筛选 / 数据可视化 / 数据导出 / 汇总统计 / 透视表 / 分组统计 / 交叉分析 / 趋势分析 / 对比分析 / 异常值检测 / 去重 / 缺失值处理 / Excel 报告 / 生成报表 / analyze Excel / data analysis / data cleaning / pivot table;2用户上传或指定了 .xlsx / .xls / .csv 文件并要求分析、清洗、统计或可视化;3任务涉及多 Sheet 读取、条件筛选、分类汇总、图表生成中的任意一项;4用户要求导出带格式的 Excel 报告或下载链接。仅不用于:不涉及表格数据的纯文本处理、图片分析(使用 sn-da-image-caption)、单个公式计算的简单问答。"
---

# Excel Data Analysis Workflow

End-to-end workflow for structured Excel analysis. Each step maps to a
capability sub-skill that can be loaded for detailed patterns.

## Workflow

### Step 1 — Count rows across all sheets (lightweight, no full load)

Count rows per sheet **without loading data into memory**. Use openpyxl
`read_only` mode — this works for any file size.

```python
import openpyxl, gc

wb = openpyxl.load_workbook(file_path, read_only=True, data_only=True)
total_rows = 0
sheet_info = {}
for name in wb.sheetnames:
    ws = wb[name]
    row_count = sum(1 for _ in ws.iter_rows(min_row=2, values_only=True))
    total_rows += row_count
    sheet_info[name] = row_count
    print(f"Sheet '{name}': {row_count} rows")
wb.close()
print(f"总行数={total_rows}")
```

⚠️ **Do NOT use `pd.read_excel()` to count rows** — it loads all data into
memory, which will OOM on large files.

→ capability: `excel-reading/multi-sheet-reading`

### Step 2 — Large file gate (CRITICAL — choose strategy by row count)

| total_rows | Strategy | What to do |
|-----------|----------|------------|
| < 10k | Direct read | `df = pd.read_excel(file_path, sheet_name=target_sheet)` |
| 10k – 100k | Parquet cache | `pd.read_excel()` once → `df.to_parquet()` → all later reads from Parquet |
| **>= 100k** | **STOP. Load `sn-da-large-file-analysis` skill** | Read its SKILL.md, then follow its streaming read + Parquet pattern. **Do NOT use `pd.read_excel()` at all** — it will OOM or timeout on 100k+ rows. |

**For >= 100k rows:**
```
read_file(path="<skills_base>/sn-da-large-file-analysis/SKILL.md")
```
Then use `stream_excel_to_parquet()` from that skill — it reads via
openpyxl `iter_rows` in 50k-row chunks with constant memory.

**For 10k – 100k rows (only):**
```python
import pandas as pd
parquet_path = "/tmp/_auto_parquet.parquet"
df = pd.read_excel(file_path, sheet_name=target_sheet)
df.to_parquet(parquet_path, engine="pyarrow")
del df; gc.collect()
df = pd.read_parquet(parquet_path)
```

→ capability: `excel-reading/large-excel-reading`

### Step 3 — Inspect schema & data types

Preview target sheet structure. **For large files (>= 10k rows), only read
a small sample — never full load just to inspect.**

```python
# For any file size — read only first N rows for inspection
df_head = pd.read_excel(file_path, sheet_name=target_sheet, nrows=20)
print(f"Columns: {df_head.columns.tolist()}")
print(f"Dtypes:\n{df_head.dtypes}")
print(df_head.head(10))
```

→ capability: `excel-reading/range-reading`

### Step 4 — Data cleaning

Handle missing values, normalize text, clean invalid characters.

```python
# Missing values
null_count = df[col].isna().sum()

# Text cleaning: keep only Chinese characters
import re
def clean_text(val):
    if pd.isna(val): return val
    return "".join(re.findall(r"[\u4e00-\u9fff]", str(val))) or ""

df[col] = df[col].apply(clean_text)
```

⚠️ **Large file rule**: When `total_rows >= 100k`, do NOT use `df.apply(lambda...)`.
Use vectorized operations or `np.where()` instead. See `sn-da-large-file-analysis` skill
for the vectorized cheat sheet.

→ capabilities:
  - `excel-data-cleaning/missing-value-handling`
  - `excel-data-cleaning/invalid-data-cleaning`
  - `excel-data-cleaning/text-normalization`

### Step 5 — Filter & extract

Apply condition or category filters, aggregate results.

```python
# Condition filter
mask = df[col].astype(str).str.strip() == target_value
filtered = df[mask]

# Category extraction (for headerless layouts)
df_raw = pd.read_excel(file_path, sheet_name=sheet, header=None)
# Walk rows to find category markers, collect items until next marker
```

→ capabilities:
  - `excel-data-filtering/condition-filtering`
  - `excel-data-filtering/category-filtering`
  - `excel-data-filtering/threshold-filtering`

### Step 6 — Export results

Save filtered/cleaned data as Excel or CSV. Provide download link.

```python
output_path = "/mnt/data/result.xlsx"
result_df.to_excel(output_path, index=False)
print(f"[Download](sandbox:{output_path})")
```

→ capabilities:
  - `excel-result-export/single-sheet-export`
  - `excel-result-export/formatted-export`

## Key rules

- **Always count rows first** — gate large-file logic on the 10k threshold.
- **>= 100k rows → MUST load `sn-da-large-file-analysis` skill** — do not attempt to handle with `pd.read_excel()`.
- **Column names may contain spaces** (e.g. `'是否通 过'`) — use exact string indexing.
- **Headerless sheets** — use `header=None` and positional indexing.
- **Prohibited on large files (>= 100k rows)**:
  - `pd.read_excel()` for full load (use streaming read → Parquet)
  - `df.apply(lambda...)` or `df.iterrows()` (use vectorized ops or `itertuples()`)
  - `fc-list`, `find ... fonts`, `subprocess` to search fonts, or `pip install` (use fixed font paths below)
  - Printing all unique values or full DataFrames (use `.head()`, `.value_counts().head()`)

## CJK Font Setup (mandatory for charts)

When generating charts with matplotlib, **copy this block as-is**. Do NOT search for fonts.

```python
import os
import matplotlib
import matplotlib.pyplot as plt
import matplotlib.font_manager as fm

_FONT_PATHS = [
    '/mnt/afs_agents/SimHei.ttf',
    '/mnt/afs_agents/mnt/data/SimHei.ttf',
    os.path.expanduser('~/.fonts/SimHei.ttf'),
    '/usr/share/fonts/truetype/wqy/wqy-zenhei.
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha 657860e4d389full audit observations/trust-audit/skill/opensensenova__sn-da-excel-workflow.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-07657860e4d389SAFEB89first audit
05

Questions

What does the Sn Da Excel Workflow skill do?

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

Is Sn Da Excel Workflow safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Sn Da Excel Workflow access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (657860e4d389), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement