KnowledgeSAFE
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Overview
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
3e53c07ac202OBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: datachain-knowledge
description: Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.
triggers:
# Discovery
- "what datasets exist"
- "show me the schema"
- "list datasets"
- "datachain knowledge"
- "update the knowledge base"
- "refresh dataset docs"
- "what's in this bucket"
- "explore bucket"
- "scan bucket"
- "bucket overview"
- "what files are in s3://"
- "what files are in gs://"
# Creation & mutation
- "create dataset"
- "save dataset"
- "delete dataset"
- "new dataset"
- "build dataset"
- "make dataset"
- "generate dataset"
# Pipeline output
- "save the results"
- "save to dataset"
# Storage references
- "s3://"
- "gs://"
- "az://"
- "read_storage"
- "from bucket"
- "from s3"
- "from gcs"
# Data processing
- "process files"
- "extract metadata"
- "filter dataset"
- "query dataset"
# Script execution (may create datasets as side effects)
- "run script"
- "run pipeline"
- "python scan"
- "run scan"
---
Maintain a knowledge base at `dc-knowledge/`. `.md` files are the persistent
output. `.json` files are intermediate (generated in Step 3, consumed in
Step 4, then deleted).
`{core_skill_dir}/SDK.md` owns how the pipeline code is written — dataset
shape, row grain, provenance, naming. This file owns the knowledge base
and how the work is run: what already exists, what a run will cost, and
where the results land.
## Critical Rules
1. **Path is `dc-knowledge/`** — NOT `.datachain/`. The `.datachain/` directory is the internal database; the knowledge base lives at `dc-knowledge/`.
2. **Never pass `update=True`** to `dc.read_storage()` in query or exploration code unless the user explicitly asks to refresh the listing. Build scripts that read storage are the exception — they pass `update=True, delta=True`.
3. **Prefer DataChain operations** over plain Python for all metadata analysis.
4. **Bounded output** — JSON and markdown files stay small regardless of data size.
5. **Stop on auth/connection errors** — `bucket_scan.py` runs a fast access check. If it exits with an error JSON on stderr, **stop immediately** and show the error to the user. Do not retry with different regions, profiles, or endpoints — ask for the missing credentials.
6. **Follow the enrichment prompt template literally** in Step 4. Downstream tooling (`render_index.py`) parses the exact frontmatter the prompt prescribes.
## Common gotchas in UDF scripts
- **`parallel=N` vs `workers=N`.** `parallel=N` is local multiprocessing (works anywhere). `workers=N` is Studio-only and MUST be guarded: `chain = chain.settings(parallel=N); if dc.is_studio(): chain = chain.settings(workers=N)`.
- **No `from __future__ import annotations` in UDF modules.** It stringifies type hints and DataChain's signal-schema resolution rejects the string-vs-class mismatch.
- **Type the UDF return precisely.** `Iterator[object]` / `Iterator[Any]` / bare `dict` fail schema resolution. Return a specific `Iterator[T]`, a Pydantic `BaseModel`, or a primitive.
- **Generators aren't subscriptable.** Iterators returned by file APIs do not support `[:N]`. Use `enumerate` + `break`, or `list(...)` only when the result is genuinely small.
- **Use `datachain.__version__` to get the package version** (e.g. `dc.__version__`).
---
## Running the work
### Reuse before building
Read `dc-knowledge/index.md` first. When an existing dataset covers the task —
even partially — read it with `dc.read_dataset(...)` and filter / merge / extend
from there instead of going back to raw storage. Re-running a pass that already
ran is the most expensive mistake available here. Say which dataset was reused
and what it saved.
### Save what was expensive
A UDF that ran a model, decoded file bodies, or called a paid API produces rows
worth keeping: save that operation's **full** output, unfiltered, under a
descriptive name with a `description=`. Chains that only list, filter, or select
are cheap to recompute and need no dataset. Cost is the only criterion — there is
no hierarchy of datasets that has to be built.
### Estimate before a long run
Quote a number before starting anything that may run for minutes:
```
wall ≈ files × per-row × 1.5 / parallel
```
| Op class | Per-row |
|---|---|
| header / metadata parse (bounded-prefix reads) | ~1 ms |
| file-body decode | size / 10-50 MB/s |
| small CPU model (text, light CV) | 5-50 ms |
| mid CPU model (detection, segmentation) | 50-500 ms |
| streaming CPU model (ASR, audio) | 0.1-0.5× realtime |
| local GPU | 10-100× faster than the CPU row |
| paid API (LLM / VLM) | $0.001-0.01 per row + 0.5-2 s, rate-limited |
Measure instead of estimating when the implementation is untested, the model or
library has no row in the table, or files are large enough that decode dominates:
run 3-5 items with `.persist()` (never `.save()`), budget 60 s, and extrapolate
`wall_full = (wall_sample / N) × total_files × 1.5`. Kill at 60 s and fall back to
the estimate.
Label which is which — `estimated ~X` or `measured on N=5: ~X`. Never present an
estimate as a measurement, and write `not measured` literally when nothing was.
### Watch the first minutes
For any run estimated over 5 minutes, read the throughput line DataChain prints
(`Processed: N rows [elapsed, rate]`) over the first 60-90 s:
- at or above ~0.66× the expected rate → carry on;
- below ~0.5× → kill it, report the gap and the revised estimate;
- no throughput line within 2 minutes → kill it and investigate (model download,
auth retry, startup cost).
### Results land in datasTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (1 observation(s))
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
3e53c07ac202full audit observations/trust-audit/skill/datachain-ai__knowledge.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 3e53c07ac202 | SAFE | B | 89 | first audit |
Questions
What does the Knowledge skill do?
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Is Knowledge safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Knowledge access on my machine?
The audit observed that it reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: none found.
How current is this page?
The grade is for one exact copy of the source (3e53c07ac202), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.