Atlas / Skills / datachain-ai / Knowledge

KnowledgeSAFE

skills/datachain-ai/knowledge

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
Apache-2.0
Stars
2,823
01

Overview

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Read from source at commit 3e53c07ac202OBSERVED · 2026-10-08
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: datachain-knowledge
description: Use whenever datasets, cloud storage buckets, or data pipelines are mentioned — creating, saving, querying, listing, exploring, deleting, or processing data in S3, GCS, Azure Blob, or local storage. Also use when running any script that may create datasets as a side effect. Maintains a knowledge base at dc-knowledge/ (JSON + markdown). ALWAYS use this skill when the user creates a dataset, saves pipeline output, runs a data script, or references any storage bucket.
triggers:
  # Discovery
  - "what datasets exist"
  - "show me the schema"
  - "list datasets"
  - "datachain knowledge"
  - "update the knowledge base"
  - "refresh dataset docs"
  - "what's in this bucket"
  - "explore bucket"
  - "scan bucket"
  - "bucket overview"
  - "what files are in s3://"
  - "what files are in gs://"
  # Creation & mutation
  - "create dataset"
  - "save dataset"
  - "delete dataset"
  - "new dataset"
  - "build dataset"
  - "make dataset"
  - "generate dataset"
  # Pipeline output
  - "save the results"
  - "save to dataset"
  # Storage references
  - "s3://"
  - "gs://"
  - "az://"
  - "read_storage"
  - "from bucket"
  - "from s3"
  - "from gcs"
  # Data processing
  - "process files"
  - "extract metadata"
  - "filter dataset"
  - "query dataset"
  # Script execution (may create datasets as side effects)
  - "run script"
  - "run pipeline"
  - "python scan"
  - "run scan"
---

Maintain a knowledge base at `dc-knowledge/`. `.md` files are the persistent
output. `.json` files are intermediate (generated in Step 3, consumed in
Step 4, then deleted).

`{core_skill_dir}/SDK.md` owns how the pipeline code is written — dataset
shape, row grain, provenance, naming. This file owns the knowledge base
and how the work is run: what already exists, what a run will cost, and
where the results land.

## Critical Rules

1. **Path is `dc-knowledge/`** — NOT `.datachain/`. The `.datachain/` directory is the internal database; the knowledge base lives at `dc-knowledge/`.
2. **Never pass `update=True`** to `dc.read_storage()` in query or exploration code unless the user explicitly asks to refresh the listing. Build scripts that read storage are the exception — they pass `update=True, delta=True`.
3. **Prefer DataChain operations** over plain Python for all metadata analysis.
4. **Bounded output** — JSON and markdown files stay small regardless of data size.
5. **Stop on auth/connection errors** — `bucket_scan.py` runs a fast access check. If it exits with an error JSON on stderr, **stop immediately** and show the error to the user. Do not retry with different regions, profiles, or endpoints — ask for the missing credentials.
6. **Follow the enrichment prompt template literally** in Step 4. Downstream tooling (`render_index.py`) parses the exact frontmatter the prompt prescribes.

## Common gotchas in UDF scripts

- **`parallel=N` vs `workers=N`.** `parallel=N` is local multiprocessing (works anywhere). `workers=N` is Studio-only and MUST be guarded: `chain = chain.settings(parallel=N); if dc.is_studio(): chain = chain.settings(workers=N)`.
- **No `from __future__ import annotations` in UDF modules.** It stringifies type hints and DataChain's signal-schema resolution rejects the string-vs-class mismatch.
- **Type the UDF return precisely.** `Iterator[object]` / `Iterator[Any]` / bare `dict` fail schema resolution. Return a specific `Iterator[T]`, a Pydantic `BaseModel`, or a primitive.
- **Generators aren't subscriptable.** Iterators returned by file APIs do not support `[:N]`. Use `enumerate` + `break`, or `list(...)` only when the result is genuinely small.
- **Use `datachain.__version__` to get the package version** (e.g. `dc.__version__`).

---

## Running the work

### Reuse before building

Read `dc-knowledge/index.md` first. When an existing dataset covers the task —
even partially — read it with `dc.read_dataset(...)` and filter / merge / extend
from there instead of going back to raw storage. Re-running a pass that already
ran is the most expensive mistake available here. Say which dataset was reused
and what it saved.

### Save what was expensive

A UDF that ran a model, decoded file bodies, or called a paid API produces rows
worth keeping: save that operation's **full** output, unfiltered, under a
descriptive name with a `description=`. Chains that only list, filter, or select
are cheap to recompute and need no dataset. Cost is the only criterion — there is
no hierarchy of datasets that has to be built.

### Estimate before a long run

Quote a number before starting anything that may run for minutes:

```
wall ≈ files × per-row × 1.5 / parallel
```

| Op class | Per-row |
|---|---|
| header / metadata parse (bounded-prefix reads) | ~1 ms |
| file-body decode | size / 10-50 MB/s |
| small CPU model (text, light CV) | 5-50 ms |
| mid CPU model (detection, segmentation) | 50-500 ms |
| streaming CPU model (ASR, audio) | 0.1-0.5× realtime |
| local GPU | 10-100× faster than the CPU row |
| paid API (LLM / VLM) | $0.001-0.01 per row + 0.5-2 s, rate-limited |

Measure instead of estimating when the implementation is untested, the model or
library has no row in the table, or files are large enough that decode dominates:
run 3-5 items with `.persist()` (never `.save()`), budget 60 s, and extrapolate
`wall_full = (wall_sample / N) × total_files × 1.5`. Kill at 60 s and fall back to
the estimate.

Label which is which — `estimated ~X` or `measured on N=5: ~X`. Never present an
estimate as a measurement, and write `not measured` literally when nothing was.

### Watch the first minutes

For any run estimated over 5 minutes, read the throughput line DataChain prints
(`Processed: N rows [elapsed, rate]`) over the first 60-90 s:

- at or above ~0.66× the expected rate → carry on;
- below ~0.5× → kill it, report the gap and the revised estimate;
- no throughput line within 2 minutes → kill it and investigate (model download,
  auth retry, startup cost).

### Results land in datas
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
declared (1 observation(s))
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 3e53c07ac202full audit observations/trust-audit/skill/datachain-ai__knowledge.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-083e53c07ac202SAFEB89first audit
05

Questions

What does the Knowledge skill do?

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Is Knowledge safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Knowledge access on my machine?

The audit observed that it reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: none found.

How current is this page?

The grade is for one exact copy of the source (3e53c07ac202), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement