CoreSAFE
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Overview
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
3e53c07ac202OBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: datachain-core
description: Use ONLY for abstract DataChain SDK questions — API usage, method signatures, or code patterns — when no specific dataset or bucket is referenced. If the request mentions creating, saving, listing, exploring datasets or buckets, use datachain-knowledge instead.
---
Read `{skill_dir}/SDK.md` in full before answering DataChain SDK questions or generating DataChain Python code. It holds the SDK rules: API usage, UDF signatures, settings, delta semantics, materialization patterns, saving, exporting. The last section below holds the steps that need a local checkout and the `dc-knowledge/` knowledge base.
## Scope of this skill
`SDK.md` owns how DataChain code is written — API usage, UDF signatures, the shape of a saved dataset, saving and exporting. It is self-sufficient on its own.
The **datachain-knowledge** skill owns the knowledge base at `dc-knowledge/`: what datasets already exist, what each one holds, how long a run will take, and keeping that record current. When it is loaded, it drives the session and calls the rules here to write the code.
## Before writing any pipeline code
1. If `dc-knowledge/index.md` exists, read it **first**.
2. When the user's task overlaps with an existing dataset, read its `.md` under `dc-knowledge/datasets/` for schema, code patterns, and lineage.
3. **Bucket access: anonymous or authenticated?** Check `dc-knowledge/buckets/` for a `.md` file with `anon: true/false` in frontmatter. If none, run `datachain bucket status <uri>` to detect. If `denied` or `not found`, stop and ask the user.
Never create or modify files under `dc-knowledge/` — that directory is owned by the `knowledge` skill.Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (1)
**Provenance: every emitted row carries its source file, typed.** A UDF that fans one file out into many rows puts the originating file on each emitted row as `dc.File` or a subclass (`dc.VideoFile`,
Gates applied: no_behavioural_pass.
3e53c07ac202full audit observations/trust-audit/skill/datachain-ai__core.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 3e53c07ac202 | SAFE | B | 89 | first audit |
Questions
What does the Core skill do?
The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure
Is Core safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Core access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (3e53c07ac202), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.