Atlas / Skills / datachain-ai / Core

CoreSAFE

skills/datachain-ai/core

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
Apache-2.0
Stars
2,823
01

Overview

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Read from source at commit 3e53c07ac202OBSERVED · 2026-10-08
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: datachain-core
description: Use ONLY for abstract DataChain SDK questions — API usage, method signatures, or code patterns — when no specific dataset or bucket is referenced. If the request mentions creating, saving, listing, exploring datasets or buckets, use datachain-knowledge instead.
---

Read `{skill_dir}/SDK.md` in full before answering DataChain SDK questions or generating DataChain Python code. It holds the SDK rules: API usage, UDF signatures, settings, delta semantics, materialization patterns, saving, exporting. The last section below holds the steps that need a local checkout and the `dc-knowledge/` knowledge base.

## Scope of this skill

`SDK.md` owns how DataChain code is written — API usage, UDF signatures, the shape of a saved dataset, saving and exporting. It is self-sufficient on its own.

The **datachain-knowledge** skill owns the knowledge base at `dc-knowledge/`: what datasets already exist, what each one holds, how long a run will take, and keeping that record current. When it is loaded, it drives the session and calls the rules here to write the code.

## Before writing any pipeline code

1. If `dc-knowledge/index.md` exists, read it **first**.
2. When the user's task overlaps with an existing dataset, read its `.md` under `dc-knowledge/datasets/` for schema, code patterns, and lineage.
3. **Bucket access: anonymous or authenticated?** Check `dc-knowledge/buckets/` for a `.md` file with `anon: true/false` in frontmatter. If none, run `datachain bucket status <uri>` to detect. If `denied` or `not found`, stop and ask the user.

Never create or modify files under `dc-knowledge/` — that directory is owned by the `knowledge` skill.
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (1)

INFOPrompt injection · prompt.credential_read · CWE-94, CWE-1427
SDK.md:154
**Provenance: every emitted row carries its source file, typed.** A UDF that fans one file out into many rows puts the originating file on each emitted row as `dc.File` or a subclass (`dc.VideoFile`,
Why it matters. asks the agent to read credentials

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 3e53c07ac202full audit observations/trust-audit/skill/datachain-ai__core.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-083e53c07ac202SAFEB89first audit
05

Questions

What does the Core skill do?

The Context Layer for unstructured data: typed, versioned datasets over S3, GCS, Azure

Is Core safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Core access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (3e53c07ac202), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement