Kedro Pipeline GuideSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Install
Commands as the repository documents them. They are shown, not run.
pip install kedro
pip install -e ".[dev]"
pip install kedro-viz
pip install kedro-airflow
Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: kedro-pipeline-guide
description: "Build reproducible data science pipelines with Kedro for research projects"
metadata:
openclaw:
emoji: "🔧"
category: "research"
subcategory: "automation"
keywords: ["pipeline", "reproducibility", "data-science", "workflow", "automation", "mlops"]
source: "https://github.com/kedro-org/kedro"
---
# Kedro Pipeline Guide
## Overview
Kedro is an open-source Python framework for creating reproducible, maintainable, and modular data science pipelines. Developed originally at McKinsey's QuantumBlack labs, Kedro provides an opinionated project structure and a set of conventions that transform ad-hoc analysis scripts into production-quality code that can be tested, versioned, and shared across research teams.
In academic research, reproducibility is both a scientific imperative and a practical challenge. Jupyter notebooks and standalone scripts often become tangled webs of dependencies that are difficult to re-run months later when responding to reviewer comments or extending prior work. Kedro addresses this by separating data processing logic from data access, enforcing explicit pipeline definitions, and providing built-in data versioning and experiment tracking.
With over 11,000 GitHub stars, Kedro has gained adoption across industry and academia. Its design philosophy aligns naturally with the needs of computational research: clear data lineage, parameterized experiments, and the ability to scale from a laptop to a cluster without rewriting code.
## Installation and Setup
Install Kedro via pip:
```bash
pip install kedro
```
Create a new project using the Kedro starter:
```bash
kedro new --name my-research-project --tools lint,test,docs
cd my-research-project
```
This generates a standardized project structure:
```
my-research-project/
conf/
base/
catalog.yml # Data source definitions
parameters.yml # Experiment parameters
local/ # Local overrides (gitignored)
src/
my_research_project/
pipelines/
data_processing/
nodes.py # Pure Python functions
pipeline.py # Pipeline definition
modeling/
nodes.py
pipeline.py
pipeline_registry.py
data/ # Local data directory
notebooks/ # Jupyter notebooks
tests/ # Unit tests
```
Install project dependencies:
```bash
pip install -e ".[dev]"
```
## Core Concepts
**Nodes**: The fundamental units of computation in Kedro. Each node is a pure Python function with explicitly declared inputs and outputs:
```python
# src/my_research_project/pipelines/data_processing/nodes.py
import pandas as pd
from sklearn.preprocessing import StandardScaler
def clean_raw_data(raw_data: pd.DataFrame) -> pd.DataFrame:
"""Remove missing values and outliers from raw experimental data."""
cleaned = raw_data.dropna(subset=["measurement", "condition"])
q1 = cleaned["measurement"].quantile(0.01)
q99 = cleaned["measurement"].quantile(0.99)
return cleaned[cleaned["measurement"].between(q1, q99)]
def normalize_features(
cleaned_data: pd.DataFrame, parameters: dict
) -> pd.DataFrame:
"""Standardize feature columns specified in parameters."""
feature_cols = parameters["feature_columns"]
scaler = StandardScaler()
result = cleaned_data.copy()
result[feature_cols] = scaler.fit_transform(cleaned_data[feature_cols])
return result
```
**Pipelines**: Chains of nodes connected through named datasets:
```python
# src/my_research_project/pipelines/data_processing/pipeline.py
from kedro.pipeline import Pipeline, node, pipeline
from .nodes import clean_raw_data, normalize_features
def create_pipeline(**kwargs) -> Pipeline:
return pipeline([
node(
func=clean_raw_data,
inputs="raw_experiment_data",
outputs="cleaned_data",
name="clean_data_node",
),
node(
func=normalize_features,
inputs=["cleaned_data", "params:preprocessing"],
outputs="normalized_data",
name="normalize_node",
),
])
```
**Data Catalog**: A declarative registry that maps logical dataset names to physical storage:
```yaml
# conf/base/catalog.yml
raw_experiment_data:
type: pandas.CSVDataset
filepath: data/01_raw/experiment_results.csv
cleaned_data:
type: pandas.ParquetDataset
filepath: data/02_intermediate/cleaned.parquet
normalized_data:
type: pandas.ParquetDataset
filepath: data/03_primary/normalized.parquet
versioned: true
```
The `versioned: true` flag automatically creates timestamped versions of outputs, enabling exact reproduction of prior runs.
**Parameters**: Experiment configuration separated from code:
```yaml
# conf/base/parameters.yml
preprocessing:
feature_columns:
- temperature
- pressure
- concentration
outlier_method: iqr
modeling:
algorithm: random_forest
n_estimators: 500
max_depth: 10
test_size: 0.2
random_seed: 42
```
## Running and Visualizing Pipelines
Execute the full pipeline:
```bash
kedro run
```
Run a specific pipeline or node:
```bash
kedro run --pipeline data_processing
kedro run --nodes clean_data_node
```
Visualize the pipeline dependency graph:
```bash
pip install kedro-viz
kedro viz run
```
This launches an interactive web visualization showing the complete data flow, making it easy to understand and communicate your analytical pipeline to collaborators and reviewers.
## Research Workflow Integration
**Experiment Reproducibility**: Every Kedro run uses explicit parameters and versioned data. Store parameter files in Git alongside code to create a complete record of every experiment configuration.
**Reviewer Response**: When peer reviewers request additional analyses or modified parameters, change `parameters.yml` and re-run. The pipeline automatically reprocesses only affected downstream nodes.
**Team CollaborTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__kedro-pipeline-guide.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Kedro Pipeline Guide skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Kedro Pipeline Guide safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Kedro Pipeline Guide access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Kedro Pipeline Guide work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.