Atlas / Skills / brycewang-stanford / Dataset Finder Guide

Dataset Finder GuideSAFE

skills/brycewang-stanford/dataset-finder-guide

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
1 documented
License
NOASSERTION
Stars
4,535
01

Overview

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Read from source at commit e1ba289846fdOBSERVED · 2026-10-08
02

Install

Commands as the repository documents them. They are shown, not run.

pip install kaggle
03

Host compatibility

What the documentation claims. We have not run a compatibility test.

HostStatusNotes
openclawmentioned
04

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: dataset-finder-guide
description: "Search and download research datasets from Kaggle, HuggingFace, and repos"
metadata:
  openclaw:
    emoji: "🗄️"
    category: "tools"
    subcategory: "scraping"
    keywords: ["dataset", "Kaggle", "data download", "HuggingFace", "data repository", "open data"]
    source: "wentor-research-plugins"
---

# Dataset Finder Guide

Search, evaluate, and download research datasets from major repositories including Kaggle, Hugging Face, Google Dataset Search, Zenodo, UCI Machine Learning Repository, and domain-specific archives. This skill helps researchers locate the right data for their experiments efficiently.

## Overview

Finding suitable datasets is often one of the most time-consuming phases of empirical research. Datasets are scattered across dozens of platforms, each with different APIs, licensing terms, download mechanisms, and metadata standards. A single research project might require datasets from Kaggle for benchmarking, Hugging Face for NLP tasks, Zenodo for supplementary materials from published papers, and government open data portals for demographic or economic variables.

This skill provides a unified approach to dataset discovery: formulating search queries, evaluating dataset quality and suitability, understanding licensing implications, and efficiently downloading and organizing data. It covers both general-purpose repositories and domain-specific archives that researchers in various fields need.

The emphasis is on reproducibility -- every dataset used in research should be citable, versioned, and documented. This skill includes patterns for recording dataset provenance, creating data cards, and managing dataset versions across experiments.

## Dataset Repositories

### General-Purpose Repositories

| Repository | Strengths | API | Citation Support |
|------------|-----------|-----|-----------------|
| Kaggle | ML benchmarks, competitions, community kernels | REST + CLI | DOI via dataset cards |
| Hugging Face Datasets | NLP, CV, audio; streaming support | Python library | Built-in citation |
| Zenodo | Any research data, DOI minting, EU-funded | REST API | Automatic DOI |
| Google Dataset Search | Meta-search across repositories | Web only | Links to source |
| UCI ML Repository | Classic ML benchmarks | Direct download | BibTeX provided |
| Figshare | Figures, datasets, media, preprints | REST API | DOI per item |
| Dryad | Ecology, biology, environmental science | REST API | DOI per dataset |
| ICPSR | Social science survey data | Restricted API | Persistent IDs |
| Harvard Dataverse | Multi-discipline, institutional | REST API | DOI per dataset |

### Domain-Specific Archives

| Domain | Repository | Notable Datasets |
|--------|-----------|-----------------|
| Genomics | NCBI GEO, ENA | Gene expression, sequencing data |
| Astronomy | NASA archives, SDSS | Sky surveys, spectral data |
| Economics | FRED, World Bank, IMF | Time series, macro indicators |
| Climate | NOAA, CMIP6 | Temperature, precipitation records |
| Linguistics | LDC, CLARIN | Corpora, treebanks |
| Medical | PhysioNet, MIMIC | Clinical records, ECG/EEG |
| Chemistry | PubChem, ChEMBL | Molecular structures, bioassays |

## Searching for Datasets

### Kaggle CLI

```bash
# Install and configure
pip install kaggle
# Place kaggle.json in ~/.kaggle/

# Search datasets
kaggle datasets list -s "sentiment analysis" --sort-by votes
kaggle datasets list -s "medical imaging" --file-type csv --min-size 100MB

# Get dataset details
kaggle datasets metadata -d stanford/imdb-review-dataset

# Download dataset
kaggle datasets download -d stanford/imdb-review-dataset -p ./data/
unzip ./data/imdb-review-dataset.zip -d ./data/imdb/

# Download competition data
kaggle competitions download -c titanic -p ./data/
```

### Hugging Face Datasets

```python
from datasets import load_dataset, list_datasets

# Search for datasets by task
from huggingface_hub import HfApi
api = HfApi()
datasets = api.list_datasets(
    search="scientific papers",
    sort="downloads",
    direction=-1,
    limit=20
)
for ds in datasets:
    print(f"{ds.id}: {ds.downloads} downloads")

# Load a dataset (with streaming for large datasets)
dataset = load_dataset("scientific_papers", "arxiv", streaming=True)

# Inspect structure
print(dataset["train"].features)
print(f"Number of examples: {dataset['train'].num_rows}")

# Load specific split and subset
validation = load_dataset(
    "scientific_papers", "arxiv",
    split="validation[:1000]"
)
```

### Google Dataset Search (Programmatic)

```python
import requests
from bs4 import BeautifulSoup

def search_google_datasets(query, num_results=10):
    """Search Google Dataset Search and extract results."""
    url = f"https://datasetsearch.research.google.com/search"
    params = {"query": query, "docid": ""}
    # Note: Google Dataset Search does not have an official API
    # Use the web interface or alternative approaches
    print(f"Search at: {url}?query={query.replace(' ', '+')}")
    return url
```

### Zenodo API

```python
import requests

def search_zenodo(query, resource_type="dataset", size=10):
    """Search Zenodo for research datasets."""
    url = "https://zenodo.org/api/records"
    params = {
        "q": query,
        "type": resource_type,
        "size": size,
        "sort": "mostrecent",
        "access_right": "open"
    }
    response = requests.get(url, params=params)
    results = response.json()

    for hit in results.get("hits", {}).get("hits", []):
        meta = hit["metadata"]
        print(f"Title: {meta['title']}")
        print(f"DOI: {meta.get('doi', 'N/A')}")
        print(f"License: {meta.get('license', {}).get('id', 'N/A')}")
        print(f"Size: {sum(f['size'] for f in hit.get('files', []))/1e6:.1f} MB")
        print("---")

    return results
```

## Dataset Evaluation Checklist

Before using a dataset in research, verify the following:

### Quality Assessment

- **Completeness**: What percentage of values ar
05

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__dataset-finder-guide.json · Report an issue / request a re-scan
06

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-08e1ba289846fdSAFEB89first audit
07

Questions

What does the Dataset Finder Guide skill do?

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Is Dataset Finder Guide safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Dataset Finder Guide access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

Which assistants does Dataset Finder Guide work with?

Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.

How current is this page?

The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement