Atlas / Skills / brycewang-stanford / Ai Scientist V2 Guide

Ai Scientist V2 GuideSAFE

skills/brycewang-stanford/ai-scientist-v2-guide

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
1 documented
License
NOASSERTION
Stars
4,535
01

Overview

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Read from source at commit e1ba289846fdOBSERVED · 2026-10-08
02

Install

Commands as the repository documents them. They are shown, not run.

git clone https://github.com/SakanaAI/AI-Scientist-v2.git
pip install -r requirements.txt
03

Host compatibility

What the documentation claims. We have not run a compatibility test.

HostStatusNotes
openclawmentioned
04

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: ai-scientist-v2-guide
description: "Automated scientific discovery via agentic tree search by Sakana AI"
metadata:
  openclaw:
    emoji: "🧪"
    category: "research"
    subcategory: "automation"
    keywords: ["scientific-discovery", "automation", "tree-search", "paper-generation", "experiment-design", "sakana-ai"]
    source: "https://github.com/SakanaAI/AI-Scientist-v2"
---

# AI Scientist v2 Guide

## Overview

AI-Scientist-v2 is an open-source system developed by Sakana AI with over 2,000 GitHub stars that automates the full scientific research pipeline -- from idea generation through experimentation to paper writing. Building on the original AI Scientist, version 2 introduces an agentic tree search approach that systematically explores the space of research ideas, designs and runs experiments, analyzes results, and produces workshop-level scientific papers with minimal human intervention.

The key innovation in v2 is the tree search mechanism. Rather than pursuing a single research direction linearly, the system maintains a tree of possible research trajectories. At each node, the agent can branch into multiple experimental variations, evaluate the results, and prune unpromising directions while doubling down on successful ones. This mirrors how experienced researchers navigate the research landscape -- exploring broadly at first, then focusing resources on the most promising leads.

AI-Scientist-v2 has demonstrated the ability to generate novel, valid research papers in machine learning subfields including diffusion models, language model training, and optimization. While the generated papers are currently at workshop acceptance level, the system represents a significant step toward autonomous scientific discovery and is an invaluable tool for researchers looking to automate the more mechanical aspects of their research workflow.

## Installation and Setup

```bash
# Clone the repository
git clone https://github.com/SakanaAI/AI-Scientist-v2.git
cd AI-Scientist-v2

# Create a conda environment
conda create -n ai-scientist python=3.11
conda activate ai-scientist

# Install dependencies
pip install -r requirements.txt
```

### Prerequisites

AI-Scientist-v2 requires several components:

```bash
# LLM API access (required for ideation, analysis, and writing)
export OPENAI_API_KEY=$OPENAI_API_KEY
# Or Anthropic
export ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY

# GPU access for running ML experiments
# Recommended: at least one NVIDIA GPU with 24GB+ VRAM

# LaTeX installation for paper compilation
# Ubuntu/Debian
sudo apt-get install texlive-full

# macOS
brew install --cask mactex
```

### Configuration

Set up your research configuration:

```python
# config.yaml
llm:
  provider: "openai"
  model: "gpt-4o"
  temperature: 0.7

search:
  max_depth: 5          # Maximum tree depth
  branching_factor: 3   # Number of branches per node
  pruning_threshold: 0.3  # Prune branches below this score

experiment:
  gpu_ids: [0, 1]       # Available GPUs
  timeout_hours: 2      # Max time per experiment
  num_seeds: 3          # Random seeds per experiment

paper:
  template: "icml"      # Paper template (icml, neurips, iclr)
  max_pages: 8          # Maximum paper length
```

## Core Research Pipeline

### Phase 1: Idea Generation

The system generates research ideas by analyzing existing literature and identifying gaps or extensions:

```python
from ai_scientist import IdeaGenerator

generator = IdeaGenerator(
    research_area="efficient_transformers",
    seed_papers=[
        "path/to/related_paper_1.pdf",
        "path/to/related_paper_2.pdf",
    ],
    num_ideas=10,
)

ideas = generator.generate()
for idea in ideas:
    print(f"Title: {idea.title}")
    print(f"Hypothesis: {idea.hypothesis}")
    print(f"Novelty score: {idea.novelty_score}")
    print(f"Feasibility score: {idea.feasibility_score}")
```

### Phase 2: Agentic Tree Search

The tree search mechanism explores the research space systematically:

```python
from ai_scientist import TreeSearchResearcher

researcher = TreeSearchResearcher(
    idea=ideas[0],  # Start with the top-ranked idea
    base_code="templates/efficient_transformer/",
    config="config.yaml",
)

# Run the tree search
result = researcher.run()

# The search tree tracks all explorations
print(f"Tree depth reached: {result.max_depth}")
print(f"Total experiments run: {result.total_experiments}")
print(f"Best result: {result.best_node.metrics}")
```

The tree search works as follows:

1. **Root node**: The initial research idea and baseline implementation
2. **Expansion**: At each node, the agent proposes 2-4 modifications (hyperparameter changes, architectural tweaks, new training strategies)
3. **Evaluation**: Each modification is implemented and evaluated experimentally
4. **Selection**: Promising branches are selected for further exploration using UCB (Upper Confidence Bound) or similar strategies
5. **Pruning**: Branches that underperform the baseline or show diminishing returns are pruned

### Phase 3: Experiment Execution

Experiments are executed in isolated environments with proper controls:

```python
# Each experiment node contains:
class ExperimentNode:
    hypothesis: str          # What we're testing
    code_changes: list       # Specific code modifications
    config_changes: dict     # Hyperparameter changes
    results: dict            # Experimental results
    analysis: str            # LLM-generated analysis
    children: list           # Branch experiments
```

The system automatically handles experiment boilerplate including random seed management, metric logging, checkpoint saving, and result visualization. Each experiment is run with multiple seeds to ensure statistical significance.

### Phase 4: Paper Generation

After the tree search completes, the system generates a scientific paper:

```python
from ai_scientist import PaperWriter

writer = PaperWriter(
    research_result=result,
    template="neurips",
    sections=[
        "
05

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__ai-scientist-v2-guide.json · Report an issue / request a re-scan
06

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-08e1ba289846fdSAFEB89first audit
07

Questions

What does the Ai Scientist V2 Guide skill do?

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Is Ai Scientist V2 Guide safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Ai Scientist V2 Guide access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

Which assistants does Ai Scientist V2 Guide work with?

Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.

How current is this page?

The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement