Ai Scientist V2 GuideSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Install
Commands as the repository documents them. They are shown, not run.
git clone https://github.com/SakanaAI/AI-Scientist-v2.git
pip install -r requirements.txt
Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: ai-scientist-v2-guide
description: "Automated scientific discovery via agentic tree search by Sakana AI"
metadata:
openclaw:
emoji: "🧪"
category: "research"
subcategory: "automation"
keywords: ["scientific-discovery", "automation", "tree-search", "paper-generation", "experiment-design", "sakana-ai"]
source: "https://github.com/SakanaAI/AI-Scientist-v2"
---
# AI Scientist v2 Guide
## Overview
AI-Scientist-v2 is an open-source system developed by Sakana AI with over 2,000 GitHub stars that automates the full scientific research pipeline -- from idea generation through experimentation to paper writing. Building on the original AI Scientist, version 2 introduces an agentic tree search approach that systematically explores the space of research ideas, designs and runs experiments, analyzes results, and produces workshop-level scientific papers with minimal human intervention.
The key innovation in v2 is the tree search mechanism. Rather than pursuing a single research direction linearly, the system maintains a tree of possible research trajectories. At each node, the agent can branch into multiple experimental variations, evaluate the results, and prune unpromising directions while doubling down on successful ones. This mirrors how experienced researchers navigate the research landscape -- exploring broadly at first, then focusing resources on the most promising leads.
AI-Scientist-v2 has demonstrated the ability to generate novel, valid research papers in machine learning subfields including diffusion models, language model training, and optimization. While the generated papers are currently at workshop acceptance level, the system represents a significant step toward autonomous scientific discovery and is an invaluable tool for researchers looking to automate the more mechanical aspects of their research workflow.
## Installation and Setup
```bash
# Clone the repository
git clone https://github.com/SakanaAI/AI-Scientist-v2.git
cd AI-Scientist-v2
# Create a conda environment
conda create -n ai-scientist python=3.11
conda activate ai-scientist
# Install dependencies
pip install -r requirements.txt
```
### Prerequisites
AI-Scientist-v2 requires several components:
```bash
# LLM API access (required for ideation, analysis, and writing)
export OPENAI_API_KEY=$OPENAI_API_KEY
# Or Anthropic
export ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY
# GPU access for running ML experiments
# Recommended: at least one NVIDIA GPU with 24GB+ VRAM
# LaTeX installation for paper compilation
# Ubuntu/Debian
sudo apt-get install texlive-full
# macOS
brew install --cask mactex
```
### Configuration
Set up your research configuration:
```python
# config.yaml
llm:
provider: "openai"
model: "gpt-4o"
temperature: 0.7
search:
max_depth: 5 # Maximum tree depth
branching_factor: 3 # Number of branches per node
pruning_threshold: 0.3 # Prune branches below this score
experiment:
gpu_ids: [0, 1] # Available GPUs
timeout_hours: 2 # Max time per experiment
num_seeds: 3 # Random seeds per experiment
paper:
template: "icml" # Paper template (icml, neurips, iclr)
max_pages: 8 # Maximum paper length
```
## Core Research Pipeline
### Phase 1: Idea Generation
The system generates research ideas by analyzing existing literature and identifying gaps or extensions:
```python
from ai_scientist import IdeaGenerator
generator = IdeaGenerator(
research_area="efficient_transformers",
seed_papers=[
"path/to/related_paper_1.pdf",
"path/to/related_paper_2.pdf",
],
num_ideas=10,
)
ideas = generator.generate()
for idea in ideas:
print(f"Title: {idea.title}")
print(f"Hypothesis: {idea.hypothesis}")
print(f"Novelty score: {idea.novelty_score}")
print(f"Feasibility score: {idea.feasibility_score}")
```
### Phase 2: Agentic Tree Search
The tree search mechanism explores the research space systematically:
```python
from ai_scientist import TreeSearchResearcher
researcher = TreeSearchResearcher(
idea=ideas[0], # Start with the top-ranked idea
base_code="templates/efficient_transformer/",
config="config.yaml",
)
# Run the tree search
result = researcher.run()
# The search tree tracks all explorations
print(f"Tree depth reached: {result.max_depth}")
print(f"Total experiments run: {result.total_experiments}")
print(f"Best result: {result.best_node.metrics}")
```
The tree search works as follows:
1. **Root node**: The initial research idea and baseline implementation
2. **Expansion**: At each node, the agent proposes 2-4 modifications (hyperparameter changes, architectural tweaks, new training strategies)
3. **Evaluation**: Each modification is implemented and evaluated experimentally
4. **Selection**: Promising branches are selected for further exploration using UCB (Upper Confidence Bound) or similar strategies
5. **Pruning**: Branches that underperform the baseline or show diminishing returns are pruned
### Phase 3: Experiment Execution
Experiments are executed in isolated environments with proper controls:
```python
# Each experiment node contains:
class ExperimentNode:
hypothesis: str # What we're testing
code_changes: list # Specific code modifications
config_changes: dict # Hyperparameter changes
results: dict # Experimental results
analysis: str # LLM-generated analysis
children: list # Branch experiments
```
The system automatically handles experiment boilerplate including random seed management, metric logging, checkpoint saving, and result visualization. Each experiment is run with multiple seeds to ensure statistical significance.
### Phase 4: Paper Generation
After the tree search completes, the system generates a scientific paper:
```python
from ai_scientist import PaperWriter
writer = PaperWriter(
research_result=result,
template="neurips",
sections=[
"Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__ai-scientist-v2-guide.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Ai Scientist V2 Guide skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Ai Scientist V2 Guide safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Ai Scientist V2 Guide access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Ai Scientist V2 Guide work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.