Atlas / Skills / brycewang-stanford / Paper Critique Framework

Paper Critique FrameworkSAFE

skills/brycewang-stanford/paper-critique-framework

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
1 documented
License
NOASSERTION
Stars
4,537
01

Overview

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Read from source at commit e1ba289846fdOBSERVED · 2026-10-08
02

Host compatibility

What the documentation claims. We have not run a compatibility test.

HostStatusNotes
openclawmentioned
03

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: paper-critique-framework
description: "Structured framework for writing peer review reports and paper critiques"
metadata:
  openclaw:
    emoji: "📋"
    category: "research"
    subcategory: "paper-review"
    keywords: ["peer review", "paper critique", "referee report", "academic review", "manuscript evaluation", "constructive feedback"]
    source: "wentor-research-plugins"
---

# Paper Critique and Peer Review Framework

## Overview

Writing constructive peer reviews is a core academic skill. This framework provides a systematic approach to evaluating manuscripts — from initial read-through to the final referee report. It covers what reviewers should assess, how to structure feedback, and how to calibrate between different review outcomes (accept, revise, reject). Applicable to conference papers, journal articles, and internal lab reviews.

## The Three-Pass Review Method

### Pass 1: Orientation (15-20 minutes)

Read only these elements:
- Title, abstract, and keywords
- Introduction (first and last paragraphs)
- Section headings and figure captions
- Conclusion

After Pass 1, answer:
```
□ What is the main claim?
□ What type of contribution? (empirical, theoretical, system, survey)
□ Is it within the venue's scope?
□ Does the abstract accurately represent the content?
□ Initial impression: novel or incremental?
```

### Pass 2: Detailed Read (60-90 minutes)

Read the full paper. Annotate as you go:

```markdown
Annotation symbols:
  ? = I don't understand this
  ! = This is interesting / strong point
  X = I disagree / see a problem
  → = This needs more evidence or justification
  ≈ = This is similar to [existing work] — check novelty
```

Focus on:
- **Claims vs. evidence**: Is every major claim supported by data?
- **Methodology**: Are the methods appropriate for the research question?
- **Experimental design**: Are baselines fair? Are ablations sufficient?
- **Figures and tables**: Do they support the narrative? Are they readable?
- **Writing quality**: Is it clear, concise, and well-organized?

### Pass 3: Verification (30-60 minutes)

For papers you're seriously evaluating:
- Check key references — do they say what the authors claim?
- Verify mathematical derivations (spot-check, not exhaustive)
- Examine statistical claims (p-values, confidence intervals, effect sizes)
- Check for cherry-picking in results (only best runs? selected metrics?)
- Look for missing baselines that should have been compared

## Review Report Structure

```markdown
## Summary (3-5 sentences)
[Describe what the paper does, the approach, and the main finding.
 Demonstrate that you understood the paper.]

## Strengths (bulleted list)
- S1: [Specific strength with evidence from the paper]
- S2: [Another strength]
- S3: [Another strength]

## Weaknesses (bulleted list, ordered by severity)
- W1 (Major): [Specific weakness + why it matters + suggestion to fix]
- W2 (Major): [Another major weakness]
- W3 (Minor): [A less critical issue]
- W4 (Minor): [Another minor issue]

## Questions for Authors
- Q1: [Something you'd like clarified]
- Q2: [A concern that the authors might be able to address]

## Detailed Comments
[Page/line-specific comments, typos, suggestions]

## Overall Assessment
Recommendation: [Strong Accept / Accept / Weak Accept / Borderline /
                  Weak Reject / Reject / Strong Reject]
Confidence: [High / Medium / Low]
```

## Assessment Criteria by Dimension

| Dimension | Questions to Ask | Weight |
|-----------|-----------------|--------|
| **Novelty** | Is the idea new? Is the contribution beyond incremental? | High |
| **Significance** | Would this matter to the community? Does it advance the field? | High |
| **Soundness** | Are the methods correct? Are conclusions supported? | High |
| **Clarity** | Is it well-written? Can it be understood and reproduced? | Medium |
| **Completeness** | Are related works covered? Are experiments thorough? | Medium |
| **Reproducibility** | Could someone replicate this? Code/data available? | Medium |

### Calibration Guide

```
Strong Accept: Significant contribution, technically sound, well-written.
  Would be a highlight of the venue.

Accept: Solid contribution with minor issues. Advances the field.
  Worth publishing as-is or with minor revisions.

Weak Accept: Has merit but notable weaknesses. Contribution is real but modest.
  Borderline for this venue; would be accepted at a less selective venue.

Borderline: Equal arguments for and against. Significant weaknesses offset
  by some novelty. Depends on other reviews.

Weak Reject: Interesting direction but fundamental issues not addressed.
  Major revisions needed that likely require a new submission cycle.

Reject: Significant problems in novelty, soundness, or relevance.
  Not suitable for this venue even with revisions.

Strong Reject: Fundamental flaws. Clearly below threshold.
```

## Common Review Pitfalls to Avoid

| Pitfall | Better Approach |
|---------|----------------|
| "The writing needs improvement" (vague) | Give 2-3 specific examples with suggested fixes |
| Rejecting for not solving YOUR problem | Evaluate the paper on its own stated goals |
| Demanding impossible experiments | Suggest feasible improvements within scope |
| Ignoring supplementary material | Check appendix — authors may have addressed your concern |
| Being harsh without being constructive | Every weakness should include a suggestion for improvement |
| Reviewing too quickly | Block dedicated time; a rushed review harms both authors and science |
| Citing only your own work as "missing" | Only cite if genuinely relevant, not self-promotion |

## Reviewing Different Paper Types

### Empirical Papers
- Are datasets described completely? (Size, source, splits, preprocessing)
- Are baselines appropriate and fairly tuned?
- Statistical significance: error bars, multiple runs, significance tests
- Ablation studies: which components contribute to the gain?

### Systems Papers
- Is the system actually 
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__paper-critique-framework.json · Report an issue / request a re-scan
05

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-08e1ba289846fdSAFEB89first audit
06

Questions

What does the Paper Critique Framework skill do?

🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.

Is Paper Critique Framework safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Paper Critique Framework access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

Which assistants does Paper Critique Framework work with?

Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.

How current is this page?

The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement