Automated Review GuideSAFE
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Overview
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
e1ba289846fdOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| openclaw | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: automated-review-guide
description: "AI-assisted peer review tools, workflows, and quality standards"
metadata:
openclaw:
emoji: "🤖"
category: "research"
subcategory: "paper-review"
keywords: ["automated review", "AI peer review", "LLM review", "review quality", "manuscript screening", "editorial workflow"]
source: "wentor-research-plugins"
---
# Automated Review Guide
A skill for leveraging AI-assisted tools in the peer review process, covering both author-side self-review and editor-side manuscript screening. Addresses tool selection, prompt engineering for review tasks, limitations and biases of LLM-generated reviews, quality assurance workflows, and ethical guidelines for AI use in peer review.
## Overview of AI in Peer Review
### Current Landscape
AI-assisted peer review tools operate at multiple stages of the publication pipeline. Understanding where automation adds genuine value and where it introduces risk is essential for responsible adoption.
```
Where AI assists in peer review:
Author-side (pre-submission):
- Grammar and style checking (Grammarly, Writefull)
- Statistical result verification (statcheck, GRIM/SPRITE)
- Reference completeness checking
- Plagiarism detection (iThenticate, Turnitin)
- Readability scoring
- Structural completeness (IMRAD compliance)
Editor-side (triage and assignment):
- Desk rejection screening (scope, quality threshold)
- Reviewer matching (expertise alignment)
- Conflict of interest detection
- Duplicate submission detection
- Plagiarism and image manipulation screening
Reviewer-side (review assistance):
- Paper summarization for rapid assessment
- Statistical claim verification
- Reference checking (do cited papers support claims)
- Comparison with related work
- Structured review template generation
Post-review:
- Decision consistency analysis
- Review quality assessment
- Revision compliance checking
```
## Self-Review with AI Before Submission
### Structured Self-Review Prompts
```
Pre-submission AI review checklist:
1. Abstract completeness check:
Prompt: "Analyze this abstract. Does it contain:
(a) background/motivation, (b) research gap,
(c) methodology summary, (d) key results with
numbers, (e) conclusion/implication? Identify
any missing elements."
2. Claim-evidence alignment:
Prompt: "For each claim in the Discussion section,
identify the specific result (table, figure, or
statistical test) that supports it. Flag any claims
without corresponding evidence in the Results."
3. Methods reproducibility:
Prompt: "Read the Methods section and list every
piece of information that another researcher would
need to replicate this study. Identify any gaps:
missing sample sizes, unspecified parameters,
ambiguous procedures, unnamed software versions."
4. Statistical reporting:
Prompt: "Check all statistical results in this paper
for completeness. Each test should report: test name,
test statistic, degrees of freedom, p-value, and
effect size. List any incomplete reports."
5. Reference audit:
Prompt: "For each citation in the Introduction, verify
that the cited claim matches the in-text description.
Flag any cases where the citation might not support
the specific claim being made."
```
### Automated Statistical Checking
```python
import re
def check_statistical_reporting(text):
"""
Check for common statistical reporting issues.
Verifies:
- p-values are reported with test statistics
- Degrees of freedom are included
- Effect sizes are reported
- Exact p-values (not just p < .05)
"""
issues = []
# Find p-value reports
p_pattern = r'p\s*[<=<>]\s*\.?\d+'
p_matches = re.finditer(p_pattern, text, re.IGNORECASE)
for match in p_matches:
# Check context (100 chars before) for test statistic
start = max(0, match.start() - 100)
context = text[start:match.end()]
has_test_stat = any(
stat in context for stat in
["t(", "F(", "chi", "r(", "r =", "z =",
"U =", "W =", "H(", "d =", "eta"]
)
if not has_test_stat:
issues.append({
"location": match.start(),
"text": context[-50:],
"issue": "p-value without test statistic"
})
# Check for "p < .05" without exact values
vague_p = re.findall(r'p\s*<\s*\.05(?!\d)', text)
if vague_p:
issues.append({
"issue": f"Found {len(vague_p)} instances of 'p < .05' "
"without exact p-values. APA recommends exact values."
})
return issues
```
## Limitations and Biases of AI Reviews
### Known Failure Modes
```
AI review limitations to be aware of:
1. Hallucinated references:
- LLMs may claim a paper cites X when it does not
- Always verify any reference claims made by AI
- LLMs cannot actually read PDFs behind paywalls
2. False confidence in statistical judgments:
- LLMs may incorrectly flag valid statistical approaches
- They may miss subtle errors that require domain expertise
- Statistical verification tools (statcheck) are more reliable
3. Novelty assessment failures:
- LLMs have knowledge cutoff dates and cannot assess true novelty
- They may flag well-known methods as novel or novel methods as
well-known, depending on training data coverage
- Human expertise is essential for novelty evaluation
4. Disciplinary bias:
- LLMs trained primarily on English text from well-resourced fields
- May apply STEM conventions to humanities papers inappropriately
- May not recognize valid methodologies in underrepresented fields
5. Sycophancy:
- Tendency to agree with the framing of the prompt
- "Review this excellent paper" vs "Review this paper" yields
systematically different feedback
- Use neutral prompts and ask for both strengths and weaknesses
6. ReproTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e1ba289846fdfull audit observations/trust-audit/skill/brycewang-stanford__automated-review-guide.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | e1ba289846fd | SAFE | B | 89 | first audit |
Questions
What does the Automated Review Guide skill do?
🔬 A curated collection of 23,000+ agent skills for empirical research across 8 social science disciplines. | 精选 23,000+ AI Agent 技能库,覆盖8大社会科学学科的实证研究。CoPaper.AI 20分钟完成一篇可复现的规范实证论文,并支持用户上传 Skills。-- Maintained by CoPaper.AI from Stanford REAP.
Is Automated Review Guide safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Automated Review Guide access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Automated Review Guide work with?
Its documentation mentions openclaw. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (e1ba289846fd), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.