Atlas / Skills / affaan-m / Scientific Thinking Scholar Evaluation

Scientific Thinking Scholar EvaluationSAFE

skills/affaan-m/scientific-thinking-scholar-evaluation

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
MIT
Stars
264,831
01

Overview

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Read from source at commit fb3fb10d9622OBSERVED · 2026-09-22
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: scholar-evaluation
description: 論文、提案書、文献レビュー、方法論セクション、証拠の質、引用サポート、研究論文フィードバックのための構造化された学術的作業評価。
origin: community
---

# Scholar Evaluation

このスキルを使用して、再現可能なルーブリックで学術的または科学的な作業を評価します。

## 使用するタイミング

- 研究論文、提案書、論文章、または文献レビューのレビュー。
- 主張が引用された証拠によって支持されているかの確認。
- 方法論、研究デザイン、分析、または限界の評価。
- 品質または関連性について2つ以上の論文を比較する。
- 改訂のための構造化されたフィードバックの作成。

## 評価の範囲

まず成果物を特定します:

- 実証的研究論文
- 理論論文
- 技術レポート
- システマティックまたはナラティブ文献レビュー
- 研究提案書
- 論文または学位論文の章
- 学会アブストラクトまたは短い論文

次に範囲を選択します:

- **包括的**: すべてのルーブリック次元
- **的を絞った**: 1つまたは2つの次元(方法や引用など)
- **比較的**: 同じルーブリックに対して複数の作品をランク付け

## ルーブリック

該当する各次元を1から5でスコアリング:

- 5: 優れている;明確、厳密で、出版の準備ができている
- 4: 良好;軽微な改善が必要
- 3: 適切;意味のあるギャップがあるが使用可能
- 2: 弱い;実質的な改訂が必要
- 1: 不十分;主要な有効性または明確性の問題

適用されない次元には `N/A` を使用します。

### 1. 問題と研究質問

- 問題は明確かつ具体的か?
- 貢献は意義があるか?
- 範囲と前提は明示的か?
- 質問は主張された貢献と一致しているか?

### 2. 文献とコンテキスト

- 関連する先行研究はカバーされているか?
- 作業はソースを単に列挙するのではなく統合しているか?
- ギャップは正確に特定されているか?
- 最近のソースと基礎的なソースはバランスが取れているか?

### 3. 方法論

- 方法は研究質問に答えるか?
- デザインの選択は正当化されているか?
- 変数、データセット、参加者、または材料は明確に説明されているか?
- 別の研究者が研究を再現できるか?
- 倫理的および実際的な制約は認識されているか?

### 4. データと証拠

- データソースは信頼性があり適切か?
- サンプルサイズまたはコーパスのカバレッジは十分か?
- 包含、除外、前処理の決定は文書化されているか?
- 欠損データとバイアスのリスクは議論されているか?

### 5. 分析

- 統計的、定性的、または計算的方法は適切か?
- ベースラインとコントロールは公正か?
- 必要に応じて不確実性、感度、またはロバスト性のチェックが含まれているか?
- 代替の説明は考慮されているか?

### 6. 結果と解釈

- 結果は明確に提示されているか?
- 主張は証拠の範囲内に留まっているか?
- 図、表、メトリクスは理解しやすいか?
- 否定的または帰無的な結果は正直に扱われているか?

### 7. 限界と妥当性への脅威

- 限界は一般的でなく具体的か?
- 内部、外部、構成概念、結論の妥当性リスクは対処されているか?
- 論文は投機的な結果と実証された結果を区別しているか?

### 8. 文章と構造

- 論証は理解しやすいか?
- セクションは研究質問を中心に構成されているか?
- 定義と表記は明確か?
- トーンは正確で学術的か?

### 9. 引用

- 引用された論文はそれに付けられた主張を支持しているか?
- 可能な限り一次ソースが使用されているか?
- レビューはレビューとしてラベル付けされているか?
- プレプリントはプレプリントとしてラベル付けされているか?
- 引用メタデータとリンクは正しいか?

## レビュープロセス

1. 主張された貢献についてアブストラクト、序論、図、結論を読む。
2. 証拠の質について方法と結果を読む。
3. 最も強い主張を引用されたソースと照合する。
4. 該当する各次元をスコアリングする。
5. 重大なブロッカーと改訂提案を分離する。
6. 具体的な次の編集で終了する。

## 出力テンプレート

```markdown
# Scholar Evaluation: <成果物>

## 総合評価

- 総合スコア: <1-5 または N/A>
- 信頼度: <高 | 中 | 低>
- サマリー: <3-5 文>

## 次元スコア

| 次元 | スコア | 証拠 | 改訂優先度 |
| --- | ---: | --- | --- |
| 問題と質問 |  |  |  |
| 文献とコンテキスト |  |  |  |
| 方法論 |  |  |  |
| データと証拠 |  |  |  |
| 分析 |  |  |  |
| 結果と解釈 |  |  |  |
| 限界 |  |  |  |
| 文章と構造 |  |  |  |
| 引用 |  |  |  |

## 重大な問題

## 推奨される改訂

## 必要な証拠確認
```

## 落とし穴

- スコアを具体的なフィードバックの代替として使用しない。
- 範囲外の次元の省略で論文にペナルティを与えない。
- 引用数、出版場所、または著者の評判を品質の証拠として扱わない。
- アブストラクトに現れるからといって根拠のない主張を受け入れない。
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-09-22 · audit v0.4.1 · source sha fb3fb10d9622full audit observations/trust-audit/skill/affaan-m__scientific-thinking-scholar-evaluation.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-09-22fb3fb10d9622SAFEB89first audit
05

Questions

What does the Scientific Thinking Scholar Evaluation skill do?

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Is Scientific Thinking Scholar Evaluation safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Scientific Thinking Scholar Evaluation access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (fb3fb10d9622), read on 2026-09-22. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement