ai-scientist-evaluator
Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks. Trigger when the user asks to evaluate notebooks, code, figures, analyses, manuscripts, software, or final reports produced by AI scientists; compare multiple AI scientists on the same task; judge publication readiness; or audit rigor, reproducibility, novelty, and task completion. Do not use this skill to perform the original research task itself unless the user is explicitly asking for a reviewer-style audit of already produced outp
npx skills add BioTender-max/awesome-bio-agent-skills --skill ai-scientist-evaluator --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# AI Scientist Evaluator Use this skill when Codex should behave like a skeptical reviewer panel rather than a research generator. Evaluate completed outputs, not just plans. ## Instructions 1. Confirm the request is evaluative. Use this skill to audit or compare existing outputs, not to perform the original research task. 2. Restate the exact task in one or two sentences so the review stays anchored to the real objective and required deliverables. 3. Inventory the submitted artifacts and note what is missing. Prefer primary artifacts over summaries: - notebooks, code, scripts, and workflow files - environment files, package versions, and runtime logs - figures, tables, and manuscript drafts - data provenance, accession lists, database versions, and citations - benchmark results, hardware notes, and task constraints 4. Choose the closest task profile from [`references/task_profiles.md`](references/task_profiles.md) and load the matching weights from [`assets/default_weight_profiles.yaml`](assets/default_weight_profiles.yaml). Use the primary scientific profile first for composite tasks, then add manuscript comments as a secondary layer. 5. Review with a four-person panel and synthe
- Instructions
- Quick Reference
- Input Requirements
- Output
- Quality Gates
- Examples
- Example 1: Compare five AI scientist submissions
- Example 2: Audit one submission for publication readiness
- Example 3: Rank finished JSON evaluations
- Troubleshooting
- Related Skills
python scripts/aggregate_reviews.py review_a.json review_b.json --out_md leaderboard.md
What does the ai-scientist-evaluator skill do?
Critically review, score, compare, and rank one or more AI scientist outputs for biology, bioinformatics, computational life science, or adjacent research tasks. Trigger when the user asks to evaluate notebooks, code, figures, analyses, manuscripts, software, or final reports produced by AI scientists; compare multiple AI scientists on the same task; judge publication readiness; or audit rigor, reproducibility, novelty, and task completion. Do not use this skill to perform the original research task itself unless the user is explicitly asking for a reviewer-style audit of already produced outp
How do I install it?
Run `npx skills add BioTender-max/awesome-bio-agent-skills --skill ai-scientist-evaluator --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From BioTender-max/awesome-bio-agent-skills, a repository with 135 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
