aaai-experiments
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human evaluation, AI-for-Social-Impact and alignment/safety evidence, compute and cost reporting, and reproducibility-checklist alignment for Phase-1 survival.
npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aaai-experiments --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# AAAI Experiments Use this before submission to ensure empirical evidence supports the AI contribution. AAAI reviewers may come from adjacent AI subfields, so experiments must be interpretable beyond one benchmark community. ## Experiment audit - Map every experimental block to a claim in the introduction. - Compare against strong, recent, and fairly tuned baselines. - Include ablations that isolate mechanisms rather than removing multiple components at once. - Report uncertainty, variance, and statistical tests when small differences matter. - Test robustness to data split, prompt, seed, environment, user population, or distribution shift when relevant. - For human evaluation, document task, instructions, annotator pool, quality control, aggregation, and ethics/IRB status. - Report compute, hardware, data access, model size, and training/inference cost. ## Claim-to-evidence ledger Build this table before adding new experiments. It keeps the AAAI evidence package aligned with the main text and with the reproducibility checklist. | Manuscript claim | Required evidence | Phase-1 risk if missing | Checklist hook | | --- | --- | --- | --- | | New AI capability | benchmark + qualitativ
- Experiment audit
- Claim-to-evidence ledger
- AAAI-specific review pressure
- Pre-rebuttal freeze rule
- Evidence triage table
- Common AAAI experiment rejects
- Worked vignette
- Output format
What does the aaai-experiments skill do?
Use when designing or auditing AAAI experiments for the broad-AI program committee, including baselines, ablations, statistical significance, robustness, human evaluation, AI-for-Social-Impact and alignment/safety evidence, compute and cost reporting, and reproducibility-checklist alignment for Phase-1 survival.
How do I install it?
Run `npx skills add brycewang-stanford/Awesome-Journal-Skills --skill aaai-experiments --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From brycewang-stanford/Awesome-Journal-Skills, a repository with 909 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.