Agent skill

ln-34-benchmark-comparator

Compares tools or implementations through reproducible A/B workloads, correctness oracles, and controlled measurements. Use to choose alternatives; not to optimize a known bottleneck.

levnikolaevich528★ · 1 repos on radarProfile →
claude-codecodexMIT
Install
npx skills add levnikolaevich/claude-code-skills --skill ln-34-benchmark-comparator --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 1
SKILL.md size: 11 KB
Bundled scripts: none
Path: plugins/optimization-suite/skills/ln-34-benchmark-comparator/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 528
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# Benchmark Comparator **Goal:** Compare alternatives under controlled, reproducible conditions. Correctness comes before speed, and measured data must remain separate from estimates, setup cost, and interpretation. **Execution contract:** Treat the ordered checkbox workflow below as this skill's Definition of Done. Work through every item in order, and mark it complete only when its action and required evidence are complete. `N/A`, skipped, unavailable, or delegated items remain incomplete. Before returning, apply this skill's verdict, decision, and approval rules to every incomplete item and prepend **Checklist: X/Y complete**<br>**Incomplete: None | section/item — reason; outcome impact; exact next action**; list every incomplete item. ## Tool Routing | Need | Preferred tool | Use it when | Fallback | |---|---|---|---| | Canonical workload and oracle | Repository fixtures, tests, expected diffs, schemas, or independently specified outcomes | Defining what success means before either candidate runs | Create the smallest deterministic fixture that represents the decision | | Isolation | Clean Git worktrees, temporary directories, controlled environment, fixed seeds, and resettable

What's inside
Steps it walks through
  1. Tool Routing
  2. Evidence Rules
  3. Checklist
  4. 1. Define the Decision and Experiment
  5. 2. Build a Symmetric Harness
  6. 3. Execute and Capture Evidence
  7. 4. Analyze Validity and Results
  8. 5. Decide, Preserve, and Clean Up
  9. Output Contract
More from claude-code-skills
All skills →
About this skill
What does the ln-34-benchmark-comparator skill do?

Compares tools or implementations through reproducible A/B workloads, correctness oracles, and controlled measurements. Use to choose alternatives; not to optimize a known bottleneck.

How do I install it?

Run `npx skills add levnikolaevich/claude-code-skills --skill ln-34-benchmark-comparator --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From levnikolaevich/claude-code-skills, a repository with 528 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going