ai-evals
Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating test cases, designing rubrics, or trying to systematically measure AI output quality.
npx skills add majiayu000/claude-skill-registry --skill ai-evals-refoundai-lenny-skills-2 --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
# AI Evals Help the user create systematic evaluations for AI products using insights from AI practitioners. ## How to Help When the user asks for help with AI evals: 1. **Understand what they're evaluating** - Ask what AI feature or model they're testing and what "good" looks like 2. **Help design the eval approach** - Suggest rubrics, test cases, and measurement methods 3. **Guide implementation** - Help them think through edge cases, scoring criteria, and iteration cycles 4. **Connect to product requirements** - Ensure evals align with actual user needs, not just technical metrics ## Core Principles ### Evals are the new PRD Brendan Foody: "If the model is the product, then the eval is the product requirement document." Evals define what success looks like in AI products—they're not optional quality checks, they're core specifications. ### Evals are a core product skill Hamel Husain & Shreya Shankar: "Both the chief product officers of Anthropic and OpenAI shared that evals are becoming the most important new skill for product builders." This isn't just for ML engineers—product people need to master this. ### The workflow matters Building good evals involves error analysis, open
- How to Help
- Core Principles
- Evals are the new PRD
- Evals are a core product skill
- The workflow matters
- Questions to Help Users
- Common Mistakes to Flag
- Deep Dive
- Related Skills
What does the ai-evals skill do?
Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating test cases, designing rubrics, or trying to systematically measure AI output quality.
How do I install it?
Run `npx skills add majiayu000/claude-skill-registry --skill ai-evals-refoundai-lenny-skills-2 --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
