Agent skill · Testing & QA

ai-evals

Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating test cases, designing rubrics, or trying to systematically measure AI output quality.

majiayu000github.com/majiayu000GitHub ↗
claude-codeMIT
Install
npx skills add majiayu000/claude-skill-registry --skill ai-evals-refoundai-lenny-skills-2 --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 3 KB
Bundled scripts: none
Path: skills/ai-llm/ai-evals-refoundai-lenny-skills-2/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 534
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# AI Evals Help the user create systematic evaluations for AI products using insights from AI practitioners. ## How to Help When the user asks for help with AI evals: 1. **Understand what they're evaluating** - Ask what AI feature or model they're testing and what "good" looks like 2. **Help design the eval approach** - Suggest rubrics, test cases, and measurement methods 3. **Guide implementation** - Help them think through edge cases, scoring criteria, and iteration cycles 4. **Connect to product requirements** - Ensure evals align with actual user needs, not just technical metrics ## Core Principles ### Evals are the new PRD Brendan Foody: "If the model is the product, then the eval is the product requirement document." Evals define what success looks like in AI products—they're not optional quality checks, they're core specifications. ### Evals are a core product skill Hamel Husain & Shreya Shankar: "Both the chief product officers of Anthropic and OpenAI shared that evals are becoming the most important new skill for product builders." This isn't just for ML engineers—product people need to master this. ### The workflow matters Building good evals involves error analysis, open

What's inside
Steps it walks through
  1. How to Help
  2. Core Principles
  3. Evals are the new PRD
  4. Evals are a core product skill
  5. The workflow matters
  6. Questions to Help Users
  7. Common Mistakes to Flag
  8. Deep Dive
  9. Related Skills
Ships with 1 file
  • metadata.json
More from claude-skill-registry
All skills →
About this skill
What does the ai-evals skill do?

Help users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating test cases, designing rubrics, or trying to systematically measure AI output quality.

How do I install it?

Run `npx skills add majiayu000/claude-skill-registry --skill ai-evals-refoundai-lenny-skills-2 --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going