Agent skill · AI & Agents

phoenix-evals

Build and run evaluators for AI/LLM applications using Phoenix.

GitHub68,948★ · +463/wk · 2 repos on radarProfile →
copilotMIT
Install
npx skills add github/awesome-copilot --skill phoenix-evals --agent copilot

Same command for any agent — swap --agent for claude-code, codex, cursor.

Facts
Files in the skill folder: 35
SKILL.md size: 4 KB
Bundled scripts: none
Version: 1.0.0
Declared author: oss@arize.com
Requires: Requires Phoenix server. Python skills need phoenix and openai packages; TypeScript skills need @arizeai/phoenix-client.
Path: skills/phoenix-evals/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 37,432 · +281 this week
Language: Python

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

From the SKILL.md

# Phoenix Evals Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans. ## Quick Reference | Task | Files | | ---- | ----- | | Setup | [setup-python](references/setup-python.md), [setup-typescript](references/setup-typescript.md) | | Decide what to evaluate | [evaluators-overview](references/evaluators-overview.md) | | Choose a judge model | [fundamentals-model-selection](references/fundamentals-model-selection.md) | | Use pre-built evaluators | [evaluators-pre-built](references/evaluators-pre-built.md) | | Build code evaluator | [evaluators-code-python](references/evaluators-code-python.md), [evaluators-code-typescript](references/evaluators-code-typescript.md) | | Build LLM evaluator | [evaluators-llm-python](references/evaluators-llm-python.md), [evaluators-llm-typescript](references/evaluators-llm-typescript.md), [evaluators-custom-templates](references/evaluators-custom-templates.md) | | Batch evaluate DataFrame | [evaluate-dataframe-python](references/evaluate-dataframe-python.md) | | Run experiment | [experiments-running-python](references/experiments-running-python.md), [experiments-running-typescript](references/experiments-running-ty

What's inside
Steps it walks through
  1. Quick Reference
  2. Workflows
  3. Reference Categories
  4. Key Principles
Ships with 24 files
  • references/axial-coding.md
  • references/common-mistakes-python.md
  • references/error-analysis-multi-turn.md
  • references/error-analysis.md
  • references/evaluate-dataframe-python.md
  • references/evaluators-code-python.md
  • references/evaluators-code-typescript.md
  • references/evaluators-custom-templates.md
  • references/evaluators-llm-python.md
  • references/evaluators-llm-typescript.md
  • references/evaluators-overview.md
  • references/evaluators-pre-built.md
  • references/evaluators-rag.md
  • references/experiments-datasets-python.md
  • references/experiments-datasets-typescript.md
  • references/experiments-overview.md
  • references/experiments-running-python.md
  • references/experiments-running-typescript.md
  • references/experiments-synthetic-python.md
  • references/experiments-synthetic-typescript.md
  • references/fundamentals-anti-patterns.md
  • references/fundamentals-model-selection.md
  • references/fundamentals.md
  • references/observe-sampling-python.md
first 24 of 35
More from awesome-copilot
All skills →
About this skill
What does the phoenix-evals skill do?

Build and run evaluators for AI/LLM applications using Phoenix.

How do I install it?

Run `npx skills add github/awesome-copilot --skill phoenix-evals --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From github/awesome-copilot, a repository with 37,432 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going