hugging-face-evaluation-manager
Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
npx skills add majiayu000/claude-skill-registry --skill hf-model-evaluation --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
What it does
The skill provides methods to add structured evaluation results to Hugging Face model cards. It supports three main approaches: extracting existing evaluation tables from README content, importing benchmark scores from Artificial Analysis, and running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai). It also updates model-index metadata to enable leaderboard integration and is designed to work with Papers with Code specifications.
How it works
- Inspect and extract evaluation tables from README: uses markdown-it-py to parse tables, select a specific table with --table N, optionally override model headers with --model-name-override, identify model columns, and generate YAML in model-index format with a specified --task-type.
- Import from Artificial Analysis: fetches benchmark scores via AA API, formats results into model-index entries, preserves metadata attribution, and can create pull requests to publish updates.
- Model-index management: generates properly formatted YAML entries, merges evaluations into existing model cards without overwriting, validates against Papers with Code specs, and supports batch processing for multiple models.
- Run evaluations on HF Jobs: supports running evaluations via uv run and hf jobs integration, with options for CPU or GPU hardware, and secure handling of tokens and secrets.
- Run custom vLLM evaluations: execute local or HF Jobs-based evaluations using vLLM or accelerate backends, with prerequisites and specific workflow steps (including lighteval and inspect-ai options).
When to use it
Use when you need to add or update evaluation results in Hugging Face model cards, either by importing external scores, extracting existing README tables, or running bespoke evaluations on local hardware or HF infrastructure. Trigger steps include inspecting tables, extracting a specific table, applying changes, importing AA data (with optional PR creation), or running a vLLM/inspect-ai evaluation job.
What it can touch
- Commands and scripts in scripts/ (e.g., evaluation_manager.py, inspect-tables, extract-readme, import-aa, show, validate, get-prs).
- HF model-index metadata and model cards via integration points described in the workflow.
- External APIs: Artificial Analysis API for importing scores.
- Evaluation backends: vLLM, lighteval, accelerate, inspect-ai, and HF Jobs infrastructure.
Caveats
- Requires environment setup with HF_TOKEN (and AA_API_KEY for Artificial Analysis) and possibly a local GPU for vLLM workflows.
- The workflow cautions to check for existing open PRs before creating new ones to avoid duplicates.
- The skill relies on specific dependencies and PEP 723 header behavior for automatic dependency installation when using uv run.
# Overview This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data: - Extracting existing evaluation tables from README content - Importing benchmark scores from Artificial Analysis - Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai) ## Integration with HF Ecosystem - **Model Cards**: Updates model-index metadata for leaderboard integration - **Artificial Analysis**: Direct API integration for benchmark imports - **Papers with Code**: Compatible with their model-index specification - **Jobs**: Run evaluations directly on Hugging Face Jobs with `uv` integration - **vLLM**: Efficient GPU inference for custom model evaluation - **lighteval**: HuggingFace's evaluation library with vLLM/accelerate backends - **inspect-ai**: UK AI Safety Institute's evaluation framework # Version 1.3.0 # Dependencies ## Core Dependencies - huggingface_hub>=0.26.0 - markdown-it-py>=3.0.0 - python-dotenv>=1.2.1 - pyyaml>=6.0.3 - requests>=2.32.5 - re (built-in) ## Inference Provider Evaluation - inspect-ai>=0.3.0 - inspect-evals - openai ## vLLM Custom Model Evaluation (GPU req
- Integration with HF Ecosystem
- Core Dependencies
- Inference Provider Evaluation
- vLLM Custom Model Evaluation (GPU required)
- ⚠️ CRITICAL: Check for Existing PRs Before Creating New Ones
- 1. Inspect and Extract Evaluation Tables from README
- 2. Import from Artificial Analysis
- 3. Model-Index Management
- 4. Run Evaluations on HF Jobs (Inference Providers)
- 5. Run Custom Model Evaluations with vLLM (NEW)
- Before running the script
- Running the script
- Features
- Prerequisites
uv run scripts/evaluation_manager.py get-prs --repo-id "username/model-name" uv run scripts/evaluation_manager.py --help uv run scripts/evaluation_manager.py inspect-tables --help uv run scripts/evaluation_manager.py extract-readme --help uv run scripts/train_sft_example.py uv run scripts/evaluation_manager.py inspect-tables --repo-id "username/model" uv run scripts/evaluation_manager.py extract-readme \ or Create .env file echo "AA_API_KEY=your-api-key" >> .env
What does the hugging-face-evaluation-manager skill do?
Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
How do I install it?
Run `npx skills add majiayu000/claude-skill-registry --skill hf-model-evaluation --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
