Agent skill · Backend & API

hugging-face-evaluation

Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.

majiayu000github.com/majiayu000GitHub ↗
claude-codeMIT
Install
npx skills add majiayu000/claude-skill-registry --skill antigravity-hugging-face-evaluation --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 22 KB
Bundled scripts: none
Path: skills/ai-ml/antigravity-hugging-face-evaluation/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 534
Language: HTML

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Adds structured evaluation results to Hugging Face model cards using multiple methods: extracting existing evaluation tables from README content, importing benchmark scores from Artificial Analysis, and running custom model evaluations with vLLM/accelerate backends. It supports updating model-index metadata for leaderboard integration and aligning with Papers with Code specifications.

How it works

  • Inspect and extract evaluation data from README: identifies and parses Markdown tables in a README, selects a target table with --table N, optionally matches model columns via --model-column-index or --model-name-override, and converts the table to model-index YAML with a specified --task-type.
  • Import from Artificial Analysis: fetches benchmark scores via API, formats them into model-index entries, preserves source attribution, and can create pull requests to update cards.
  • Model-index management: generates properly formatted model-index YAML, merges evaluations into existing model cards, validates against Papers with Code specs, and supports batch processing.
  • Run evaluations: supports running evaluations on Hugging Face Jobs with uv, or locally via vLLM/accelerate backends (lighteval, inspect-ai) for custom models; includes guidance for hardware selection and secret/token handling.

When to use it

Use when you need to add or update evaluation results for a model in Hugging Face model cards, drawing from README tables, external benchmark sources, or custom evaluations.

What it can touch

  • CLI tools and scripts in scripts/evaluation_manager.py (as described in the CLI workflows): extract-readme, import-aa, inspect-tables, show, validate, get-prs, and run_eval_job equivalents.
  • Environment variables and secrets: HF_TOKEN, AA_API_KEY as described in the usage sections.
  • Model-index YAML and repository PRs in a Hugging Face model repository.

Caveats

  • CRITICAL: Check for Existing PRs Before Creating New Ones: use get-prs prior to any --create-pr or --apply actions to avoid duplicates.
  • Dependencies include huggingface_hub, markdown-it-py, python-dotenv, pyyaml, requests, and other HF-related packages; vLLM-related workflows require GPU hardware and uv installation as noted.
  • The skill notes risk as unknown and is sourced as community-provided; license is MIT.
From the SKILL.md

# Overview This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data: - Extracting existing evaluation tables from README content - Importing benchmark scores from Artificial Analysis - Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai) ## Integration with HF Ecosystem - **Model Cards**: Updates model-index metadata for leaderboard integration - **Artificial Analysis**: Direct API integration for benchmark imports - **Papers with Code**: Compatible with their model-index specification - **Jobs**: Run evaluations directly on Hugging Face Jobs with `uv` integration - **vLLM**: Efficient GPU inference for custom model evaluation - **lighteval**: HuggingFace's evaluation library with vLLM/accelerate backends - **inspect-ai**: UK AI Safety Institute's evaluation framework # Version 1.3.0 # Dependencies ## Core Dependencies - huggingface_hub>=0.26.0 - markdown-it-py>=3.0.0 - python-dotenv>=1.2.1 - pyyaml>=6.0.3 - requests>=2.32.5 - re (built-in) ## Inference Provider Evaluation - inspect-ai>=0.3.0 - inspect-evals - openai ## vLLM Custom Model Evaluation (GPU req

What's inside
Steps it walks through
  1. Integration with HF Ecosystem
  2. Core Dependencies
  3. Inference Provider Evaluation
  4. vLLM Custom Model Evaluation (GPU required)
  5. ⚠️ CRITICAL: Check for Existing PRs Before Creating New Ones
  6. 1. Inspect and Extract Evaluation Tables from README
  7. 2. Import from Artificial Analysis
  8. 3. Model-Index Management
  9. 4. Run Evaluations on HF Jobs (Inference Providers)
  10. 5. Run Custom Model Evaluations with vLLM (NEW)
  11. Before running the script
  12. Running the script
  13. Features
  14. Prerequisites
Ships with 1 file
  • metadata.json
Commands it runs
uv run scripts/evaluation_manager.py get-prs --repo-id "username/model-name"
uv run scripts/evaluation_manager.py --help
uv run scripts/evaluation_manager.py inspect-tables --help
uv run scripts/evaluation_manager.py extract-readme --help
uv run scripts/train_sft_example.py
uv run scripts/evaluation_manager.py inspect-tables --repo-id "username/model"
uv run scripts/evaluation_manager.py extract-readme \
or
Create .env file
echo "AA_API_KEY=your-api-key" >> .env
More from claude-skill-registry
All skills →
About this skill
What does the hugging-face-evaluation skill do?

Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.

How do I install it?

Run `npx skills add majiayu000/claude-skill-registry --skill antigravity-hugging-face-evaluation --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going