hugging-face-evaluation
Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
npx skills add majiayu000/claude-skill-registry --skill antigravity-hugging-face-evaluation --agent claude-code
Same command for any agent — swap --agent for codex, cursor, copilot.
Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.
What it does
Adds structured evaluation results to Hugging Face model cards using multiple methods: extracting existing evaluation tables from README content, importing benchmark scores from Artificial Analysis, and running custom model evaluations with vLLM/accelerate backends. It supports updating model-index metadata for leaderboard integration and aligning with Papers with Code specifications.
How it works
- Inspect and extract evaluation data from README: identifies and parses Markdown tables in a README, selects a target table with --table N, optionally matches model columns via --model-column-index or --model-name-override, and converts the table to model-index YAML with a specified --task-type.
- Import from Artificial Analysis: fetches benchmark scores via API, formats them into model-index entries, preserves source attribution, and can create pull requests to update cards.
- Model-index management: generates properly formatted model-index YAML, merges evaluations into existing model cards, validates against Papers with Code specs, and supports batch processing.
- Run evaluations: supports running evaluations on Hugging Face Jobs with uv, or locally via vLLM/accelerate backends (lighteval, inspect-ai) for custom models; includes guidance for hardware selection and secret/token handling.
When to use it
Use when you need to add or update evaluation results for a model in Hugging Face model cards, drawing from README tables, external benchmark sources, or custom evaluations.
What it can touch
- CLI tools and scripts in scripts/evaluation_manager.py (as described in the CLI workflows): extract-readme, import-aa, inspect-tables, show, validate, get-prs, and run_eval_job equivalents.
- Environment variables and secrets: HF_TOKEN, AA_API_KEY as described in the usage sections.
- Model-index YAML and repository PRs in a Hugging Face model repository.
Caveats
- CRITICAL: Check for Existing PRs Before Creating New Ones: use get-prs prior to any --create-pr or --apply actions to avoid duplicates.
- Dependencies include huggingface_hub, markdown-it-py, python-dotenv, pyyaml, requests, and other HF-related packages; vLLM-related workflows require GPU hardware and uv installation as noted.
- The skill notes risk as unknown and is sourced as community-provided; license is MIT.
# Overview This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data: - Extracting existing evaluation tables from README content - Importing benchmark scores from Artificial Analysis - Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai) ## Integration with HF Ecosystem - **Model Cards**: Updates model-index metadata for leaderboard integration - **Artificial Analysis**: Direct API integration for benchmark imports - **Papers with Code**: Compatible with their model-index specification - **Jobs**: Run evaluations directly on Hugging Face Jobs with `uv` integration - **vLLM**: Efficient GPU inference for custom model evaluation - **lighteval**: HuggingFace's evaluation library with vLLM/accelerate backends - **inspect-ai**: UK AI Safety Institute's evaluation framework # Version 1.3.0 # Dependencies ## Core Dependencies - huggingface_hub>=0.26.0 - markdown-it-py>=3.0.0 - python-dotenv>=1.2.1 - pyyaml>=6.0.3 - requests>=2.32.5 - re (built-in) ## Inference Provider Evaluation - inspect-ai>=0.3.0 - inspect-evals - openai ## vLLM Custom Model Evaluation (GPU req
- Integration with HF Ecosystem
- Core Dependencies
- Inference Provider Evaluation
- vLLM Custom Model Evaluation (GPU required)
- ⚠️ CRITICAL: Check for Existing PRs Before Creating New Ones
- 1. Inspect and Extract Evaluation Tables from README
- 2. Import from Artificial Analysis
- 3. Model-Index Management
- 4. Run Evaluations on HF Jobs (Inference Providers)
- 5. Run Custom Model Evaluations with vLLM (NEW)
- Before running the script
- Running the script
- Features
- Prerequisites
uv run scripts/evaluation_manager.py get-prs --repo-id "username/model-name" uv run scripts/evaluation_manager.py --help uv run scripts/evaluation_manager.py inspect-tables --help uv run scripts/evaluation_manager.py extract-readme --help uv run scripts/train_sft_example.py uv run scripts/evaluation_manager.py inspect-tables --repo-id "username/model" uv run scripts/evaluation_manager.py extract-readme \ or Create .env file echo "AA_API_KEY=your-api-key" >> .env
What does the hugging-face-evaluation skill do?
Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
How do I install it?
Run `npx skills add majiayu000/claude-skill-registry --skill antigravity-hugging-face-evaluation --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.
Where does this skill come from?
From majiayu000/claude-skill-registry, a repository with 534 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.
Is a popular skill a good skill?
Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.
