Agent skills

Testing & QA skills

Read straight from the source repositories, not from submitted listings. Every skill shows what it does, what is inside, where it came from — and whether attention around its source is actually growing.

Toolclaude-code 29,140codex 4,755cursor 3,111copilot 976windsurf 55cline 34
CategoryWorkflow & Productivity 4,979AI & Agents 3,037Data & Analytics 2,345Code Review & Quality 1,376Backend & API 1,244Security 1,194Design & Presentation 1,154Documentation 965Content & Marketing 916Testing & QA 777DevOps & Cloud 576Databases 550Frontend 469Business & Finance 328Media & Video 257Other 9,833
1,723 found
529576 · page 12 / 36
agentic-evalPatterns and techniques for evaluating and improving AI agent outputs. Use this skill when implementing self-critique and…majiayu000agents-mdUse when creating or updating AGENTS.md files to guide AI coding agents. Covers file structure, placement, content guidelines…majiayu000agf-writing-github-issueUse whenever a user, product-lead, or qa-engineer wants to create a GitHub issue in the project repo — including phrases like…pcliangxai-evalsCreate an AI Evals Pack (eval PRD, test set, rubric, judge plan, results + iteration loop). Use for LLM evaluation, benchmarks…majiayu000ai-evalsHelp users create and run AI evaluations. Use when someone is building evals for LLM products, measuring model quality, creating…majiayu000alibaba-image-modelsPlan, prompt, call, edit, iterate, and productionize Alibaba Cloud Model Studio image generation with current Wan 2.7 Image…calesthioAllure Test ReportingAllure test reporting framework for comprehensive test result visualizationa5c-aiwritesamazon-pollyUse Amazon Polly for production text-to-speech work: selecting Standard, Neural, Long-form, or Generative engines and compatible…calesthioamazon-rekognitionUse this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR…calesthioanalysis-diagnosePerform systematic root cause investigation through 4-phase process. Use when encountering bugs, test failures, or unexpected…majiayu000analyze-and-planMUST USE when investigating a bug, CI failure, test failure, regression, incident, broken behavior, root cause, RCA, or debug-why…Optim-Agentanalyze-sessionAnalyze a Figma MCP test session transcript. Reads raw session data (JSON or HTML) and produces a structured analysis document…majiayu000android-playstore-api-validationCreate and run validation script to test Play Store API connectionmajiayu000anomaly-investigationUse when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number…gaasherapi-builderGenerate complete FastAPI backend scaffolds from OpenAPI 3.x specifications. Automatically creates SQLAlchemy models, Pydantic…majiayu000api-design-principlesMaster REST and GraphQL API design principles to build intuitive, scalable, and maintainable APIs that delight developers and…majiayu000api-devScaffold, test, document, and debug REST and GraphQL APIs. Use when the user needs to create API endpoints, write integration…majiayu000api-developmentAPI development workflow and best practices for career_ios_backend FastAPI project. Automatically invoked when user wants to…majiayu000writesapi-investigatornote.com APIの調査を支援します。mitmproxyとPlaywrightを使用してHTTPトラフィックをキャプチャ・分析し、API動作を解明します。majiayu000writesapi-test-suite-builderUse when the user asks to generate API tests, create integration test suites, test REST endpoints, or build contract tests.alirezarezvaniapi-testerTest and document API endpoints - validate responses, check status, generate examplesgooseworks-aiarchitectural-analysis"Performs deep architectural analysis of a specified module, directory, or feature area by examining structural coupling, data…testdoubleasset-continuity-managementProvider-independent asset continuity and version management for generated-media production. Use when an agent must track…calesthioassumption-trackerExplicitly track, test, and validate assumptions - prevent blind spotsmajiayu000atlas-reconDocumentation reconnaissance for takeover — find all docs, assess accuracy, freshness, coverage, and discoverability, and…jeremylongshorewritesauditing-safe-harbor-checklistVerify OpenMed de-identified output against all 18 HIPAA Safe Harbor identifier categories and report residual re-identification…maziyarpanahiauto-claudeAutonomous multi-agent coding with git worktree isolation, QA validation, and memory. Use for complex features requiring…majiayu000autoresearch-agentAutonomous experiment loop that optimizes any file by a measurable metric. Inspired by Karpathy's autoresearch. The agent edits a…majiayu000autoresearchAutonomous experiment loop: edit code, commit, run benchmark, extract metrics, keep improvements or revert, repeat forever. Use…majiayu000avatar-spokesperson-productionProvider-independent production workflow for AI avatar spokesperson videos, including presenter briefs, consent and likeness…calesthioaxe-accessibilityDeep integration with axe-core for automated accessibility testing. Execute accessibility scans, interpret WCAG violations…a5c-aiwritesazure-speechUse Microsoft Azure Speech in Foundry Tools for media-production speech workflows: speech-to-text, fast and batch transcription…calesthiobats-test-scaffolderGenerate BATS test structure and fixtures for shell script testing with setup/teardown, assertions, and mocking.a5c-aiwritesbats-testing-patternsMaster Bash Automated Testing System (Bats) for comprehensive shell script testing. Use when writing tests for shell scripts…sickn33bats-testing-patternsMaster Bash Automated Testing System (Bats) for comprehensive shell script testing. Use when writing tests for shell scripts…wshobsonbenchmark-fetcherFetch benchmark performance data from 6 leaderboard websites using Playwright MCP and update model manifests with the latest…majiayu000better-codexBehavioral guardrails for Codex coding work based on common user complaints. Use when Codex is asked to implement, modify, debug…noobnoocbook-idea-validatorStress-test book concepts against existing research before committing to architecture. Use when the user has a Book Concept…majiayu000brainstormingUse when defining new features, product behavior, UI/component design, architecture choices, contract changes, or ambiguous…GanyuanRanbrooks-healthCombined codebase health dashboard that scores a project across all four quality dimensions — PR quality, architecture, tech…hyhmrrightbrooks-sweepFull-sweep mode: runs a unified analysis across all quality dimensions — code decay, architecture, tech debt, and test quality —…hyhmrrightbrooks-sweepFull-sweep mode: runs a unified analysis across all quality dimensions — code decay, architecture, tech debt, and test quality —…sickn33brooks-testTest quality review drawing on twelve classic engineering books — with primary focus on xUnit Test Patterns, The Art of Unit…hyhmrrightbrooks-testTest quality review drawing on twelve classic engineering books — with primary focus on xUnit Test Patterns, The Art of Unit…sickn33bugbashSystematically explore and test any software project (CLI, API, Backend, Library, etc.) to find bugs, usability issues, and edge…avbuild-testRun the project's build / typecheck / lint / test commands and emit the build.passing + tests.passing signals devloop convergence…nexu-iobusiness-brainstormWhen you want to pressure-test a potential new business, product, or side project against the serial-founder filter. Not…coreyhaines31byteplus-seed-speech-ttsProduction guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or…calesthio
← Prev12 / 36Next →
How the catalog works
What is an agent skill?

A folder with a SKILL.md inside — instructions, and often scripts and assets, that an AI agent loads when the task matches. Claude Code, Codex, Cursor and Copilot all read the same format, so one skill usually works across them.

Where does this catalog come from?

We read 660 source repositories straight from their file trees rather than from submitted listings — what you see is what is actually published. 98 repositories were rejected because they advertise skills but contain none: link lists, not folders.

Why is there no install counter?

Because install counts live in the registry that serves `npx skills add`, and that is not ours — publishing a number we cannot verify would be worse than showing none. Instead we show where a skill comes from and whether attention around its source is actually growing, measured from our own weekly snapshots.

Do you deduplicate?

Yes, and it matters more than expected. Aggregator repositories republish the same skill in several places — one source carried 6,317 SKILL.md files for 2,001 actual skills. We collapse by folder name and keep the canonical copy, so the catalog counts things, not copies.

Keep going