Agent skills

Media & Video skills

Read straight from the source repositories, not from submitted listings. Every skill shows what it does, what is inside, where it came from — and whether attention around its source is actually growing.

Toolclaude-code 29,139codex 4,760cursor 3,079copilot 978windsurf 55cline 34
CategoryWorkflow & Productivity 4,987AI & Agents 3,033Data & Analytics 2,347Code Review & Quality 1,377Backend & API 1,243Security 1,197Design & Presentation 1,129Documentation 967Content & Marketing 917Testing & QA 777DevOps & Cloud 576Databases 550Frontend 469Business & Finance 328Media & Video 258Other 9,845
520 found
241288 · page 6 / 11
openai-realtime-voiceBuild production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket…calesthioopenai-whisperLocal Whisper: audio transcription, multi-language, word timestamps, speaker diarizationmajiayu000openrouter-transcribeTranscribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).majiayu000openrouter-transcribeTranscribe audio files via OpenRouter using audio-capable models (Gemini, GPT-4o-audio, etc).majiayu000performance-budget-labUse for website performance planning, Lighthouse/web-vitals checks, bundle/media/font budgets, Core Web Vitals triage…TheGoat395podcast-productionProduce audio-first podcast episodes with generative tools — design the show and episode format (interview, narrative, news…calesthioprocedural-audioProcedural sound skill for synthesis and dynamic sound design.a5c-aiwritesproduct-ad-video-scriptTurn a product into ad angles, hooks, scenes, and CTA.0xslinequantum-musicQuantum computer music composition and performance using quantum circuits, ZX-calculus notation, and quantum instrumentsmajiayu000qwen-asrTranscribe audio files using Qwen ASR. Use when the user sends voice messages and wants them converted to text.majiayu000resemble-detectDeepfake detection and media safety — detect AI-generated audio, images, video, and text, trace synthesis sources, apply…davepoonresemble-detectDeepfake detection and media safety — detect AI-generated audio, images, video, and text, trace synthesis sources, apply…Prat011reverse-traceIdentify the source of an image or video frame — TV show episode, movie scene, geographic location, or original publication. This…majiayu000transformersThis skill should be used when working with pre-trained transformer models for natural language processing, computer vision…majiayu000senior-computer-visionWorld-class computer vision skill for image/video processing, object detection, segmentation, and visual AI systems. Expertise in…majiayu000spatial-designDesigning for visionOS — spatial layout and ergonomics (60pt eye targets, field-of-view placement, dynamic scale), eyes-and-hands…rshankraswritesspeech-pathology-aiExpert speech-language pathologist specializing in AI-powered speech therapy, phoneme analysis, articulation visualization, voice…majiayu000writesspeech-pathology-aiExpert speech-language pathologist specializing in AI-powered speech therapy, phoneme analysis, articulation visualization, voice…majiayu000writesspeech-to-textTranscribe audio to text with Whisper models via inference.sh CLI. Models: Fast Whisper Large V3, Whisper V3 Large. Capabilities…majiayu000tiktok-ad-analysisAUTOMATIC video market analysis with hybrid transcription (TikTok captions → Whisper fallback). Auto-triggers when videos exist…majiayu000transformersHugging Face Transformers for loading Hub models, running pipeline inference, text generation, and Trainer fine-tuning on NLP…majiayu000twinmind-incident-runbookIncident response for TwinMind failures: transcription not starting, audio not captured, sync failures, and calendar disconnect.…jeremylongshorewritesunreal-metasoundsUnreal Engine MetaSounds skill for procedural audio, real-time synthesis, and advanced audio graphs.a5c-aiwritesvideo-analysis-workflowGuides video analysis for CMJ and drop jump. Use when processing athlete videos, debugging pose detection, troubleshooting…majiayu000video-clipperRepurposes long-form video (podcasts, interviews, talks) into short-form vertical clips for Instagram Reels, TikTok, and YouTube…gooseworks-aiwritesvideo-digestTriage new videos, generate watch recommendationsmajiayu000video-metadata-inspectorUse when asked to inspect video file metadata, get video duration, resolution, codec information, frame rate, or bitrate.majiayu000video-polishTakes an existing screen recording or demo video and adds professional zoom/pan effects synchronized to the narration. Uses…gooseworks-aiwritesvideo-shortformShort-form video generation skill — 3-10 second clips for product reveals, motion teasers, ambient loops. Defaults to Seedance 2…nexu-ioVideo Transcript AnalyzerAnalyze customer interview transcripts (SRT or plain text) to generate thematic breakdowns with summary, quotes, topics…majiayu000writesvideo-upscaleUpscale and restore video in ComfyUI — both the quick local path (per-frame ESRGAN like 4x_foolhardy_Remacri via…majiayu000Voice SynthesisGenerate voice audio using Google Cloud TTS with enhancementmajiayu000wavecap-hallucinationConfigure WaveCap hallucination detection and prevention. Use when Whisper outputs gibberish, repeated phrases, or phantom text…majiayu000wedding-immortalistTransform thousands of wedding photos and hours of footage into an immersive 3D Gaussian Splatting experience with theatre mode…majiayu000writeswhisperOpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language…Orchestra-ResearchwwiseWwise integration skill for sound banks, RTPC, and interactive music.a5c-aiwritesyaml-jazzYAML is sheet music. The LLM is the jazz musician. Comments are soul.majiayu000youtube-comment-analysisUse when user requests YouTube comments. Run standalone for comment analysis or sequential with youtube-to-markdown for…majiayu000writesyoutubeAnalyze a YouTube video and create a vault note from its transcript. Use this skill whenever a YouTube URL is pasted — even with…majiayu000writesyoutube-processorProcess YouTube videos into summarized Obsidian notes. Use when given a YouTube URL to summarize, extract insights, or turn…majiayu000writesyoutube-summarizerAutomatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms.…majiayu00004-script-videoViet script video ngan cho TikTok, Reels, YouTube Shorts — 2 ban A/B, co hook, CTA, huong dan quay chi tietminhnv080706-ugc-egc-brief-globalDetailed brief for UGC (customers), EGC (employees), KOC (paid creators) for global markets — shooting guide, do/don't, usage…minhnv0807abridge-common-errorsDiagnose and fix common Abridge clinical AI integration errors. Use when encountering EHR connectivity failures, note generation…jeremylongshorewritesabridge-performance-tuningOptimize Abridge clinical AI integration performance for high-volume deployments. Use when reducing note generation latency…jeremylongshorewritesacmmm-related-workUse when building or auditing the related-work section of an ACM MM (ACM Multimedia) paper — covering the multimedia literature…brycewang-stanfordadvanced-dubbing-studioSubmit audio or video for multilingual dubbing, poll status, and download dubbed audio. Use when the user asks for dubbing…opensquillaadvmat-results-framingUse when an Advanced Materials result is sound but its central advance and significance are not yet sharp, or when a…brycewang-stanford
← Prev6 / 11Next →
How the catalog works
What is an agent skill?

A folder with a SKILL.md inside — instructions, and often scripts and assets, that an AI agent loads when the task matches. Claude Code, Codex, Cursor and Copilot all read the same format, so one skill usually works across them.

Where does this catalog come from?

We read 660 source repositories straight from their file trees rather than from submitted listings — what you see is what is actually published. 98 repositories were rejected because they advertise skills but contain none: link lists, not folders.

Why is there no install counter?

Because install counts live in the registry that serves `npx skills add`, and that is not ours — publishing a number we cannot verify would be worse than showing none. Instead we show where a skill comes from and whether attention around its source is actually growing, measured from our own weekly snapshots.

Do you deduplicate?

Yes, and it matters more than expected. Aggregator repositories republish the same skill in several places — one source carried 6,317 SKILL.md files for 2,001 actual skills. We collapse by folder name and keep the canonical copy, so the catalog counts things, not copies.

Keep going