Agent skill · Testing & QA

audio-mixing-mastering

Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social content. Use when planning, mixing, repairing, mastering, QCing, or delivering dialogue, music, ambience, and sound effects, including loudness/true-peak targets, intelligibility, accessibility, stems, stereo/immersive decisions, platform/client specs, and final audio QA.

Calesthio43,316★ · +2,384/wk · 2 repos on radarProfile →
claude-codecodexcopilotcursorships scriptsMIT
Install
npx skills add calesthio/generative-media-skills --skill audio-mixing-mastering --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 4
SKILL.md size: 27 KB
Bundled scripts: yes
Path: skills/production/audio-craft/audio-mixing-mastering/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 112 · +8 this week
Language: Python
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Directs an AI agent on how to plan, mix, repair, master, QC, and deliver dialogue, music, ambience, and sound effects across videos and audio-centric content, including considerations for loudness targets, intelligibility, accessibility, and stems, with attention to platform/client specifications and final QA.

How it works

  • Start with an audio contract for the project, outlining delivery context, foreground hierarchy, required deliverables, target spec, monitoring assumptions, and source risks.
  • Follow evidence categories and use bundled loudness measurement tools when needed to assess integrated loudness, true peaks, and related metrics without altering input.
  • Practice session prep to lock picture, align audio to video, preserve originals, and organize tracks by role.
  • Build balance and hierarchy from foreground (dialogue/VO) outward, prioritizing intelligibility and scene-appropriate dynamics rather than chasing maximum loudness.
  • Apply targeted processing for dialogue (gain, high-pass, de-essing, compression stages, minimal reverb), music (ducking, space for speech, and licensing checks), SFX/ambience (careful use, perspective matching, and accessibility considerations), and general EQ/compression/limiting aimed at preserving codec-friendly headroom.
  • Use limiter guidance to set ceilings based on delivery specs or conservative defaults (-1 dBTP or -2 dBTP) and re-measure after encoding when possible.
  • Prioritize intelligibility for speech-first materials, preserve short-term dynamics for ads/trailers, and maintain consistent dialogue levels for podcasts.

When to use it

Use when audio must survive diverse playback environments and delivery specs, including phone speakers, earbuds, laptops, TVs, cinema, podcasts, client previews, broadcast, and social feeds. Employ before final mastering as a production decision, not a last-minute normalization step.

What it can touch

  • Tools referenced: Python 3.11+, ffmpeg with loudnorm, optional scripts/measure_loudness.py for measurements. The skill describes using these for repeatable local measurement and analysis.
  • It instructs generating a short contract file, session prep steps, and a balance/hierarchy plan, plus detailed processing directions for dialogue, music, SFX, ambience, and EQ/compression/limiting tasks.

Caveats

  • The skill emphasizes adherence to documented loudness standards and platform specs, and warns against blindly mastering to loudest references; it includes stated delivery targets and tolerance guidelines.
  • It notes that results depend on the FFmpeg build and analysis mode when using loudnorm.
  • It requires proper calibration and validation through measurement tools rather than assuming target results from processing alone.
From the SKILL.md

# Audio Mixing and Mastering Direction Use this skill when audio must survive real playback: phone speakers, earbuds, laptops, TV soundbars, cinema-style trailers, podcasts, client review links, broadcast deliveries, and social feeds. Treat the mix as a production decision, not as a last-minute normalization step. The job is to make the audience understand the foreground, feel the intended energy,

More from generative-media-skills
All skills →
About this skill
What does the audio-mixing-mastering skill do?

Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social content. Use when planning, mixing, repairing, mastering, QCing, or delivering dialogue, music, ambience, and sound effects, including loudness/true-peak targets, intelligibility, accessibility, stems, stereo/immersive decisions, platform/client specs, and final audio QA.

How do I install it?

Run `npx skills add calesthio/generative-media-skills --skill audio-mixing-mastering --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From calesthio/generative-media-skills, a repository with 112 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going