Agent skill · Testing & QA

video-to-audio-foley

Use for video-conditioned audio and automated Foley production: adding synchronized sound effects, ambience, impacts, footsteps, cloth, prop, and environmental audio to silent or under-sounded video using video-to-audio models, text-to-sound tools, manual spot lists, editing, rights checks, and QA.

Calesthio43,316★ · +2,384/wk · 2 repos on radarProfile →
claude-codecodexcopilotcursorMIT
Install
npx skills add calesthio/generative-media-skills --skill video-to-audio-foley --agent claude-code

Same command for any agent — swap --agent for codex, cursor, copilot.

Facts
Files in the skill folder: 2
SKILL.md size: 23 KB
Bundled scripts: none
Path: skills/production/audio-craft/video-to-audio-foley/SKILL.md
Open the folder on GitHub →
Where it comes from
Stars: 112 · +8 this week
Language: Python
Read our review of the source →

Weekly change comes from our own snapshots, not the repository page — it measures attention, not adoption.

Review
written from the skill's own SKILL.md · Aug 5, 2026

What it does

Turn a silent or under-sounded video clip into a credible sound-effects bed for various video formats, treating V2A as a production accelerator to decide automation scope, prompts, segmentation, sync preservation, and QA.

How it works

  • Defines boundaries for when to use automated V2A, and when not to rely on it, guiding the user to consider frame-exact sync, asset delivery format, and rights.
  • Provides a comprehensive workflow: define deliverables, prepare and segment video, build a spot list, generate a first V2A bed per segment, review against picture, layer precise sounds, build the final mix, and perform QA/packaging.
  • Emphasizes prompt construction with layered inputs (scene bed, hero actions, materials/mechanics, perspective/density, exclusions) and gives example phrasing to steer the model.
  • Recommends tool selection based on needs (hosted V2A with prompt control, prompt-inference from video, or local/open-source options) and includes decision rules to choose the simplest adequate path.
  • Outlines sync and editing guidance for hero impacts, footsteps, cloth, props, vehicles, UI, and animation, stressing timing discipline and manual adjustments when necessary.
  • Enumerates rights, consent, and safety requirements prior to upload or publication, including license checks, provenance, and avoiding real-voice cloning or trademarked sounds.
  • Provides a QA checklist focusing on semantic, temporal, acoustic, and technical fit to ensure alignment with visuals and output quality.

When to use it

Use automated V2A when the clip is short enough for the model, aims for a believable first-pass Foley bed, and the user can tolerate iteration and editorial repair. Avoid automated V2A when frame-exact sync is mission-critical, when clean stems are required, or when the source content includes sensitive or restricted material. If the user requests different deliverables (mixed track, stems, assets), follow the guidance to choose a suitable workflow and provide separation accordingly.

What it can touch

The skill discusses multiple hosted and local tools for V2A, including text prompts, negative prompts, reference audio options, duration, seeds, inference steps, and model IDs. It instructs inspecting the current tool registry, provider docs, licenses, input schema, and safety terms before choosing a tool suitable to the production need.

Caveats

  • The guidance emphasizes that V2A is a production accelerator and may require manual spot work and QA to ensure timing and auditory quality.
  • It cautions against relying on automated V2A for frame-precise assets or where licensed material or privacy constraints exist, and it requires rights clearance and provenance.
  • It notes that model outputs may vary with duration limits, prompt interpretation, and tool terms, requiring live rechecking before paid calls or deliverables.
From the SKILL.md

# Video-to-audio Foley Use this skill to turn a silent, generated, animated, game, ad, social, or edit-ready video clip into a credible sound-effects bed. Treat video-to-audio (V2A) as a production accelerator, not as a finished mix. The job is to decide what should be automated, what must be hand-spotted, how to prompt and segment the model, how to preserve sync, and how to verify the result agai

More from generative-media-skills
All skills →
About this skill
What does the video-to-audio-foley skill do?

Use for video-conditioned audio and automated Foley production: adding synchronized sound effects, ambience, impacts, footsteps, cloth, prop, and environmental audio to silent or under-sounded video using video-to-audio models, text-to-sound tools, manual spot lists, editing, rights checks, and QA.

How do I install it?

Run `npx skills add calesthio/generative-media-skills --skill video-to-audio-foley --agent claude-code` — it drops the skill into your project so the agent can pick it up. Swap the --agent value for codex, cursor or copilot if you use one of those.

Where does this skill come from?

From calesthio/generative-media-skills, a repository with 112 stars. We read it straight from the repository tree rather than a submitted listing, so what you see here is what is actually published.

Is a popular skill a good skill?

Not necessarily. Stars measure attention, not adoption — a repository can trend for a week and be abandoned. That is why we show the weekly change from our own snapshots next to the total, instead of a single flattering number.

Keep going