Audio
Generate score, dialogue, ambience, and effects that are ready to edit.
Start by deciding what role the sound plays in the cut. A score prompt describes musical development; a dialogue prompt provides exact words and a performance; an effects prompt describes a physical source in a space. Choosing that role first makes the model list easier to navigate.
Audio generation can produce score, songs, foley, ambience, dialogue, voiceover, or a complete scene bed. Music models build cues and songs. Sound-effects models create discrete sounds and textures. Speech models perform written dialogue. All-in-one models combine several layers into one result.
Music
- MiniMax Music: generate a complete track or instrumental score from one prompt.
- Lyria 3: generate 30-second jingles, vocal sketches, and quick music ideas.
- Lyria 3 Pro: generate polished songs up to about three minutes, including cinematic instrumentals and vocals.
- ElevenLabs Music: control the duration and build a song section by section.
- Stable Audio 2.5: generate short music, loops, ambience, foley, and sound-design textures.
Write music prompts like cue sheets. Name the tempo, meter, instrumentation, performance, and edit points: “0–8s sparse low strings; 8–20s brushed drums enter; 20–32s muted brass swell; no vocals.” For score, describe its job under picture: “Hold tension under dialogue; no melody until the reveal.”
Sound Effects
- ElevenLabs SFX v2: generate short foley, ambience, impacts, and atmospheric sounds.
- Beatoven SFX: generate longer stereo effects for creatures, vehicles, impacts, sci-fi, and abstract textures.
- CassetteAI SFX: generate fast placeholders and draft variations.
- Stable Audio 2.5: build longer ambience, texture beds, and music-adjacent sound design.
Describe a sound as a source in a scene. Name its distance, space, surface, and timing: “Single leather bootstep on wet concrete, close mic, parking-garage reverb, no music or voice.” Put multiple events in the order they should happen.
Dialogue
- ElevenLabs TTS v3: generate expressive voiceover, ADR, and character dialogue from exact wording.
- MiniMax Speech-02 HD: generate clean multilingual narration.
- Seed Audio 1.0: generate quick English or Chinese speech alongside soundtrack work.
Put the exact spoken words first, then add a short performance direction such as “calm documentary narrator, close mic, slower pace” or “teen character whispering, nervous, one breath before the last line.” Generate each speaker separately when you need editorial control.
All-in-one and native audio
- Seed Audio 1.0: compose dialogue, sound effects, and music into one scene bed.
- Native-audio video models: generate sound that follows the action in the shot.
For Seed Audio, describe the whole soundtrack in sequence: dialogue first, a door latch at two seconds, then a low synth pad underneath. Use native audio when the sound needs to follow visible action in the generated shot.
Each audio result lands on the canvas. Play it before changing the prompt, rename useful takes, then drag the select to an audio track in the timeline. Keep score, dialogue, and ambience on separate tracks so you can trim, balance, and bridge them across cuts. Generate extra handles when timing matters. Use Stable Audio 2.5 Inpaint when you need to repair one section without replacing the whole track.
Match the prompt to the duration. Ask for one clear event in short sound effects. Describe the full arc of a music cue. Keep dialogue to one sentence or beat when timing matters.
For a first pass, build a rough score bed and cut picture against it. Add effects for visible actions, then record or generate dialogue. Replace placeholders as the edit settles, and keep alternate takes nearby until the mix direction is clear.