Sound and dialogue
Can sound drive the picture rather than decorate it?
Sound drives the picture when it's an event, and not when it's a rhythm: an off-screen noise reliably arrives just before the on-screen reaction, but cuts land no closer to the beat than random.
What we ran
60 takes, three bins of 20, all 21:9, 5 to 10 seconds. Every take carries audio FLUX.3 generated in the same pass as the picture. One bin puts a noise off screen and asks whether anyone reacts. One asks for cutting and movement on the beat. One is nothing but lines of dialogue — no setting, no camera, no lighting.
1A quiet square with pigeons. A distant gunshot. The whole flock erupts at once.I can't hear any of it. So every audio claim here is a number. For all 60 takes I pulled the RMS envelope at 100 Hz, per-frame picture motion, per-frame brightness, and ffmpeg scene detection for cuts, then tested the two tracks against each other with a null baseline drawn from the same clips. I also read frames from 36 takes to check that the numbers meant what I thought.
Two things this pass does not test: nothing was transcribed, so I can't say whether a line was delivered as written, and I can't judge voice quality, accent, or performance. Anything about the picture is read from frames and marked as such.
What happened
The off-screen noise works, and it works on time. Measured: take the moment the picture moves most, then look at the audio level in the half-second before it against the 0.9 seconds before that. Across the off-screen-sound bin the level rises a median +4.13 dB, positive in 16 of 19 measurable takes. Sampling a random moment in the same clips instead gives +0.11 dB (p < 0.001, 2000 draws). The music bin gives +1.16 dB. The dialogue bin gives −0.66 dB — nothing. Sound leads picture only where the prompt made sound an event.
1A woman reads at a kitchen table. Off-screen, a door slams hard. She flinches and looks up toward the sound.Frames confirm it. In n=1 the woman is head-down in her book at −0.5s and looking up and off-frame 0.1s after the transient. In n=9 — "A kettle whistles off-screen. She closes the laptop" — the laptop is open before and shut after. In n=10 the pigeons are settled, then the whole square is motion blur. n=7 (doorbell) puts the dog on the sofa at 1.9s, mid-scramble at 3.1s, and gone at 4.4s, with a 12 dB rise into it.
The beat does nothing. Asking for cuts on the beat produced almost no cutting: 12 of 20 takes in that bin are a single unbroken shot, and the 18 cuts across 162 seconds work out to 6.7 per minute — the same rate as the dialogue bin, which was never asked to cut at all. Where cuts exist they are not on anything. Median distance from a cut to the nearest audio onset is 0.390s; dropping the same number of cuts at random times in the same clips gives 0.273s. The real cuts are further from the audio than chance (permutation p = 0.935). Against a beat grid fitted to each clip's own tempo it's 0.170s versus a 0.157s null. Nothing.
There's no continuous coupling either. Correlating the audio onset envelope against picture motion frame by frame gives a median r of −0.005 across all 60 takes. Pairing each take's audio with a different take's picture — 972 mismatched combinations — gives −0.019. Identical. And the picture is periodic at the music's own detected tempo in 3 of 20 takes.
Dialogue alone builds the scene, and cuts it. Ten of 20 dialogue-only takes contain real cuts, and the frames show actual coverage. n=47 goes from a woman holding a letter to an insert of the envelope on a table to a close-up of her eyes — because the lines name a letter. n=54 cuts to the ticket stub. n=42 cuts to a crack in the wall, and the only wall in the prompt is the line "that's coming from inside the walls". n=48 invented a tidal shoreline out of the word "tide". n=59 invented a graveside, a hearse and a widow out of four lines of small talk. Where my level-based turn detection could resolve the exchange (n=42, n=54, n=59), the cuts land a median 0.070s from a line boundary against a 0.393s null — it cuts on the line change. Five cuts is a small sample; treat that one as a strong hint, not a result.
1"No, I'm looking at it right now."2"...Because it's addressed to you. In your handwriting."3"I never mailed this letter. I burned it."
The pattern across all three bins is that the picture follows the nouns and the events in the prompt, not the audio track. A cut goes wherever an object gets named. Both tracks appear to be rendered around the same planned event, which is why a door slam syncs to a flinch and a 120 bpm grid syncs to nothing.
Two honest failures. In n=12 — "An alarm rings in another room" — the audio sits flat at −44 dB through the whole take. He wakes up; the alarm was never made. And the sparse-sound scenes come back quiet: the off-screen bin averages −42.5 dBFS against −27.2 for the music bin.
1A needle drops on vinyl: crackle first, then the room's color warms with the first chord.Notable takes
n=27| E3-047 · cutting/action on music | "the room's color warms with the first chord" — crackle at −51 dB, the needle thumps at 4.09s, near-silence, then the chord enters at 5.11s and the room's brightness leaves its flat 44.9 at 5.17s and reaches 79.7 by 5.29s. Sixty milliseconds between the sound and the light. The best coupling in the family, and it's a lighting ramp, not a cut.n=40| E3-060 · cutting/action on music | "until one sharp stop — held silence" — the picture obeys perfectly: motion peaks at 6.40 and drops to 0.37 as the dancer freezes mid-pose. The final second is the loudest in the take at −11.8 dB. It can stop the picture. It cannot make silence.n=1| E3-001 · sound causes action | "a door slams hard. She flinches and looks up" — the cleanest single measurement in the set, +17.18 dB into the reaction, and the frames show exactly the flinch.n=47| E3-087 · dialogue only — one-sided phone call | three lines about a letter produce a woman, an envelope insert, a close-up and a reverse. Nothing in the prompt describes a room, a prop or a shot.n=10| E3-010 · sound causes action | "A distant gunshot. The whole flock erupts at once" — a 20 dB jump in one second and the square goes to motion blur on the same frame.n=35| E3-055 · cutting/action on music | "Rain intensifies in crescendo... both stop together" — audio climbs −44.0 to −19.2 dB and motion climbs 0.37 to 3.15, monotonic, together. Then the clip just ends at maximum. It can build. It can't land.
What this means if you're prompting FLUX.3
Write sound as an event with a consequence, in that order: the thing that makes the noise, then the person, then what they do. "A phone rings off-screen. He turns toward it before rising." That structure is what got measured coupling; both tracks get planned around the same beat.
Name the reaction. The model animates the verb you write — flinches, freezes, wheels right, closes the laptop. It does not infer one.
Give the event room. The reaction lands at a median 53% of the way through the take, so on a 5-second shot you're paying for about two seconds of setup before anything happens.
Don't ask for beat-synced cutting. You'll get a montage if you ask for a montage, but the sync is yours to do in the edit. The same applies to "on the downbeat", "one room per line", and "the camera completes one orbit per phrase" — none of them landed.
Do ask for a change in light or colour on a musical event. That's the one rhythmic instruction that worked, and it worked to the frame.
You cannot ask for silence. Every prompt requesting a stop, a hold or a hush came back at full level.
Dialogue is enough on its own. Three lines will build a location, the props they mention, and shot/reverse-shot coverage — and the cut tends to land on the line change. If you want a cutaway, name the object in the dialogue and it will cut to it.
Expect to ride gain. Quiet scenes come back around 15 dB below the music takes.
1A flamenco dancer and guitarist trade accelerating phrases until one sharp stop — held silence.1Rain intensifies in crescendo with a string section; both stop together on the same frame.Every take above is straight out of the model — nothing cut, graded or retouched — and every prompt is printed in full.
See every take on the board
All 60 from this report, plus the rest of the run.