Thirteen questions. 971 takes.
We ran FLUX.3 until it broke, then wrote down where. Thirteen experiments, each one moving a single variable and holding everything else still — 1,014 generations, 184 minutes of video, every prompt published.
The short version
FLUX.3 is better at reasoning than at arithmetic, and better at building a world than at cutting one.
It tracks cause and effect further than we expected. A seven-step Rube Goldberg machine holds all the way to the plate, an off-screen noise reliably arrives just before the on-screen reaction, and asking for twenty seconds doesn't stretch a five-second idea — it writes a longer script, with a turn the short take never had. Name a period and you get the frame shape, the title cards and the cut rhythm of that decade, not a filter over a modern shot.
It cannot count. Ten coins knocked off a stack land as thirteen; a scorebug reads perfectly and then increments wrong — 132 of 145 requested strings rendered exactly while almost anything that had to tick or total went astray. It also won't take direction on pace: you can contradict a format's setting, its subject and its lighting, but not the speed of its cutting, and cuts land no closer to a musical beat than chance. Pinned keyframes get the same treatment — first and last are targets, the ones in between are suggestions, and ten pins land less often than four.
Two things matter less than the internet says. How you format a prompt is nearly irrelevant — eight of ten writing styles came back interchangeable as long as the prompt still said who and where. And length has a ceiling: the model reads every word out to 4,700 characters, but the film stops improving at about 1,400. One thing matters more: the language you write in. Write the same scene in Cantonese and it recasts the actors and rebuilds the street — though only where the scene has a street.
And when you reroll, know what moves. The face, the lens and often the ending change; the location, the light and the look do not. A longer prompt steadies the picture without steadying the plot — which is the whole argument for drafting several takes cheaply before committing to one.
How it was run
Every take is a single generation, straight out of the model — nothing cut, graded or retouched. Each experiment fixed a scene and moved one variable across it, so a difference on screen is the variable and not the roll. Where a claim is a rate (“27 of 32”), it was counted by watching the takes, and the denominator is printed with it. Where a result could be luck, it was tested against a shuffled baseline and the report says so.
Nothing here is a benchmark score. It is thirteen production questions, answered with enough takes that the answer isn't one lucky render — and the full board is public, so you can disagree with the reading and check the footage yourself.
The 13 reports
Each one is a single question, its answer, and the takes behind it.
Formats
Can it hold a format — and what happens when you contradict it?
Say "CCTV" or "1948 film noir" and FLUX.3 hands you that format's frame shape, runtime and grain whether you asked for them or not — then it will do almost anything you demand inside that frame, except change the pace of the cutting.
On-screen text
Can it write — title cards, dates, scorebugs, lower thirds?
Yes, it can write — 132 of 145 requested strings rendered exactly, and 69 of 74 dates matched the record — but it cannot count, so anything that has to tick, total or increment goes wrong while every word around it stays perfect.
Prompt style
Which way of writing a prompt wins?
Format barely matters — what you name matters, and eight of the ten styles are interchangeable as long as the prompt still says who and where.
Multilingual scenes
How many languages can share one scene?
Five, and it doesn't fall apart — coverage actually improves as the count climbs, because the model answers more speakers by cutting to more of them.
Prompt length
Does a longer prompt make a better film?
FLUX.3 reads every word you give it, right out to 4,700 characters — but the film stops getting better at about 1,400, and past 2,400 the extra writing buys texture nobody sees.
Finished films
Can it make something that reads as an actual film?
Yes — 27 of the 32 multi-shot briefs came back with the four shots they asked for, the subtitles came back verbatim, and the scripted silences landed on the second; what it will not do is break physics or fake a frame rate.
Eras
How far back does its sense of period go, and what happens at the seams?
Its sense of period is real and structural back to 1901 — it changes the frame shape, the title cards and the cut rhythm, not just the grade — but it thins out after 1996, and a future date buys you nothing the prompt didn't already name.
Sound and dialogue
Can sound drive the picture rather than decorate it?
Sound drives the picture when it's an event, and not when it's a rhythm: an off-screen noise reliably arrives just before the on-screen reaction, but cuts land no closer to the beat than random.
Duration
Does twenty seconds buy a story?
Yes — but not the way you'd think. FLUX.3 doesn't stretch a five-second idea to fill twenty seconds; it writes a longer script. A 5s take isn't a compressed version of the 20s take, it's the first five seconds of one.
Physics and logic
Does it keep track of cause, count and reflection?
It tracks cause better than anyone expects and counts worse — a seven-step machine holds all the way to the plate, while ten coins knocked off a stack land as thirteen.
Variance
How much does the same prompt drift between takes?
A reroll swaps the face, the lens and — often — the ending. It will not move the location, the light or the look, and writing a longer prompt steadies the picture without steadying the plot.
Languages
Does the prompt's language change the film?
Change the language and FLUX.3 recasts the actors, rebuilds the street and rewrites the signs — but only where the scene has a street; offices, buses and karaoke rooms come back identical in all seven.
Keyframes
Can it hit ten pinned frames in order, on time?
FLUX.3 treats your first and last pin as targets and everything in between as a suggestion — and ten pins land less often than four.
Open the board
Every take from every report, with its prompt and settings attached.