An image lab shipping a video model first is unusual. Black Forest Labs built its reputation on FLUX.1 and FLUX.2 — still images, editing, typography. Then on July 23, 2026 they announced FLUX 3 and opened early access to the video capability, with image early access explicitly scheduled for later.
That ordering is not a marketing choice. It follows from how the model was trained, and once you see the training economics, the whole roadmap makes sense.
Sourcing note: every capability, number, and comparison below comes from Black Forest Labs' own FLUX 3 announcement and the accompanying FLUX 3 x mimic research post. Preference rates are first-party and preliminary — BFL says so themselves, and this article repeats the caveat wherever the numbers appear. No third-party benchmark is presented as fact. Last verified July 28, 2026.
What this article solves
The pain point: "FLUX 3 video" searches return capability lists with no context. You can find the feature bullets anywhere. What is missing is whether the numbers mean anything, what the limits actually are, and whether the model is available to you.
The differentiator: this piece pairs each published claim with its caveat, explains the >95%-of-compute training detail that other write-ups skip entirely, and reads the preference-rate table honestly — including the entries where FLUX 3 barely wins.
What FLUX 3 Video does
BFL's stated headline: FLUX 3 can create highly diverse videos with audio up to 20 seconds in length in a single generation. Every output comes with native audio generation — not a separate soundtrack pass.
The documented capability set:
- Text-to-video generation
- Image-to-video, either continuing from a starting frame ("animation") or using images as visual references
- Video-to-video from a reference clip, carrying central elements of a source video — for instance the same character — into a new scene or context
- Generative video-audio continuation from input video and audio
- Keyframe-to-video for controlled transitions between defined moments
- Multilingual dialogue
- Agentic chaining of individual clips into longer, multi-shot sequences
- Broad style range, from candid camcorder footage to animation and cinematics
- Strong typography generation and animated designs
BFL specifically calls out three early strengths from their own testing: capturing human facial expressions, associating sounds with physical events, and multilingual capability. They add that combining these with visual references for character consistency allows sequences lasting several minutes.
Why video came first: the training economics
This is the part that explains everything else, and it comes from the FLUX 3 x mimic post.
- Video prediction accounted for over 95% of total training compute. BFL's reasoning: to generate realistic video, a model has no choice but to learn contact, motion, weight, and cause and effect. Get any of them wrong and it looks wrong. Learning to render the world accurately means learning how the world behaves.
- Audio is comparatively trivial. BFL describes it as low-dimensional and far less detailed — less than 0.5% of the tokens in a 720p video with audio. Once a model has learned video understanding, it picks up the causal link between video and audio: speech synchronised to lip movement, effects synchronised to the physical events causing them.
- Actions have the same shape as audio — a low-dimensional representation tightly coupled to visual observation. When BFL added action prediction to the training curriculum, human ratings on text-to-video and image-to-video initially fell by up to 10%, then fully recovered after roughly 3,500 steps while the model was also predicting actions.
The conclusion BFL draws: video generation and action prediction do not need separate foundations. The same backbone carries both. That is why FLUX-mimic — built with mimic robotics on the FLUX 3 backbone — is running robots tested and deployed at Audi, and why BFL frames physical AI as an extension of the roadmap rather than a pivot.
For a creator, the practical takeaway is narrower but real: a model that spent 95% of its compute learning physics should be unusually good at motion that obeys weight and contact — falling objects, cloth, liquid, impact. That is where you should point it first.
The comparison numbers, read honestly
BFL published preference rates from early evaluations. Their stated setup: 10-second text-to-video clips at 720p with audio. Their stated caveat, quoted in substance: the model and the harness around it are still in development, the results are preliminary, and further improvements are expected during early access.
| Compared against | FLUX 3 preferred in | Reading |
|---|---|---|
| Luma Ray 3.2 | 93% | Decisive |
| Runway Gen-4.5 | 77% | Clear |
| Grok Imagine Video | up to 69% | Clear, note "up to" |
| Kling v3 Pro | 60% | Meaningful |
| Happy Horse v1 | 59% | Meaningful |
| Happy Horse 1.1 | 57% | Modest |
| Seedance 2.0 | 52% | Statistical coin flip |
| Gemini Omni Flash | 52% | Statistical coin flip |
How to interpret this table:
- It is vendor-run. First-party preference testing with an unpublished rater methodology is directional evidence, not a benchmark.
- The bottom rows matter more than the top. A 52% preference over Seedance 2.0 and Gemini Omni Flash means those models are effectively at parity in this evaluation. If you already have a working pipeline on either, FLUX 3 is not an obvious upgrade for raw output quality alone.
- "Up to 69%" is a range ceiling, not an average — worth noting since the other figures are stated flat.
- The publishing choice is a good sign. A lab optimising for headlines would print the 93% and stop. Including two coin flips suggests the set was not cherry-picked.
Where FLUX 3 plausibly separates itself is not on the preference score at all — it is on native joint audio, 20-second single generations, and video-to-video character carry. Those are structural capabilities, not quality deltas.
Storyboards do not need to wait for early access. Flux 3 AI is an independent browser workspace for the planning half of video work — generate keyframes, lock character references, build shot-by-shot visual direction. Start with the image generator and walk into your first FLUX 3 session with the frames already made.
The limits you should plan around
- Availability. FLUX 3 Video is Early Access, by request through BFL. Not self-serve, not resold.
- Length. 20 seconds per generation. Longer runs come from chaining, which introduces continuity risk that visual references are meant to mitigate.
- Evaluation resolution. BFL's published evaluations used 720p clips. That is what the numbers describe; it is not a stated maximum output resolution.
- Preliminary everything. BFL applies the "still in development" caveat to the model, the harness, and the results.
- No pricing. Nothing published. Any cost estimate you read elsewhere is invented.
- Image capability is not open. If you need still images from FLUX 3, that early access phase had not opened as of July 28, 2026.
FAQ
How long can FLUX 3 videos be? Up to 20 seconds in a single generation, with longer multi-shot sequences via agentic chaining.
Does FLUX 3 generate audio? Yes — natively, on all video outputs, including multilingual dialogue and sound tied to physical events.
Is FLUX 3 better than Kling or Runway? In BFL's own preliminary evaluations, preferred over Kling v3 Pro in 60% and Runway Gen-4.5 in 77% of comparisons. First-party, preliminary, 10-second 720p clips. Treat as directional.
Can I do image-to-video with FLUX 3? Yes — both animation from a starting frame and reference-guided generation are documented modes.
Can FLUX 3 keep a character consistent across shots? That is the stated purpose of video-to-video and visual references; BFL cites character consistency across chained scenes.
How do I get access? Request early access on the FLUX 3 model page at bfl.ai. There is no other route.
Bottom line
FLUX 3 Video is a genuinely different proposition from the current field: one model producing picture and sound together, 20 seconds at a time, from a backbone that spent almost all of its training budget learning how the physical world moves.
The published quality numbers are more modest than the headline figure suggests, and the model is gated. Both facts argue for the same plan — get your visual direction, character references, and keyframes built now, so that access, when it lands, turns into output instead of experiments. Build them in the Flux 3 AI workspace, or review the credit plans if you are working at production volume.
Sources
- FLUX 3 — Real World Models (BFL, July 23, 2026) — capability list, 20-second limit, native audio, preference rates and caveats, evaluation setup
- FLUX 3 x mimic: The Next Generation of Video-Action Models (BFL, July 23, 2026) — 95% compute on video, audio token share, action-prediction recovery curve, Audi deployment
- FLUX 3 model page — availability status and access request
- BFL announcements index — publication dates
- FLUX.2: Frontier Visual Intelligence (BFL, November 25, 2025) — prior generation scope, for the image-only contrast
- BFL API documentation — models currently exposed via API
- BFL API pricing — confirms no FLUX 3 pricing published
- FLUX.2 model page — multi-reference behaviour underpinning character consistency
- black-forest-labs/flux on GitHub — open-weight releases to date
- FLUX open weights licensing — weight-level access tiers
Scope note: Flux 3 AI is an independent creator workspace, not affiliated with Black Forest Labs and not a source of FLUX 3 access. Verify availability and terms at bfl.ai.


