Promoção de lançamento:50% de descontonos planos anuais por tempo limitado
What Is FLUX 3? Multimodal Model Explained
Jul 28, 2026

What Is FLUX 3? Multimodal Model Explained

FLUX 3 is Black Forest Labs' multimodal model, announced July 23, 2026. What it generates, how it was trained, and what you can access today.

If you searched flux 3 this week and landed on five different pages that all said something different, you are not imagining things. One site called it an image generator. Another called it a video model. A third promised a free download. Most of them were written before Black Forest Labs actually said anything.

Black Forest Labs published the FLUX 3 announcement on July 23, 2026. That post is short, specific, and contradicts a lot of what is currently ranking. This guide is built directly on it.

Where this information comes from: every capability, date, and evaluation number below is traced to a Black Forest Labs primary source — the FLUX 3 announcement, the FLUX 3 x mimic research post, the official model pages, and BFL's own documentation. Where BFL has not published something (pricing for FLUX 3, an open-weights release date, benchmark methodology), this guide says so rather than filling the gap. Last verified: July 28, 2026.

The short answer

FLUX 3 is a multimodal foundation model from Black Forest Labs. It is not an image model with video bolted on. According to BFL, it jointly learns from images, video, and audio inside a single unified architecture, and it can generate images and video-with-audio together — from a text prompt alone, or guided by reference images and video.

As of late July 2026:

  • FLUX 3 Video is in Early Access. You request access; it is not a public self-serve product.
  • FLUX 3 Image early access is coming "in the following weeks," in BFL's own words.
  • FLUX 3 Dev — an open-weight multimodal backbone — is on the roadmap, with no announced date.
  • Action prediction is going to selected research and commercial partners first, starting with mimic robotics.

If a site is currently offering you "FLUX 3 free, no signup," it is not running FLUX 3.

What this article solves

The pain point is simple: flux 3 is a low-trust query right now. The model is real and recent, official access is gated, and the SERP is full of tool pages that describe capabilities nobody has verified. People searching this term want three things — what it is, whether they can use it today, and whether the claims they are reading are real.

What you will not get from the other results: most competing pages either predate the July 23 announcement or paraphrase a press summary. This guide quotes the launch plan verbatim, includes the preliminary preference-rate numbers BFL published (with their own caveats attached), and explains the training economics from the mimic research post — which is where the genuinely interesting detail lives.

How FLUX 3 is different from FLUX.1 and FLUX.2

The FLUX line has a clean through-line, and FLUX 3 breaks it deliberately.

GenerationWhat it doesAnnounced
FLUX.1Text-to-image, plus editing tools (Fill, Canny, Depth, Redux, Kontext)2024
FLUX.2Production image generation and editing, up to 4MP, up to 10 reference imagesNovember 25, 2025
FLUX 3Image + video + audio generated jointly; action predictionJuly 23, 2026

BFL's framing in the mimic post is blunt: FLUX 1 and FLUX 2 generate images, while FLUX 3 "expands into multimodality and generates audio-visual content jointly."

The technical basis is an approach BFL calls Self-Flow, which they describe as their method for aligning multimodal generation and understanding within the same architecture. They scaled compute and data to train across video, images, and audio simultaneously rather than training separate specialists and stitching them together.

The detail most summaries skip

In the FLUX 3 x mimic post, BFL puts numbers on the training mix, and they reframe what the model actually is:

  • Video prediction accounted for over 95% of total training compute. Their reasoning: to make video look right, a model has no choice but to learn contact, motion, weight, and cause and effect.
  • Audio is comparatively cheap — BFL notes it makes up less than 0.5% of the tokens in a 720p video with audio.
  • Adding action prediction cost the model almost nothing permanently. Human ratings on text-to-video and image-to-video initially fell by up to 10%, then recovered fully after roughly 3,500 steps while the model also predicted actions.

That last point is the thesis. If one backbone can drive a video generator and a robot arm, BFL argues, it was never only a content model — it is a model of how the world behaves. FLUX-mimic, built with mimic robotics on the FLUX 3 backbone, has been tested and deployed at Audi.

What FLUX 3 can actually generate

Video (Early Access now)

BFL states FLUX 3 can create videos with audio up to 20 seconds in a single generation. The published capability list:

  • Text-to-video generation
  • Image-to-video, either continuing from a starting frame or using images as visual references
  • Video-to-video from a reference clip, carrying elements such as a character into a new scene
  • Generative video-audio continuation from input video and audio
  • Keyframe-to-video for controlled transitions between defined moments
  • Multilingual dialogue
  • Agentic chaining of clips into longer, multi-shot sequences
  • Strong typography and animated design

All outputs come with native audio generation — the audio is not a separate pass.

Images (early access "in the following weeks")

BFL says FLUX 3 can synthesize and edit images across styles, aspect ratios, and resolutions, and that in mid-training evaluations it already showed significant improvement over earlier FLUX versions on complex prompts and text generation, including high-accuracy text in multiple languages. These are their words, and they attach the same caveat: preliminary, expected to improve before release.

Action

Two routes: native action prediction inside FLUX 3, and using the pretrained video backbone as a dynamics-aware foundation that specialised action models are fine-tuned from with limited task-specific data.


Working on visual concepts while FLUX 3 access is still gated? Flux 3 AI is an independent browser workspace for exactly this gap — prompt-to-image generation, reference-image workflows, editing, and storyboard frames you can build today, no waitlist. Start in the AI image generator and keep your direction ready for when official multimodal access opens up.


The evaluation numbers, with the caveat attached

BFL published preliminary preference rates from early evaluations. Their stated method: 10-second text-to-video clips at 720p with audio. Their stated caveat: the model and the surrounding harness are still in development, results are preliminary, and they expect further improvement during early access.

With that framing, FLUX 3 was preferred over:

Compared modelFLUX 3 preferred in
Luma Ray 3.293% of comparisons
Runway Gen-4.577%
Grok Imagine Videoup to 69%
Kling v3 Pro60%
Happy Horse v159%
Happy Horse 1.157%
Seedance 2.052%
Gemini Omni Flash52%

How to read these: they are first-party, preference-based, and preliminary. A 52% preference rate is close to a coin flip — BFL publishing it alongside the 93% figure is a reasonable sign they did not cherry-pick, but it is still vendor-run evaluation. Treat it as directional, not as a benchmark result.

BFL also names specific strengths their early testing surfaced: capturing human facial expressions, associating sounds with physical events, and multilingual capability. Combined with visual references for character consistency, they say sequences lasting several minutes can be assembled.

How to tell a real FLUX 3 claim from a fake one

A practical filter, since this is the main way people get burned on this keyword:

  1. "Download FLUX 3 weights" — false today. Open-weight access ("FLUX 3 Dev") is listed in the launch plan with no date. What you can download is FLUX.1 and FLUX.2 open weights.
  2. "FLUX 3 API, $X per image" — BFL has published no FLUX 3 pricing. The prices on the official pricing page are for the FLUX.2 family and FLUX Tools.
  3. "Unlimited free FLUX 3" — early access is request-gated. A site with no waitlist is running something else.
  4. "FLUX 3 Pro / Ultra / 4K" — those tier names come from the FLUX.1 and FLUX.2 families, not from any FLUX 3 announcement.
  5. A release date more precise than BFL's own — the launch plan says "over the next few weeks and months," phased, each capability after its own early access period.

FAQ

Is FLUX 3 released? It was announced on July 23, 2026 and FLUX 3 Video is available in Early Access by request. It is not generally available, and the model page still carries a "coming soon" label.

Is FLUX 3 free? BFL has not announced pricing or a free tier. Early access is request-based.

Can I run FLUX 3 locally? Not yet. An open-weight multimodal backbone called FLUX 3 Dev appears in the launch plan, but no weights have been released and no date given. FLUX.1 and FLUX.2 open weights are available today on Hugging Face.

Is FLUX 3 better than FLUX.2? For video and audio, there is no comparison — FLUX.2 does not do them. For images, BFL reports significant mid-training improvements over earlier FLUX versions, but FLUX.2 is the shipping production model and FLUX 3 Image is not open yet.

What does "multimodal" mean here specifically? One model, one architecture, trained jointly on images, video, and audio — plus action prediction — rather than separate models per modality.

Who is Black Forest Labs? The German AI lab behind the FLUX model family, founded in 2024 by researchers who worked on widely used open image models. They raised a $300M Series B and operate an "open core" strategy: open-weight models alongside commercial APIs.

The realistic take

FLUX 3 is the most interesting thing BFL has announced, and it is also the least usable thing they have announced. That combination is normal for a frontier release, and it is exactly why the search results are so noisy right now.

The sensible move for a creator or a team: build your prompt library and visual direction now on tools you can actually access, and treat FLUX 3 as a capability that arrives in stages over the coming months. When FLUX 3 Image opens up, the people with a working prompt discipline will get results on day one; the people who spent the wait reading launch rumours will start from zero.

You can start building that library in the Flux 3 AI workspace today, or see what a credit pack covers if you are running volume.

Sources

All primary, all first-party or official documentation:

  1. FLUX 3 — Real World Models (Black Forest Labs announcement, July 23, 2026) — release status, capability list, 20-second video, launch plan, preference rates
  2. FLUX 3 x mimic: The Next Generation of Video-Action Models (BFL, July 23, 2026) — training compute split, action-prediction recovery, Audi deployment
  3. FLUX 3 model page (Black Forest Labs) — current availability label, early access request
  4. FLUX.2: Frontier Visual Intelligence (BFL, November 25, 2025) — FLUX.2 release date and capabilities
  5. Black Forest Labs announcements index — publication dates for all posts cited
  6. BFL API documentation — current recommended model family
  7. BFL API pricing — confirms published pricing covers FLUX.2 and FLUX Tools only
  8. black-forest-labs/flux on GitHub — open-weight model list and licences
  9. Black Forest Labs on Hugging Face — downloadable weights that actually exist
  10. FLUX open weights licensing — commercial licence tiers

Scope note: this article describes FLUX 3 as documented by Black Forest Labs as of July 28, 2026. Flux 3 AI is an independent creator workspace and is not affiliated with, endorsed by, or reselling Black Forest Labs models. For official availability, pricing, and API terms, always check bfl.ai directly.

Comece a criar com o FLUX 3 Gerador de Imagens com IA

Experimente grátis o FLUX 3 AI: descreva uma imagem, carregue uma referência e gere resultados prontos a apresentar.