FLUX 3 Video Now Talks and Lipsyncs

FLUX 3 Video is being released with a built-in voice.
Black Forest Labs, the German AI company founded by the creators of Stable Diffusion, launched FLUX 3 on Wednesday. The update is not just a video generation tool. It is a unified multimodal model that generates images, Full HD video clips with native audio, lip-synced dialogue in more than 14 languages, and even predicts robot actions. The video component, called FLUX 3 Video, is what is grabbing attention first.
The voice gap in AI video
Most AI video generation models today produce silent clips. Sora, Runway Gen-3, Meta's Make-a-Video, and Pika all generate visuals alone. Adding audio requires a second pass through a separate sound model, and matching lip movements to a spoken script remains a notoriously hard unsolved problem in computer graphics.
FLUX 3 Video collapses this pipeline into one step. It outputs Full HD (1080p) video at up to 20 seconds with native audio baked in. The dialogue is lip-synced to the speaker in over 14 languages. The model can also render typography directly inside the scene, including text on a sign, a subtitle overlaid on a clip, or a logo in a product demo, all without a post-processing layer.
Black Forest Labs' internal Elo rankings place FLUX 3 Video ahead of Seedance 2.0 and Google's Gemini Omni Flash on overall quality.

More than video
If the audio-video sync were the only addition, FLUX 3 would still be a significant release. But Black Forest Labs went wider. The FLUX 3 platform is what the company calls a "single unified multimodal flow model": one architecture handling image generation, video generation, audio generation, and physical AI tasks simultaneously.
The physical AI component, confirmed in coverage from Bloomberg and Decrypt, performs robot action prediction. A model trained on the same FLUX 3 architecture can take a visual scene and output the sequence of actions a robot arm should execute in that environment. This aligns with a broader industry shift toward "world models" that understand physics and motion, not just pixels and text.
Black Forest Labs has also signaled it will release an open-weight version of FLUX 3 Video, following the pattern of its earlier FLUX.1 and FLUX.2 models which helped popularize open image generation. An open-weight video model at this quality level would be a first.
What FLUX 3 Video is good for now
The practical use cases are the obvious ones first: short-form advertising, social media content, product demonstrations, multilingual marketing materials where a single generated clip replaces a full production shoot. The typography rendering means a brand can generate a video with its logo and on-screen text in a single pass, cutting out an entire editing step.
The lip-sync in 14 languages is the feature most likely to reshape content workflows. A company that previously needed separate shoots for English, Spanish, Mandarin, and Arabic audiences can now generate all four from one prompt. Whether the lip-sync quality holds up under close inspection, which is the genuine test for professional use, is not yet independently verified, but even functional quality at this stage is a leap over the current state of the art.
The grounded take
FLUX 3 Video is genuinely impressive for a single-release announcement. The audio-native design, the lip-sync across 14 languages, the typography rendering, and the unified multimodal architecture together represent a bigger step forward than any single video generation release this year.
The caveats are three. First, all benchmark claims come from Black Forest Labs' own internal Elo rankings; they have not been independently reproduced. Second, BFL's earlier FLUX.2 model was strong but never displaced Sora or Runway in real-world usage, suggesting the gap between demo quality and practical reliability is not closed yet. Third, video generation at Full HD with audio is computationally expensive, and the company has not published pricing or latency figures for general availability.
Still, the trajectory is clear: the line between generated video and produced video is thinning fast. FLUX 3 Video did not just add audio. It made audio central to the design. That is a different philosophy than adding sound as a post-hoc layer, and it may be the one that matters.
Sources
- Black Forest Labs makes FLUX 3 Video generally available (The Decoder)
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio (VentureBeat)
- Black Forest Labs Releases FLUX 3 (MarkTechPost)
- Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video And Robot Hands (Decrypt)
- Black Forest Labs Unveils First Model for Robotics (Bloomberg)