Skip to content

FLUX 3 vs Seedance 2.5, Gemini Omni and Kling 3: where does the newcomer really stand?

Black Forest Labs releases a model that generates image, sound and video in a single pass. But in a market where six models compete across four different axes, it does not come first on any of them.

Advertisement
The essentials in 30 seconds ⚡
On 23 July 2026, German lab Black Forest Labs launched FLUX 3, a model trained simultaneously on image, video and audio, capable of producing twenty seconds of sound video from a single prompt. Its singularity is architectural, not quantitative: on duration, price or measured quality, it is not first anywhere. In a market where six models share four different axes, the real question is no longer which is best, but which fits your use case.

Every new video model is presented as a revolution. To find out what is really going on, you have to place it in a landscape that has shifted enormously in six months. Here is what FLUX 3 offers, and above all where it stands against the ones you may already be using.

Who is Black Forest Labs

The name is less well known than Google or ByteDance, but it carries weight. Black Forest Labs is a German research company based in Freiburg, founded by the team that was at the heart of Stable Diffusion. In other words, the people who largely popularised open source image generation. Director Martin Scorsese is among its advisers. Until now, the FLUX family was reserved for still images, with a solid reputation for prompt adherence.

Its real singularity: one brain for three senses

Here is what truly sets FLUX 3 apart, and it is not a number. Most tools that produce sound video actually assemble several separate models behind a common interface: one model generates the image, another animates it, a third makes the sound, and a software layer stitches the pieces together.

FLUX 3 was trained jointly on image, video and audio, in a single architecture, via an approach called Self-Flow. It learns spatial structure, motion, sound and physical interactions together rather than as distinct tasks. It is a polished example of what we described in our article on multimodal models.

Practical consequence: the sound is not added afterwards, it is native. The model matches audio to visible events, handles multilingual dialogue, and produces more coherent facial expressions because it has learned the link between a face speaking and the sound that comes out of it.

The real landscape, in July 2026

Let us move on to what most articles on the subject miss: the comparison. Here is where the main models stand on the axes that really matter.

Model Duration in one pass Indicative price Its strong point
Seedance 2.5 (ByteDance) 30 s, no cut Varies by platform Longest native duration, 4K, up to 50 references
Gemini Omni Flash (Google) ~10 s (capped preview) ~$0.10/second Top of all four Artificial Analysis rankings
FLUX 3 (Black Forest Labs) 20 s Not disclosed Unified image, sound, video, action architecture
Kling 3.0 (Kuaishou) 15 s ~$0.11 to $0.14/s (Turbo) Native 4K at 60 fps, multi-shot storyboards, per-character lip-sync
Veo 3.1 (Google) ~8 s $0.05 to $0.40/s depending on tier Best all-rounder, dialogue at 48 kHz
Runway Gen-4.5 5 to 10 s Depends on subscription Finest control: motion brushes, pro ecosystem

This table tells a simple story: no model wins on every axis. Seedance 2.5 dominates duration with its thirty seconds without a cut. Gemini Omni Flash dominates measured quality and price. Kling 3.0 dominates technical finish in 4K at 60 frames per second. Runway dominates creative control. And FLUX 3 arrives in the middle of the pack on duration, with no public price.

The blind ranking, mid-July 2026 📊
On the Artificial Analysis text-to-video-with-audio ranking, where humans vote without knowing which model produced what, the order was as follows: Gemini Omni Flash at the top with an Elo of around 1,240, then Seedance 2.0 in 720p at 1,225, Alibaba's HappyHorse-1.1 at 1,149, and Kling 3.0 Pro at 1,110. Runway Gen-4.5, number one at its launch in late 2025 with 1,247, has dropped out of the top 10. The pace of renewal in this market is brutal.

What Black Forest Labs' numbers are worth

The lab announces flattering comparisons from its own preference tests: FLUX 3 would have been preferred to Luma Ray 3.2 in 93% of cases and to Runway Gen-4.5 in 77%.

But the most honest figures are the ones the lab publishes against its most serious competitors. Against Seedance 2.0 and Gemini Omni Flash, the result hovers around 52%, which means that on a ten-second video generated from text, the models are statistically indistinguishable. Beating the trailing models while matching the leaders is an honourable performance, not a domination. And as we remind readers in our article on benchmarks, these comparisons come from the vendor and have not been audited.

The absence of a price, a real problem

That is the red cell in the table, and it is disqualifying in the short term. Without a published pricing grid, it is impossible to calculate a cost per project or to compare seriously. When Gemini Omni Flash shows $0.10 per second and Kling Turbo around $0.12, a model without a price cannot enter a purchasing decision.

Two other gaps add to this. Deployment is partial: only FLUX 3 Video is in early access, the image part is to follow. And the weights are not open, which is notable for a lab whose reputation was built on open source. An open-weights version is promised later in the year, without a date. As we wrote about Kimi K3, a promise with a timeline is not a file on your hard drive.

The hidden bet: robotics

If FLUX 3 is not aiming for the top of the rankings, what is it aiming for? The answer lies in the most unexpected part of the announcement. The same architecture that generates videos powers FLUX-mimic, a video-action model developed with the company mimic robotics, currently being tested by Audi on a production line.

The logic holds: to generate a credible video, a model must understand how objects move, fall, deform and sound. That is exactly what a robot needs to manipulate the physical world. Black Forest Labs is betting that physical AI and content creation rest on the same foundation. Its competitors optimise for prettier videos; it is building a world model of which video is just one output among others. That is a long-term bet, impossible to assess today.

So, which one to choose?

The right question is not which is the best model, but which axis matters to you. If you need a long shot without a cut, Seedance 2.5. If you want the best value for money with conversational back-and-forth, Gemini Omni Flash. If you need broadcast 4K and multi-character dialogue, Kling 3.0. If fine creative control comes first, Runway. If you want the safest all-rounder, Veo 3.1.

And FLUX 3? For now, it is a model to watch rather than adopt. Its unified architecture is intellectually the most ambitious of the lot, and its robotic ambition could make it relevant in a way the others are not even aiming for. But as long as it has no public price, no open weights, and no full deployment, it remains a well-built promise in a market where six competitors are already shipping.

What strikes you, looking at this table, is the speed at which the hierarchy reshuffles. Runway was number one eight months ago and is no longer in the top 10. Today's leader may be overtaken by the autumn. For a creator, the practical conclusion is perhaps this: do not build your work around a model, build it around what you want to tell, and switch tools when a better one arrives. They arrive fast.

Advertisement