# Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

> Source: <https://www.machinebrief.com/news/flux-3-generates-videos-with-native-audio-up-to-20-seconds-l-mfzs>
> Published: 2026-07-23 18:03:01+00:00

# Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs

[The Decoder](https://the-decoder.com)

Black Forest Labs has released Flux 3, a [multimodal](/glossary/multimodal) [foundation model](/glossary/foundation-model) that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a [world model](/glossary/world-model) and is already testing Flux 3 on [robotics](/category/robotics) tasks.

The article [Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs](https://the-decoder.com/flux-3-generates-videos-with-native-audio-up-to-20-seconds-long-a-first-for-black-forest-labs/) appeared first on [The Decoder](https://the-decoder.com).

Get AI news in your inbox

Daily digest of what matters in AI.

## Key Terms Explained

[Decoder](/glossary/decoder)

The part of a neural network that generates output from an internal representation.

[Foundation Model](/glossary/foundation-model)

A large AI model trained on broad data that can be adapted for many different tasks.

[Multimodal](/glossary/multimodal)

AI models that can understand and generate multiple types of data — text, images, audio, video.

[World Model](/glossary/world-model)

An AI system's internal representation of how the world works — understanding physics, cause and effect, and spatial relationships.
