The Decoder Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, video, and audio and can generate video with native sound for the first time. BFL's own tests put it just ahead of market leader Seedance 2.0, though independent results aren't yet available. The company ultimately wants to build a world model and is already testing Flux 3 on robotics tasks.
The article Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs appeared first on The Decoder.
Get AI news in your inbox
Daily digest of what matters in AI.
Key Terms Explained #
Decoder The part of a neural network that generates output from an internal representation.
Foundation Model A large AI model trained on broad data that can be adapted for many different tasks.
Multimodal AI models that can understand and generate multiple types of data — text, images, audio, video.
World Model An AI system's internal representation of how the world works — understanding physics, cause and effect, and spatial relationships.