# Black Forest Labs 发布 Flux 3：可生成带原生音频的 20 秒视频，并推出机器人动作模型 Flux-mimic

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-07-24 02:03
- AIHOT 分数：68
- AIHOT 链接：https://aihot.virxact.com/items/cmrxua5jb02kcroxpfp2mvtgm
- 原文链接：https://the-decoder.com/flux-3-generates-videos-with-native-audio-up-to-20-seconds-long-a-first-for-black-forest-labs

## AI 摘要

德国 AI 公司 Black Forest Labs 发布多模态基础模型 Flux 3，可同时从图像、视频和音频中学习，首次支持生成带原生音频、最长 20 秒的视频。

## 正文

Black Forest Labs

Key Points

The German AI company Black Forest Labs has released Flux 3, a multimodal foundation model that learns from images, videos, and audio simultaneously.

The model generates videos up to 20 seconds long with native audio and offers features such as text-to-video and the ability to chain individual clips together.

In addition, the company has developed Flux-mimic, a video action model for robotics applications, which is already being tested at Audi.

German AI company Black Forest Labs (BFL) has released Flux 3, a multimodal foundation model that learns from images, video, and audio together. In BFL's early tests, it beat several rivals in video generation.

BFL describes Flux 3 as a step toward "real-world visual intelligence," which it defines as models that can "perceive, predict, and act across physical and digital environments." The company is part of a broader push to build so-called world models.

No single modality captures reality in full, BFL argues. Images show spatial structure, video captures how it changes over time, and audio can reveal links between mechanical events and the sounds they produce. Training on all three together lets them fill gaps for one another, giving the model more information than training on each modality separately.

Flux 3 adds native audio to videos up to 20 seconds long

Flux 3 can now generate videos with native audio for the first time, with clips up to 20 seconds long. It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events.

In early evaluations using 10-second clips at 720p, BFL reports that Flux 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent. The margins narrow against stronger competitors. Flux 3 was preferred over Kling v3 Pro 60 percent of the time, over Happy Horse v1 at 59 percent, over Happy Horse 1.1 at 57 percent, and over both Seedance 2.0 and Gemini Omni Flash at 52 percent each.

BFL says the results are preliminary, and no independent tests are available yet. Matching leading systems such as Seedance, which has already reached Hollywood, and Gemini Omni Flash would put Flux 3 among the top video models.

Early user preference rates for a Flux 3 video model compared with eight rivals. A score of 50 percent indicates a tie. | Image: Black Forest Labs

BFL also expects Flux 3 to improve image generation, especially for complex prompts and accurate text rendering in multiple languages. The company plans to release Flux 3 Image in early access within the next few weeks.

BFL says the model can also predict actions based on its understanding of the world. The company worked with Mimic Robotics to develop Flux-mimic, a video-action model now being tested on production tasks at Audi.

视频 · 前往原文观看

Flux 3 is based on Self-Flow, BFL's approach for teaching one model to generate and understand content at the same time. A multimodal transformer uses dedicated components to convert images, video, and audio into a shared internal representation and then turn it back into outputs.

Flux 3 uses a multimodal transformer with dedicated encoders and decoders for images, video, audio, and actions. Its action component can be extended for new uses. | Image: Black Forest Labs

A component for actions provides the foundation for robotics applications. BFL says this unified learning process delivers better results than the previously standard flow-matching method, both in generation quality and in the model's grasp of the physical world.

BFL plans a phased rollout with open-weight access

BFL is rolling out all capabilities in stages, with early-access phases for feedback and safety testing. Flux 3 Video is already available, and Flux 3 Image is set to follow in the coming weeks. Action prediction will initially be offered through select partners.

BFL also plans to release open-weight access to the multimodal backbone under the name "Flux 3 Dev." Longer term, the company is working on next-generation models that aim to combine perception, action, and language prediction in a single model.

AI News Without the Hype – Curated by Humans

BFL Blog
