Flux 3

Multimodal AI image, video, and audio generator with world understanding.

标签:
Flux 3 is a multimodal foundation model that learns jointly from images, video, and audio to understand real-world dynamics. It generates diverse videos up to 20 seconds with native audio, offers advanced image synthesis and editing, and extends into action prediction for physical AI applications. The model excels in human facial expressions, sound matching, and multilingual generation, with a staged release plan for APIs and open-weight access.

相关导航