On July 23, 2026, Black Forest Labs introduced FLUX 3, which the lab describes as a multimodal foundation model that jointly learns from images, videos, and audio within a unified architecture. It is a significant departure for a company known for still-image generation: FLUX 1 and FLUX 2 generated images, while FLUX 3 is a single model spanning video, image, and action.
On video, FLUX 3 does text-to-video with native audio up to 20 seconds, image-to-video (both animation and reference-based), video-to-video with character and element consistency, and keyframe-to-video for controlled transitions. Black Forest Labs cites multilingual dialogue support and a style range from candid camcorder footage to animation and cinematics. In preliminary human preference evaluations published with the announcement, FLUX 3 Video was preferred over Runway Gen-4.5 in 77 percent of comparisons and over Luma Ray 3.2 in 93 percent, with other comparisons falling in the 52 to 69 percent range. On the image side the lab claims better handling of complex prompts and text rendering.
The rollout is staged rather than general. FLUX 3 Video is in early access now, with APIs and private weights following; FLUX 3 Image opens early access in the following weeks; FLUX-mimic is a partnership model for robot learning built with mimic robotics; and FLUX 3 Dev, an open-weight multimodal backbone, is planned.
The action-prediction component is the strategic move worth noting. A generative model that predicts robot actions from the same backbone that generates video puts Black Forest Labs into physical AI, a market with different buyers and much longer sales cycles than creative tooling. Whether the unified architecture actually pays off across those very different domains is still unproven at announcement time, since only the video variant is in anyone’s hands.