Black Forest Labs released FLUX 3 Action, an open 7B world-action model for robots

Black Forest Labs, best known for its FLUX image and video models, released FLUX 3 Action on 2026-09-23. It is a 7-billion-parameter “world action model” derived from the company’s multimodal FLUX 3 backbone: it takes camera images, the robot’s joint state and a text instruction, and jointly predicts the next chunk of motor commands together with the video frames it expects those actions to produce. The weights are published on Hugging Face as black-forest-labs/flux-3-action-base under the FLUX Kommunity License v1.0, with fine-tuned checkpoints for the SO-101 arm and the DROID dataset integrated into Hugging Face’s LeRobot library and code on GitHub.

On NVIDIA’s RoboLab-120 simulation benchmark, the launch post reports 42.92 percent success; the model page breaks this down as 38.3 percent for the base model and 42.2 percent for a guidance-distilled variant, against 36.8 percent for NVIDIA’s Cosmos 3 Nano policy and 28.0 percent for Physical Intelligence’s pi0.5. BFL reports inference speedups of 1.34x to 2.28x over pi0.5 and 1.52x to 3.95x over Cosmos 3 Nano in FP8 across GPU types, with 41 ms latency on a B200 in BF16. The only real-robot result at launch is from a third party, Positronic Robotics, which logged 28 of 30 successful tasks (93.3 percent) on a Franka arm.

The model reflects a broader bet that video generation models are a natural foundation for robot control: a network already trained to predict how scenes evolve can be taught to predict the actions that cause those changes. BFL trained on NVIDIA GB200 systems and says FP8 and BF16 variants run on hardware down to an RTX 5090.

Why it matters: an image-generation company entering robotics with an open-weight policy that leads a public simulation leaderboard at under half the parameters of NVIDIA’s entry shows how directly generative video work now transfers to embodied AI. What it does not show: RoboLab-120 is a simulation benchmark, the headline number depends on which variant is quoted (42.92 versus 42.2 percent across BFL’s own pages), and one 30-trial third-party test on one arm is far from evidence of robust real-world performance. The license is a custom community license, not an OSI open-source license.