{"slug": "flux-3-action-a-7b-open-weight-world-action-model-for-robots", "title": "Flux 3 Action: A 7B open-weight world action model for robots", "summary": "Black Forest Labs released FLUX 3 Action, a 7B open-weight world action model that predicts future frames and actions jointly and placed first on the RoboLab-120 benchmark at 42.92% success, ahead of the 16B Cosmos3-Nano-Policy at 36.8%. The model was fine-tuned on the DROID dataset and the SO-101 arm, with both checkpoints integrated into LeRobot, and weights are published under the FLUX Kommunity License v1.0. FLUX 3 Action is a 7B diffusion transformer that encodes frames with a frozen video VAE and instructions with a frozen Qwen3-VL-4B, returning 32 actions and optionally 32 decoded frames per call.", "body_md": "[Image-Text-to-Text •  4B • Updated   •  3.43M  •  481](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct)  \n\n# \n\t\tFLUX 3 Action: a world action model you can fine-tune\n\t\n\n [Community Article](https://huggingface.co/blog/community)\n\nFine-tuned on [DROID](https://droid-dataset.github.io/), it places first on the RoboLab benchmark at 42.92% success, against 36.8% for the 16B Cosmos 3 Nano policy. We also fine-tuned it on the SO-101 arm, and both the DROID and SO-101 checkpoints are integrated into [LeRobot](https://huggingface.co/docs/lerobot/index).\n\n| Model | Open source | Type | SR% | Parameters | \n|---|---|---|---|---|\n| **FLUX 3 Action** | Yes | WAM | **42.92%** | 7B | \n| OASIS WAM | No | VLM + WAM | 39.0% | — | \n| Cosmos3-Nano-Policy | Yes | WAM | 36.8% | 16B | \n| Phoenix | No | TAMP+FM | 34.4% | — | \n| BiMind v0.1 | No | VLA | 33.3% | — | \n| π0.5 | Yes | VLA | 28.0% | [3.3B](https://www.physicalintelligence.company/download/pi05.pdf) | \n| DreamZero | Yes | WAM | 25.7% | [14B](https://arxiv.org/abs/2602.15922) | \n| Cosmos3-Edge-Policy | Yes | WAM | 22.9% | [4B](https://huggingface.co/nvidia/Cosmos3-Edge) | \n| π0-FAST | Yes | VLA | 15.5% | [3B](https://huggingface.co/lerobot/pi0fast-base) | \n| GR00T N1.6 | Yes | VLA | 7.2% | [3B](https://huggingface.co/nvidia/GR00T-N1.6-3B) | \n| π0 | Yes | VLA | 5.0% | [3.3B](https://www.physicalintelligence.company/download/pi0.pdf) | \n| paligemma-binning | Yes | VLA | 3.4% | [3B](https://huggingface.co/google/paligemma-3b-pt-224) | \n\n*Table 1: RoboLab-120 overall success rates. Types and weight availability follow the [RoboLab leaderboard](https://research.nvidia.com/labs/srl/projects/robolab/leaderboard.html).*\n\nWeights are released under the [FLUX Kommunity License v1.0](https://huggingface.co/black-forest-labs/flux-3-action-base/blob/main/LICENSE.md) in the [FLUX 3 Action collection](https://huggingface.co/collections/black-forest-labs/flux-3-action-6ab25aef555dd30ab86567f8), and the code is at [github.com/black-forest-labs/flux-action](https://github.com/black-forest-labs/flux-action).\n\n## \n\t\tWhat is FLUX 3 Action\n\t\n\nFLUX 3 Action is an open weights world action model (WAM); it predicts future frames and actions together. It's a 7B parameter diffusion transformer, trained on action data from several robot embodiments; a frozen video VAE encodes the frames and a frozen [Qwen3-VL-4B](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) encodes the instruction.\n\nActions are a second token stream in the same sequence: one token per future frame, with an input projection and output head per embodiment. Video and action tokens share a single noise level per sample and are denoised jointly.\n\nPer call, the model takes:\n\n- one or more camera frames, composited onto a canvas the VAE encodes (three DROID cameras at 544×736, two SO-101 cameras side by side, or one game frame at 512×512);\n- a state vector in the action's own space (joint positions for a robot, the last executed action for a game);\n- a caption, through the text encoder.\n\nIt returns 32 actions and, optionally, 32 decoded frames. At control time, you skip the decode, execute the first few actions, observe again, and replan.\n\nThe model was trained on NVIDIA GB200 systems, with custom kernels written in NVIDIA's CuTe DSL. We collaborated with NVIDIA on PEFT fine-tuning recipes and edge deployment on NVIDIA Jetson.\n\n## \n\t\tRunning it on a real arm through LeRobot\n\t\n\nThe LeRobot integration makes FLUX 3 Action a policy class you train and run with LeRobot's own tools. We built it with NVIDIA and Hugging Face, and it ships with parameter-efficient fine-tuning recipes for adapting the model to a task.\n\n*The SO-101 doing the task it was trained on: put the blue box into the container. Top camera, 4× speed with the pauses between plans cut. This is the policy driving the arm.*\n\nThe policy in these clips was adapted on about 200 teleoperated episodes covering a handful of related pick-and-place tasks. It was never trained to pick the objects in the following videos.\n\n| *put the white box from the green cup to the container* | *put the tissues into the container* | \n\nThe model also recovers from its own mistakes, as in this clip with the screwdriver.\n\nWe also tested it on containers it had never seen during training.\n\n| *container manipulation and containers not seen in training* | *Notebook lying on top of the container, leaving only 40% of the container visible to FLUX* | \n\nFinally, FLUX 3 Action still works when you change the camera position between training and rollout.\n\n| *camera position 1* | *camera position 2* | *camera position 3* | \n\n## \n\t\tTrain FLUX 3 Action into a policy for a specific task\n\t\n\nTo test fine-tuning beyond robotics, we trained task-specific policies for several games. A robot arm is useful, but slow for understanding how a policy behaves: evaluating a single checkpoint on hardware can take a whole afternoon, and collecting demonstrations requires a human operator.\n\nGames make evaluation faster. An agent, running a FLUX 3 Action checkpoint to play a game, can be scored in minutes against a scripted bot that sees the same frame. We also trained a FLUX 3 Action policy to control a drone.\n\nThe full guide, with the config, the dataset module and the pitfalls, is in the docs: [docs.bfl.ai/flux_3/flux3_action_overview](https://docs.bfl.ai/flux_3/flux3_action_overview)\n\n### \n\t\tVideo games\n\t\n\nWe built two small games to train it on. GRUNT is a Quake-style shooter, 256×256, four actions: move, strafe, turn, fire. VECTOR is an Out Run-style road racer with traffic, oil and gravel, three actions: steer, throttle, nitro.\n\n| *FLUX 3 Action playing GRUNT* | *FLUX 3 Action driving in VECTOR* | \n\nThe training data comes from a scripted bot in each game. The bot plays, we record what it saw and what it pressed: 800 episodes per game. The bot has one rule. It decides from the current frame and nothing else, so anything it knows, the model can see too. Take a hit in GRUNT and the screen flashes red; that flash is how the model learns to turn toward whoever shot it. Some episodes start from random positions or get shoved mid-run, so the model also sees states a clean run never reaches.\n\nThe GRUNT bot fires only when its hit cone covers an enemy sprite. The VECTOR bot drives 296 km/h on the straights and brakes to between 155 and 215 km/h into bends.\n\n**GRUNT**\n\n| *step 250: moves and fires at nothing* | *step 500: turns towards enemies but aim isn't there* | *step 3000: kills everything it meets, never dies* | \n\n| Policy | Kills | Deaths | Shots on target | \n|---|---|---|---|\n| Scripted bot | 13 | 0 | 96% | \n| **FLUX 3 Action** | **15** | 0 | 86% | \n| Random | 8 | 2 | 13% | \n\n**VECTOR**\n\n| *250 steps: wrecks in the first seconds* | *1000 steps: on the road, passes traffic, still wrecks* | *3000 steps: keeps pace with the bot, 0 off-road* | \n\n| Model | km in 60 s | Wrecks | Passes | \n|---|---|---|---|\n| Scripted bot | 3.4 | 0 | 16 | \n| FLUX 3 Action, joint model | 3.3 | 1 | 12 | \n| Random driver | 1.0 | 3.6 | 0 | \n\nThe model in both tables is one set of weights, trained on GRUNT and VECTOR together and told which game it is playing by the caption. Sixty seconds, one seed, the bot on the same seed.\n\n### \n\t\tAn indoor drone\n\t\n\nWe also trained FLUX 3 Action on controlling an indoor drone. The drone sees a 256×256 onboard camera and outputs [forward, lateral, up, yaw]. Its episodes were recorded in [NVIDIA Isaac Sim](https://developer.nvidia.com/isaac/sim): a scripted pilot flies 800 flights from sentences like \"take off, fly under the table, and land behind the bookcase\", and the model imitates them. Isaac Sim only produced this fine-tuning set; FLUX 3 Action itself was not trained on simulated data.\n\nWe tested it in rooms it had never seen, with sentences it had never read. The sofa, the bookcase, the pad and the clutter move on every flight, and the drone takes off facing a random direction, so \"fly to the sofa\" cannot be a memorized heading; it has to look. The wording moves too: it trained on \"fly to the bookcase\" and we ask it to \"go over to the bookshelf and wait there\". It still gets there.\n\n| *FLUX flies to a specific object then hovers* | *FLUX flies to two specific objects and then hovers* | \n\n## \n\t\tWrap-up\n\t\n\nFLUX 3 Action places first on RoboLab and runs on an SO-101 via LeRobot. It is also easy to train on tasks that have nothing to do with a robot arm: a video game shooter, a racer and a drone each took the same model and a few hundred recorded episodes.\n\nThe weights, the trainer and the games are public, and the fine-tuning guide is in the docs.\n\n- **Code and recipe:**[github.com/black-forest-labs/flux-action](https://github.com/black-forest-labs/flux-action)\n- **Weights:**[FLUX 3 Action collection](https://huggingface.co/collections/black-forest-labs/flux-3-action-6ab25aef555dd30ab86567f8)\n- **LeRobot integration:**[huggingface.co/docs/lerobot/flux3](https://huggingface.co/docs/lerobot/flux3)\n- **The games:**[bfl.ai/flux_3/flux3_action_games](https://bfl.ai/flux_3/flux3_action_games) covers the recording format, training config, results and a play loop.\n- **Research notes:**[bfl.ai/models/flux-3-action](https://bfl.ai/models/flux-3-action)\n- **DROID:**[droid-dataset.github.io](https://droid-dataset.github.io/)", "url": "https://wpnews.pro/news/flux-3-action-a-7b-open-weight-world-action-model-for-robots", "canonical_source": "https://huggingface.co/blog/black-forest-labs/flux-3-action", "published_at": "2026-09-27 23:45:17+00:00", "updated_at": "2026-09-28 00:01:18.731452+00:00", "lang": "en", "topics": ["robotics", "artificial-intelligence", "machine-learning", "ai-research", "ai-products"], "entities": ["Black Forest Labs", "FLUX 3 Action", "DROID", "RoboLab", "Cosmos3-Nano-Policy", "SO-101", "LeRobot", "Qwen3-VL-4B"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/flux-3-action-a-7b-open-weight-world-action-model-for-robots", "markdown": "https://wpnews.pro/news/flux-3-action-a-7b-open-weight-world-action-model-for-robots.md", "text": "https://wpnews.pro/news/flux-3-action-a-7b-open-weight-world-action-model-for-robots.txt", "jsonld": "https://wpnews.pro/news/flux-3-action-a-7b-open-weight-world-action-model-for-robots.jsonld"}}