UnifoLM-WLA-1.0 UnifoLM-ER-1, a 4B-parameter embodied reasoning model built on Qwen3-VL-4B, leads open-source models on seven of 16 multimodal perception and understanding benchmarks and matches leading proprietary models overall, according to its published benchmark results. The model was trained on more than 5 million samples covering image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering, co-trained with general image-text data. UnifoLM-ER-1-4B scored 62.4 on RoboVQA, 82.0 on Where2Place, 93.4 on BLINK, 88.6 on CV-Bench, and 54.7 on MMMU_VAL. Embodied Reasoning Across 16 multimodal perception and understanding benchmarks, UnifoLM-ER-1 leads open-source models on seven and delivers overall performance comparable to leading proprietary models. Built on Qwen3-VL-4B, UnifoLM-ER-1 is trained on more than 5 million samples spanning image point prediction, object detection, multi-image reasoning, 2D trajectory prediction, 3D object detection, and multi-image spatial question answering. These data are co-trained with general image–text data, preserving broad vision-language capabilities while substantially improving spatial understanding and reasoning in embodied environments. Benchmark results | Model | Open Source | Spatial Understanding | | | | | | | | | | | | | Multimodal Understanding | | | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | | | RoboVQA | Ego-Plan2 | RefSpatial- Bench | Where2Place | Pixmo-Point | BLINK | CV-Bench | EmbSpatial | RoboSpatial | SAT | VSI-Bench | VSR | ERQA | RealWorld QA | MME | MMMU VAL | | UnifoLM-ER-1-4B | Yes | 62.4 | 55.1 | 61.7 | 82.0 | 73.8 | 93.4