{"slug": "starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal", "title": "STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation", "summary": "Researchers from UIUC and Apple introduced STARFlow2, a unified multimodal model built on the Pretzel architecture that interleaves a frozen pretrained vision-language model with a TARFlow stream via residual skip connections under a single causal mask, enabling continuous, single-pass, and purely causal generation of interleaved text and images. The model preserves pretrained multimodal understanding while achieving high-fidelity image generation, and experiments show strong performance on image generation and multimodal understanding benchmarks.", "body_md": "Unified multimodal models that understand, reason over, and generate interleaved text–image sequences remain structurally fragmented: existing approaches either sacrifice visual fidelity through discrete tokenization, impose structural asymmetry by combining causal text generation with iterative diffusion-based denoising, or degrade pretrained understanding when adapting vision-language models for generation. We observe that autoregressive normalizing flows are autoregressive Transformers—sharing the same causal mask, KV-cache mechanism, and left-to-right structure as LLMs—making them the most natural paradigm for truly unified multimodal generation that is continuous, single-pass, and purely causal. We present STARFlow2, built on the Pretzel architecture that vertically interleaves a frozen pretrained VLM stream with a TARFlow stream via residual skip connections, both operating under the same causal mask. This design simultaneously preserves pretrained multimodal understanding, enables high-fidelity continuous image generation, and achieves structural unification under a single causal mechanism. Combined with a deep-shallow flow design and a unified FAE latent space, STARFlow2 supports cache-friendly interleaved generation where both text and visual outputs directly enter the KV-cache without re-encoding. Experiments demonstrate strong performance across image generation and multimodal understanding benchmarks, validating autoregressive flows as a viable foundation for unified multimodal modeling.\n\n- † UIUC\n- ** Work done while at Apple", "url": "https://wpnews.pro/news/starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal", "canonical_source": "https://machinelearning.apple.com/research/starflow2-multimodal-generation", "published_at": "2026-08-25 00:00:00+00:00", "updated_at": "2026-08-25 14:16:55.403013+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "large-language-models", "computer-vision"], "entities": ["UIUC", "Apple", "STARFlow2", "Pretzel", "TARFlow"], "alternates": {"html": "https://wpnews.pro/news/starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal", "markdown": "https://wpnews.pro/news/starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal.md", "text": "https://wpnews.pro/news/starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal.txt", "jsonld": "https://wpnews.pro/news/starflow2-bridging-language-models-and-normalizing-flows-for-unified-multimodal.jsonld"}}