{"slug": "scaling-categorical-flow-maps", "title": "Scaling Categorical Flow Maps", "summary": "Researchers at Apple trained a 1.7B-parameter categorical flow model on 2.1T tokens and self-distilled it into a Categorical Flow Map (CFM) that generates text in as few as 4 inference steps while maintaining near-data-level token entropy. The team, including Oscar Davis, Anastasiia Filippova, Victor Turrisi, Amitis Shidani, Pierre Ablin, Marco Cuturi, and Louis Béthune, also introduced a likelihood bound for CFMs in the semi-discrete setting, enabling scoring on standard LM benchmarks with results comparable to discrete diffusion methods. The work addresses scalability challenges and provides insights on loss weighting and time scheduling.", "body_md": "[content type paper](/research/)published August 2026\n\nScaling Categorical Flow Maps\n\nAuthorsOscar Davis†**, Anastasiia Filippova, Victor Turrisi, Amitis Shidani, Pierre Ablin, Marco Cuturi, Louis Béthune\n\nScaling Categorical Flow Maps\n\nAuthorsOscar Davis†**, Anastasiia Filippova, Victor Turrisi, Amitis Shidani, Pierre Ablin, Marco Cuturi, Louis Béthune\n\nContinuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalities, including accelerated sampling and tilting. Recently, several works have demonstrated the possibility of generating discrete data continuously by a simple flow matching process between a Gaussian and the one-hot encoded data distribution. They have further shown the feasibility of accelerated sampling via Categorical Flow Maps (CFMs), resulting in competitive sample quality in the few-step regime. However, this method had only been evaluated at relatively modest scales (< 1B), leaving the question of its scalability completely open. In this article, we train a 1.7B-parameter base flow model on 2.1T tokens and self-distill it into a CFM that generates diverse, high-quality text in as few as 4 inference steps while maintaining near-data-level token entropy. Furthermore, we introduce a likelihood bound for CFMs in the semi-discrete setting, and show that they can be used to score the model on standard LM benchmarks, achieving results in the same range as discrete diffusion methods. Finally, we uncover some of the challenges that arise from training these models at scale, and we provide prescriptive insights on loss weighting and time scheduling.\n\nScore Distillation of Flow Matching Models\n\nDecember 16, 2025[research area Computer Vision](/research/?domain=Computer%20Vision), [research area Methods and Algorithms](/research/?domain=Methods%20and%20Algorithms)\n\nDiffusion models achieve high-quality image generation but are limited by slow iterative sampling. Distillation methods alleviate this by enabling one- or few-step generation. Flow matching, originally introduced as a distinct framework, has since been shown to be theoretically equivalent to diffusion under Gaussian assumptions, raising the question of whether distillation techniques such as score distillation transfer directly. We provide a…\n\nCAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching\n\nNovember 12, 2025[research area Computer Vision](/research/?domain=Computer%20Vision)[conference NeurIPS](/research/?event=NeurIPS)\n\nConditional generative modeling aims to learn a conditional data distribution from samples containing data-condition pairs. For this, diffusion and flow-based methods have attained compelling results. These methods use a learned (flow) model to transport an initial standard Gaussian noise that ignores the condition to the conditional data distribution. The model is hence required to learn both mass transport and conditional injection. To ease the…", "url": "https://wpnews.pro/news/scaling-categorical-flow-maps", "canonical_source": "https://machinelearning.apple.com/research/scaling-categorical-flow-maps", "published_at": "2026-08-07 00:00:00+00:00", "updated_at": "2026-08-09 13:16:56.150423+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "generative-ai", "ai-research"], "entities": ["Apple", "Oscar Davis", "Anastasiia Filippova", "Victor Turrisi", "Amitis Shidani", "Pierre Ablin", "Marco Cuturi", "Louis Béthune"], "alternates": {"html": "https://wpnews.pro/news/scaling-categorical-flow-maps", "markdown": "https://wpnews.pro/news/scaling-categorical-flow-maps.md", "text": "https://wpnews.pro/news/scaling-categorical-flow-maps.txt", "jsonld": "https://wpnews.pro/news/scaling-categorical-flow-maps.jsonld"}}