{"slug": "google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a", "title": "Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model", "summary": "Google DeepMind retrofitted Gemma 4 into a text diffusion model, DiffusionGemma, using less than 10 percent of the original training budget, and it generates 256 tokens in parallel at about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks.", "body_md": "# Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model\n\n[The Decoder](https://the-decoder.com)\n\nInstead of [training](/glossary/training) a new model from scratch, Google [DeepMind](/glossary/deepmind) retrofitted Gemma 4 into a [diffusion model](/glossary/diffusion-model) using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original [autoregressive model](/glossary/autoregressive-model) in benchmarks, especially on [reasoning](/glossary/reasoning) tasks.\n\nThe article [Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model](https://the-decoder.com/googles-diffusiongemma-proves-you-dont-need-to-train-from-scratch-to-build-a-text-diffusion-model/) appeared first on [The Decoder](https://the-decoder.com).\n\nGet AI news in your inbox\n\nDaily digest of what matters in AI.\n\n## Key Terms Explained\n\n[Autoregressive Model](/glossary/autoregressive-model)\n\nA model that generates output one piece at a time, with each new piece depending on all the previous ones.\n\n[Decoder](/glossary/decoder)\n\nThe part of a neural network that generates output from an internal representation.\n\n[DeepMind](/glossary/deepmind)\n\nA leading AI research lab, now part of Google.\n\n[Diffusion Model](/glossary/diffusion-model)\n\nA generative AI model that creates data by learning to reverse a gradual noising process.", "url": "https://wpnews.pro/news/google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a", "canonical_source": "https://www.machinebrief.com/news/googles-diffusiongemma-proves-you-dont-need-to-train-from-sc-xvqz", "published_at": "2026-08-09 10:01:26+00:00", "updated_at": "2026-08-09 12:30:54.579142+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research"], "entities": ["Google DeepMind", "Gemma 4", "DiffusionGemma", "The Decoder"], "alternates": {"html": "https://wpnews.pro/news/google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a", "markdown": "https://wpnews.pro/news/google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a.md", "text": "https://wpnews.pro/news/google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a.txt", "jsonld": "https://wpnews.pro/news/google-s-diffusiongemma-proves-you-don-t-need-to-train-from-scratch-to-build-a.jsonld"}}