{"slug": "glm-5-3-flash-architecture-notes", "title": "GLM-5.3-Flash Architecture Notes", "summary": "The Ox Alpha LLM has been identified as GLM-5.3-Flash, a new model from Zhipu AI that introduces a hybrid attention architecture combining 34 Kimi Delta Attention layers and 11 Multi-head Latent Attention/DeepSeek Sparse Attention layers, and scales down its sparse MoE backbone from 744B-A40B to 320B-A18B. The model also features a DeepSeek V4-style mHC residual path with four parallel streams and a native vision encoder, according to architecture notes by Sebastian Raschka.", "body_md": "# GLM-5.3-Flash Architecture Notes\n\nNow we know: The popular Ox Alpha LLM was GLM-5.3-Flash…\n\nCompared to GLM-5.2, this new GLM-5.3-Flash model uses:\n\n-\na Kimi Linear-style 3:1 (super*) hybrid attention pattern with 34 Kimi Delta Attention layers (KDA) and 11 Multi-head Latent Attention (MLA) / DeepSeek Sparse Attention (DSA) layers;\n\n-\na scaled-down GLM-5.2-style sparse MoE backbone, going from 744B-A40B to 320B-A18B;\n\n-\na DeepSeek V4-style mHC residual path with four parallel streams;\n\n-\nplus a native vision encoder (not shown).\n\nI called it a “super hybrid” above because both KDA and MLA/DSA are “efficient” components. E.g., Kimi only uses KDA + full attention MLA, DeepSeek V3.2 uses DSA + full attention MLA.\n\nPS: I’m sorry for the excessive tech jargon. Explainers on all these components (MLA, DSA, KDA, mhC, etc.) in my [LLM Architecture Gallery](/llm-architecture-gallery/).\n\nPPS: Haha, maybe justification for getting that pricey Mac Studio M5 Ultra 256 GB / 512 GB to run this locally…\n\nSource: website version of my [Substack note](https://substack.com/@rasbt/note/c-323088504).\n\n## Read Next\n\n[How Claude's Text Watermarking Works Short illustration of how Claude's text watermarking is supposed to work based on Anthropic's released materials.](/blog/2026/claude-text-watermarking.html)\n\n[Build a Reasoning Model From Scratch Is Now on Amazon Short note on the Amazon availability of Build a Reasoning Model From Scratch and a warning about counterfeit black-and-white copies sold through Amazon In](/blog/2026/build-a-reasoning-model-from-scratch-on-amazon.html)\n\n[Muse Glimmer 30B Architecture Notes Short architecture note on Meta Muse Glimmer 30B, including gated local and global GQA, KV-cache efficiency, and release-time benchmark comparisons.](/blog/2026/muse-glimmer-30b-architecture-notes.html)", "url": "https://wpnews.pro/news/glm-5-3-flash-architecture-notes", "canonical_source": "https://sebastianraschka.com/blog/2026/glm-5-3-flash-architecture-notes.html", "published_at": "2026-08-26 10:11:51+00:00", "updated_at": "2026-08-27 13:49:44.933006+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "artificial-intelligence"], "entities": ["Ox Alpha", "GLM-5.3-Flash", "Zhipu AI", "Kimi", "DeepSeek", "Sebastian Raschka"], "alternates": {"html": "https://wpnews.pro/news/glm-5-3-flash-architecture-notes", "markdown": "https://wpnews.pro/news/glm-5-3-flash-architecture-notes.md", "text": "https://wpnews.pro/news/glm-5-3-flash-architecture-notes.txt", "jsonld": "https://wpnews.pro/news/glm-5-3-flash-architecture-notes.jsonld"}}