Focusing on Post-Training
Fireworks post-trained Kimi K3 to create Ember-1, a model that uses roughly 40% fewer tokens while maintaining comparable quality across Fireworks' evaluations, according to the company's announcement…
Fireworks post-trained Kimi K3 to create Ember-1, a model that uses roughly 40% fewer tokens while maintaining comparable quality across Fireworks' evaluations, according to the company's announcement…
Xiaomi's MiMo-V2.6 Pro ranks No. 1 on open-weight benchmarks by weighted average despite using a classic Grouped Query Attention architecture with Sliding Window Attention at a 128-token window, accor…
Sebastian Raschka argued in a blog post that Jev, the system-one model from Typesafe.ai, should not be dismissed as "just a classifier," crediting its generalization across tasks such as classifying e…
Sebastian Raschka released a 1.5-hour "AI Reasoning Models" course on LinkedIn Learning, offered as part of LinkedIn Premium, covering how reasoning models relate to conventional LLMs and how they are…
OpenAI released GPT-6 Astra, which scores 99.9% on the ARC-AGI-3 benchmark, up from GPT-5.6 Sol's 7.8%, and leads the Artificial Analysis Coding Agent Index v1.4, though independent benchmarks suggest…
OpenAI's Astra model is not a novel 'looped transformer' breakthrough, according to Sebastian Raschka, who explains that the technique merely reuses transformer layers and was already used in Nanbeige…
Sebastian Raschka published a video and accompanying note explaining how to set up the Python and PyTorch environment for building a reasoning model from scratch, using the `uv` package manager. The v…
Author Sebastian Raschka will host two free live Q&A sessions on September 3, 2026, about his new book Build a Reasoning Model (From Scratch), first with Sophia Yang's The AI Book Club at 10 am CT and…
The Ox Alpha LLM has been identified as GLM-5.3-Flash, a new model from Zhipu AI that introduces a hybrid attention architecture combining 34 Kimi Delta Attention layers and 11 Multi-head Latent Atten…
Anthropic announced that it will watermark text outputs from its Claude models, a technique explained in a new 48-minute video lecture by Sebastian Raschka, author of 'Build a Large Language Model Fro…
Sebastian Raschka, an AI researcher and author, published a tutorial on building an AI text detector from scratch, including dataset construction, model training, local deployment, and reinforcement l…
Sebastian Raschka's book 'Build a Reasoning Model (From Scratch)' is now available on Amazon, but readers in India are advised to order directly from Manning to avoid counterfeit black-and-white copie…
Meta released Muse Glimmer, a 30B open-weight multimodal reasoning model with a Gemma-like architecture, featuring a 131k context window, dense design, hybrid attention with a 3:1 sliding-window-to-gr…
The LLMs-from-scratch repository by Sebastian Raschka surpassed 100,000 stars on GitHub, marking a milestone for the open-source project that provides from-scratch implementations of large language mo…
Kimi K3, a 2.8-trillion-parameter open-weight model from Moonshot AI, is the largest open-weight model to date, scaling up their earlier Kimi Linear architecture from 48B parameters. The architecture …
Open-weight model releases this week include Nanbeige 4.2 3B with looped depth sharing, poolside's Laguna S 2.1 (118B sparse MoE, 8B active, 1M-token context), Motif-3-Beta (314B-A13B sparse MoE with …
Sebastian Raschka issued a correction for Listing 6.5 in his book 'Build a Reasoning Model From Scratch', changing the line 'torch.manual_seed(0)' to 'torch.manual_seed(5)' on page 198. The change is …
A new book, 'Build a Reasoning Model (From Scratch)' (ISBN-13 9781633434677), teaches readers how to add reasoning capabilities to a pre-trained large language model through hands-on code examples, co…
OpenAI released the GPT-5.6 model family last week, which comes in three sizes each with roughly five or six reasoning-effort settings, according to Sebastian Raschka. The article explains how reasoni…
Thinking Machines Lab released Inkling, an open-weight 975B-parameter sparse Mixture-of-Experts model with 41B active parameters and a 1M-token context window, which outperforms GLM-5.2 on IFBench (79…