cd/sources/sebastianraschka-auto-discovered· home› sources› Sebastianraschka (auto-discovered)
cat /sources/sebastianraschka-auto-discovered.feed | wc -l → 28

Sebastianraschka (auto-discovered)

articles 28 domain sebastianraschka.com → page 1/2 feed RSS
22:12
2026-09-27
sebastianraschka.com
large-language-models

Focusing on Post-Training

Fireworks post-trained Kimi K3 to create Ember-1, a model that uses roughly 40% fewer tokens while maintaining comparable quality across Fireworks' evaluations, according to the company's announcement…

13:47
2026-09-22
sebastianraschka.com
large-language-models

MiMo-V2.6 Pro Architecture and Training Notes

Xiaomi's MiMo-V2.6 Pro ranks No. 1 on open-weight benchmarks by weighted average despite using a classic Grouped Query Attention architecture with Sliding Window Attention at a 128-token window, accor…

15:17
2026-09-20
sebastianraschka.com
machine-learning

It's Easy to Dismiss Jev as Just a Classifier

Sebastian Raschka argued in a blog post that Jev, the system-one model from Typesafe.ai, should not be dismissed as "just a classifier," crediting its generalization across tasks such as classifying e…

13:11
2026-09-13
sebastianraschka.com
large-language-models

AI Reasoning Models Course on LinkedIn Learning

Sebastian Raschka released a 1.5-hour "AI Reasoning Models" course on LinkedIn Learning, offered as part of LinkedIn Premium, covering how reasoning models relate to conventional LLMs and how they are…

11:14
2026-09-09
sebastianraschka.com
artificial-intelligence

GPT-6 Astra, Looped Transformers, and Hidden Reasoning

OpenAI released GPT-6 Astra, which scores 99.9% on the ARC-AGI-3 benchmark, up from GPT-5.6 Sol's 7.8%, and leads the Artificial Analysis Coding Agent Index v1.4, though independent benchmarks suggest…

08:30
2026-09-02
sebastianraschka.com
artificial-intelligence

OpenAI Astra and Looped Transformers

OpenAI's Astra model is not a novel 'looped transformer' breakthrough, according to Sebastian Raschka, who explains that the technique merely reuses transformer layers and was already used in Nanbeige…

08:42
2026-08-30
sebastianraschka.com
artificial-intelligence

Reasoning Models From Scratch: Code Setup

Sebastian Raschka published a video and accompanying note explaining how to set up the Python and PyTorch environment for building a reasoning model from scratch, using the `uv` package manager. The v…

11:55
2026-08-29
sebastianraschka.com
large-language-models

Two Live Book Club Q&As on September 3

Author Sebastian Raschka will host two free live Q&A sessions on September 3, 2026, about his new book Build a Reasoning Model (From Scratch), first with Sophia Yang's The AI Book Club at 10 am CT and…

10:11
2026-08-26
sebastianraschka.com
large-language-models

GLM-5.3-Flash Architecture Notes

The Ox Alpha LLM has been identified as GLM-5.3-Flash, a new model from Zhipu AI that introduces a hybrid attention architecture combining 34 Kimi Delta Attention layers and 11 Multi-head Latent Atten…

11:11
2026-08-22
sebastianraschka.com
generative-ai

How Claude Watermarks AI-Generated Text

Anthropic announced that it will watermark text outputs from its Claude models, a technique explained in a new 48-minute video lecture by Sebastian Raschka, author of 'Build a Large Language Model Fro…

11:54
2026-08-15
sebastianraschka.com
artificial-intelligence

Building an AI Text Detector From Scratch

Sebastian Raschka, an AI researcher and author, published a tutorial on building an AI text detector from scratch, including dataset construction, model training, local deployment, and reinforcement l…

08:25
2026-08-12
sebastianraschka.com
artificial-intelligence

Build a Reasoning Model From Scratch Is Now on Amazon

Sebastian Raschka's book 'Build a Reasoning Model (From Scratch)' is now available on Amazon, but readers in India are advised to order directly from Manning to avoid counterfeit black-and-white copie…

09:15
2026-08-11
sebastianraschka.com
artificial-intelligence

Muse Glimmer 30B Architecture Notes

Meta released Muse Glimmer, a 30B open-weight multimodal reasoning model with a Gemma-like architecture, featuring a 131k context window, dense design, hybrid attention with a 3:1 sliding-window-to-gr…

09:40
2026-08-07
sebastianraschka.com
large-language-models

LLMs From Scratch Reaches 100,000 GitHub Stars

The LLMs-from-scratch repository by Sebastian Raschka surpassed 100,000 stars on GitHub, marking a milestone for the open-source project that provides from-scratch implementations of large language mo…

08:38
2026-07-28
sebastianraschka.com
large-language-models

Kimi K3 Architecture Notes

Kimi K3, a 2.8-trillion-parameter open-weight model from Moonshot AI, is the largest open-weight model to date, scaling up their earlier Kimi Linear architecture from 48B parameters. The architecture …

08:47
2026-07-26
sebastianraschka.com
large-language-models

A Few Notable Open-Weight Models This Week

Open-weight model releases this week include Nanbeige 4.2 3B with looped depth sharing, poolside's Laguna S 2.1 (118B sparse MoE, 8B active, 1M-token context), Motif-3-Beta (314B-A13B sparse MoE with …

08:18
2026-07-25
sebastianraschka.com
large-language-models

Correction for Listing 6.5 in Build a Reasoning Model From Scratch

Sebastian Raschka issued a correction for Listing 6.5 in his book 'Build a Reasoning Model From Scratch', changing the line 'torch.manual_seed(0)' to 'torch.manual_seed(5)' on page 198. The change is …

16:26
2026-07-21
sebastianraschka.com
large-language-models

Build a Reasoning Model (From Scratch)

A new book, 'Build a Reasoning Model (From Scratch)' (ISBN-13 9781633434677), teaches readers how to add reasoning capabilities to a pre-trained large language model through hands-on code examples, co…

11:16
2026-07-18
sebastianraschka.com
large-language-models

Controlling Reasoning Effort in LLMs

OpenAI released the GPT-5.6 model family last week, which comes in three sizes each with roughly five or six reasoning-effort settings, according to Sebastian Raschka. The article explains how reasoni…

14:03
2026-07-16
sebastianraschka.com
large-language-models

Inkling: A New Open-Weight 975B Moe with a Few Surprises

Thinking Machines Lab released Inkling, an open-weight 975B-parameter sparse Mixture-of-Experts model with 41B active parameters and a 1M-token context window, which outperforms GLM-5.2 on IFBench (79…

page 1 / 2 next →