18:45
2026-09-10
gilesthomas.com
large-language-models
Extending Raschka's GPT-2: an MoE trained from scratch on an RTX 3090
A developer extended Sebastian Raschka's GPT-2-style code from the book "Build a Large Language Model (from Scratch)" to add mixture-of-experts support and trained a 446M-parameter model with 220M act…