LLMs From Scratch Reaches 100,000 GitHub Stars The LLMs-from-scratch repository by Sebastian Raschka surpassed 100,000 stars on GitHub, marking a milestone for the open-source project that provides from-scratch implementations of large language models. Raschka, the author, announced the achievement and outlined plans to add new attention variants and architectures, while also working on a larger applied custom 'small' LLM project to be detailed in an upcoming Substack article. LLMs From Scratch Reaches 100,000 GitHub Stars Just saw that the LLMs-from-scratch repository https://github.com/rasbt/LLMs-from-scratch passed 100,000 stars on GitHub This is super cool and motivating. I am really happy to see that this open-source repo has helped so many people. Thanks also to everyone who shared ideas and opened PRs with improvements Of course, I plan to keep adding new material, including new attention variants and architectures while bigger projects like RL and Reasoning From Scratch live in their separate repositories . I am also currently working on a larger applied custom “small” LLM project. It has been keeping me super busy this month, but I will share more on that soon in an upcoming Substack mega-article It’s my longest one yet If you are new to it, some of the highlights in the repo include - Of course, the complete code path from tokenization https://github.com/rasbt/LLMs-from-scratch/blob/main/ch02/01 main-chapter-code/ch02.ipynb and attention https://github.com/rasbt/LLMs-from-scratch/blob/main/ch03/01 main-chapter-code/ch03.ipynb to pretraining https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/01 main-chapter-code/ch05.ipynb , classification https://github.com/rasbt/LLMs-from-scratch/blob/main/ch06/01 main-chapter-code/ch06.ipynb , and instruction fine-tuning https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/01 main-chapter-code/ch07.ipynb , etc. All of it FROM SCRATCH, of course RL lives in a companion repo. - From-scratch implementations of Llama https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/07 gpt to llama/standalone-llama32.ipynb , Qwen https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/11 qwen3 , Gemma https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/12 gemma3 , and Olmo https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/13 olmo3 smaller variants that run locally and can be plugged into the training scripts . - From-scratch implementations of attention alternatives and other architecture components, such as GQA https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/04 gqa , MLA https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/05 mla , sliding-window attention https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/06 swa , Gated DeltaNet https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/08 deltanet , DeepSeek Sparse Attention https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/09 dsa , cross-layer KV sharing https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/10 kv-sharing , and mixture-of-experts https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/07 moe - Materials on KV caching https://github.com/rasbt/LLMs-from-scratch/tree/main/ch04/03 kv-cache , training performance https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/10 llm-training-speed , memory-efficient weight loading https://github.com/rasbt/LLMs-from-scratch/blob/main/ch05/08 memory efficient weight loading/memory-efficient-state-dict.ipynb , DPO https://github.com/rasbt/LLMs-from-scratch/blob/main/ch07/04 preference-tuning-with-dpo/dpo-from-scratch.ipynb , evaluation https://github.com/rasbt/LLMs-from-scratch/tree/main/ch07/03 model-evaluation , and LoRA https://github.com/rasbt/LLMs-from-scratch/blob/main/appendix-E/01 main-chapter-code/appendix-E.ipynb So, if you don’t have any weekend plans yet, happy tinkering Source: website version of my Substack note https://substack.com/@rasbt/note/c-310007138 . Read Next Kimi K3 Architecture Notes Short architecture note on Kimi K3, including LatentMoE, Kimi Delta Attention, Attention Residuals, NoPE, multimodality, and inference-efficiency choices. /blog/2026/kimi-k3-architecture-notes.html A Few Notable Open-Weight Models This Week Short note on the architectures of six new open-weight models, including Nanbeige 4.2, Laguna S 2.1, Motif-3-Beta, Solar Open 2, Antares 1B, and BTL-3. /blog/2026/notable-open-weight-models-this-week.html Correction for Listing 6.5 in Build a Reasoning Model From Scratch Short correction note for the random seed in Listing 6.5 on page 198 of Build a Reasoning Model From Scratch. /blog/2026/reasoning-model-listing-6-5-correction.html