19:52
2026-07-21
github.com
large-language-models
I trained a 30M-param LLM from scratch and the scaling "floor" was a mirage
A developer trained a 30M-parameter decoder-only transformer from scratch on the TinyStories dataset using Kaggle's free T4 GPUs, achieving a validation loss of 1.401 at a learning rate of 1e-3. The pโฆ