{"slug": "a-defense-of-the-quadratic-model", "title": "A Defense of the Quadratic Model", "summary": "A new study from researchers testing the quadratic model of optimization on a 150M-parameter LLM trained on 3B tokens finds that the simple model can predict optimization dynamics over windows lasting up to 10% of training. Using Lanczos quadrature, the team analyzed the Hessian spectrum and local stability, revealing structure dependent on batch size, preconditioner, and training time, and showing that LLM optimization typically occurs at a stochastic edge of stability.", "body_md": "arXiv:2607.21716v1 Announce Type: new\nAbstract: Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model of optimization -- the quadratic model -- and show that it can be surprisingly predictive in an LLM setting with 150M parameters and 3B training tokens. Specifically, we show that Taylor expanding the model and the loss function at intermediate checkpoints through training can accurately predict the optimization dynamics over windows that can last up to 10\\% of training. Having established this agreement, we then turn to analyzing the structure of these local quadratic optimization problems through two lenses: the Hessian spectrum and local stability. Using Lanczos quadrature with extremely deep probes, we are able to estimate the Hessian spectrum deep into the tail, and we find a surprising amount of structure in both the eigenvalues and eigenvectors, which depends on the batch size, preconditioner, and training time. We also empirically test local linear stability at intermediate checkpoints and compare it to theoretical predictions to demonstrate that optimization in LLMs typically occurs at a stochastic edge of stability, whose nature is also determined by batch size. Our results indicate the quadratic model may be a theoretically tractable proxy for pretraining optimization dynamics.", "url": "https://wpnews.pro/news/a-defense-of-the-quadratic-model", "canonical_source": "https://arxiv.org/abs/2607.21716", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:09:17.894396+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "neural-networks", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/a-defense-of-the-quadratic-model", "markdown": "https://wpnews.pro/news/a-defense-of-the-quadratic-model.md", "text": "https://wpnews.pro/news/a-defense-of-the-quadratic-model.txt", "jsonld": "https://wpnews.pro/news/a-defense-of-the-quadratic-model.jsonld"}}