{"slug": "adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to", "title": "Adam isn't actually tracking the natural gradient as closely as we like to pretend.", "summary": "A study posted at arxiv.org/abs/2610.00004 finds that Adam's diagonal scaling departs from the natural gradient path by up to a factor of 10^3 on ill-conditioned loss landscapes and complex neural networks, quantified with a scale-invariant γ(Δθ) metric. The research attributes the drift to three structural approximations — diagonal truncation, empirical label substitution and temporal lag — yet reports that Adam still reaches low loss, crediting momentum smoothing rather than fidelity to natural gradient descent. The same study finds the standard empirical Fisher frequently oscillates or diverges, while the improved empirical Fisher (iEF) is significantly more stable.", "body_md": "# Adam isn't actually tracking the natural gradient as closely as we like to pretend.\n\n## Why does Adam drift from the natural gradient?\n\nThe common assumption is that Adam's diagonal scaling is just a \"cheap\" version of the Fisher Information Matrix used in NGD. In reality, Adam's update rule is a mess of approximations. It doesn't just simplify the matrix; it suffers from diagonal truncation, empirical label substitution, and temporal lag. These aren't just academic footnotes—they are the reasons why Adam doesn't actually follow the steepest descent on the Riemannian manifold.\n\nThe research uses a scale-invariant $\\gamma(\\Delta\\theta)$ metric to quantify exactly how far off the rails Adam goes. When the loss landscape is well-conditioned, like in simple linear regression, the deviation is low. But once you hit ill-conditioned settings or complex neural networks, the \"approximation\" becomes a complete departure from the natural gradient path.\n\n## Does this drift actually break the training?\n\nSurprisingly, no. You'd think a misalignment of $10^3$ would send your weights flying into orbit, but the results show that while high geometric drift correlates with slower initial optimization, it doesn't stop Adam from eventually reaching a low loss.\n\nThe real takeaway is that Adam's success isn't because it's a great approximation of NGD. Instead, it's likely a lucky balance between those structural approximation errors and the smoothing effect of momentum. It’s essentially stumbling its way to the minimum, but it does it efficiently enough that we don't care about the geometric inaccuracy.\n\n## How does the Empirical Fisher behave?\n\nThe study also looks at the standard empirical Fisher (EF) versus the improved empirical Fisher (iEF). If you're trying to track the natural gradient, the standard EF is a nightmare—it frequently oscillates or diverges entirely. The iEF is significantly more stable, which suggests that the \"standard\" way of approximating the Fisher matrix is often too volatile for actual use.\n\nIf you want to see the raw data or the $\\gamma(\\Delta\\theta)$ measurements, the full breakdown is at `https://arxiv.org/abs/2610.00004`.\n\n## What to do when optimization slows down?\n\nSince we now know that ill-conditioned landscapes cause massive geometric drift and slow down early training, you can't just blame your learning rate. If your loss is plateauing early or crawling at a snail's pace despite a reasonable LR, you're likely seeing the effect of this misalignment.\n\nSince Adam consistently reaches low loss eventually, the \"fix\" isn't necessarily to switch to a full NGD optimizer (which is computationally expensive), but to recognize that the initial slow-down is a structural feature of the optimizer's drift. If you are seeing extreme instability, checking if your Fisher approximation is oscillating like the standard EF might be the move, though for most of us, just letting Adam grind through the drift is the only practical option.\n\n[Next How do you navigate the labyrinth of AI-generated code? →](https://promptcube3.com/en/threads/9751/)\n\n## All Replies （1）\n\nWant a live back-and-forth? [Join the global AI chat room](https://promptcube3.com/en/chat/) — login to talk.\n\nThe diagonal truncation makes sense, but I'm curious how much of that drift comes from the temporal lag versus just the label substitution in practice.", "url": "https://wpnews.pro/news/adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to", "canonical_source": "https://promptcube3.com/en/threads/9752/", "published_at": "2026-10-03 16:14:10+00:00", "updated_at": "2026-10-03 16:36:15.498826+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence", "ai-research"], "entities": ["Adam", "natural gradient descent", "Fisher Information Matrix", "empirical Fisher", "improved empirical Fisher", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to", "markdown": "https://wpnews.pro/news/adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to.md", "text": "https://wpnews.pro/news/adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to.txt", "jsonld": "https://wpnews.pro/news/adam-isn-t-actually-tracking-the-natural-gradient-as-closely-as-we-like-to.jsonld"}}