cd /news/machine-learning/a-data-dependent-early-stopping-rule… · home topics machine-learning article
[ARTICLE · art-111346] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

A Data-dependent Early Stopping Rule using Rademacher Complexity with L1-norm

Researchers introduced an analytic early stopping rule for neural network training that estimates the optimal stopping time without training, using Rademacher complexity with L1-norm instead of L2-norm. The method, applicable to linear models and linear regression, avoids probabilistic assumptions on data or covariance eigenvalues and extends to nonlinear networks via linear probing, as demonstrated on MNIST classification.

read1 min views1 publishedAug 26, 2026

arXiv:2608.24210v1 Announce Type: new Abstract: Training neural networks requires balancing the trade-off between fitting the training data and achieving robust performance on unseen inputs. This ability, commonly referred to as generalizability, is determined by the gap between the empirical risk on the training set (empirical loss'') and the expected risk over the data distribution (generalization error''). Existing approaches typically estimate the generalization error numerically, requiring gradient descent training and an early stopping'' strategy. In this work, we introduce an analytic framework that estimates the optimal time of early stopping without the need for training. Several works in the literature also give such analytical estimations, but they are generally based on random matrix theory and often make assumptions on the distribution of the data or the eigenvalue distribution of the covariance matrix. In contrast, our work is based on Rademacher complexity (RC) without needing such probabilistic assumptions. For both theoretical and numerical reasons, it is more relevant to express RC with the L1- norm rather than with the L2-norm. We focus on the case of linear models and the problem of linear regression. Thanks to the linear probing'' method, our results can, however, be successfully applied to nonlinear neural networks, as illustrated in the classification MNIST example.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-data-dependent-ear…] indexed:0 read:1min 2026-08-26 ·