cd /news/machine-learning/drive-vs-decay-on-the-training-dynam… · home › topics › machine-learning › article
[ARTICLE · art-145147] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures

A new arXiv paper (2610.02344v1) develops an early-training stability theory for Joint-Embedding Predictive Architectures (JEPAs) that explains representation collapse through a per-mode stability ratio μ_i = γ_i / σ_i, separating a driving force (γ) from a decay effect (σ). The framework predicts a phase boundary confirmed empirically across more than 800 Tabular-JEPA configurations and unifies predictor scaling, masking ratio, and EMA as mechanisms for shifting μ. Guided by the analysis, the authors introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation, improving effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet, with code available at https://github.com/jose-melo/drive-vs-decay.

by read1 min views4 publishedOct 5, 2026

arXiv:2610.02344v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($\gamma$) and a decay effect ($\sigma$). Under approximate spectral decoupling, a per-mode stability ratio $\mu_i = \gamma_i / \sigma_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $\mu$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at https://github.com/jose-melo/drive-vs-decay.

── more in #machine-learning 4 stories · sorted by recency
── more on @joint-embedding predictive architectures 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/drive-vs-decay-on-th…] indexed:0 read:1min 2026-10-05 · —