{"slug": "drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive", "title": "Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures", "summary": "A new arXiv paper (2610.02344v1) develops an early-training stability theory for Joint-Embedding Predictive Architectures (JEPAs) that explains representation collapse through a per-mode stability ratio μ_i = γ_i / σ_i, separating a driving force (γ) from a decay effect (σ). The framework predicts a phase boundary confirmed empirically across more than 800 Tabular-JEPA configurations and unifies predictor scaling, masking ratio, and EMA as mechanisms for shifting μ. Guided by the analysis, the authors introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation, improving effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet, with code available at https://github.com/jose-melo/drive-vs-decay.", "body_md": "arXiv:2610.02344v1 Announce Type: new \nAbstract: Joint-Embedding Predictive Architectures (JEPAs) are prone to representation collapse, typically mitigated through empirical heuristics. We develop an early-training stability theory that unifies these heuristics. Linearising the coupled JEPA gradient flow around the trivial fixed point reveals two competing effects: a driving force ($\\gamma$) and a decay effect ($\\sigma$). Under approximate spectral decoupling, a per-mode stability ratio $\\mu_i = \\gamma_i / \\sigma_i$ factorises into independent data-side and predictor-side terms and the count of unstable modes tracks the rank of representations that can emerge. The framework predicts a phase boundary, which we confirm empirically across more than 800 Tabular-JEPA configurations. It also unifies predictor scaling, masking ratio, and EMA as distinct mechanisms for shifting $\\mu$. Guided by this analysis, we introduce ResidualPred, a transformer predictor whose attention is biased toward the identity at initialisation; it improves both effective rank and downstream accuracy on tabular benchmarks and in I-JEPA pretraining on CIFAR-10, CIFAR-100, STL-10, and ImageNet. Our framework connects empirical collapse-avoidance heuristics to an explicit dynamical picture, yielding theory-driven stabilizers. Code is available at https://github.com/jose-melo/drive-vs-decay.", "url": "https://wpnews.pro/news/drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive", "canonical_source": "https://arxiv.org/abs/2610.02344", "published_at": "2026-10-05 04:00:00+00:00", "updated_at": "2026-10-05 04:12:45.022890+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "neural-networks", "artificial-intelligence"], "entities": ["Joint-Embedding Predictive Architectures", "JEPA", "ResidualPred", "I-JEPA", "CIFAR-10", "CIFAR-100", "STL-10", "ImageNet"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive", "markdown": "https://wpnews.pro/news/drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive.md", "text": "https://wpnews.pro/news/drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive.txt", "jsonld": "https://wpnews.pro/news/drive-vs-decay-on-the-training-dynamics-of-joint-embedding-predictive.jsonld"}}