cd /news/artificial-intelligence/the-jepa-predictor-a-transferable-op… · home topics artificial-intelligence article
[ARTICLE · art-66390] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

The JEPA Predictor: A Transferable Operator for Occluded Feature Completion

Researchers show that the frozen predictor from Joint-Embedding Predictive Architectures (JEPAs) can be transferred across encoder families to improve occluded feature completion. In experiments, pairing CLIP with the I-JEPA predictor lifted fine-grained Stanford Dogs accuracy from 15.9% to 52.1% (+36 percentage points) at heavy occlusion, using only 500 ImageNet-1k images to fit a linear projection between feature spaces. The portable operator requires no retraining of either model and provides a benefit that grows monotonically with mask fraction.

read1 min views2 publishedJul 21, 2026

arXiv:2607.16274v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and reads features from the encoder alone. The predictor is, by construction, a learned operator from visible-context features to features at masked positions, the structure a partial-view classifier needs. We show that this operator is portable across encoder families. We first establish that, at heavy mask, retaining the frozen predictor on a JEPA encoder substantially closes the accuracy gap against the strongest non-JEPA discriminative baselines. We then bolt the frozen predictors of I-JEPA and V-JEPA 2 onto four non-JEPA hosts (CLIP, DINOv3, DINOv2, MAE) through a single linear projection between feature spaces, fit in closed form on 500 ImageNet-1k images. Across both ImageNet-9 and Stanford Dogs and across three mask fractions, the lift over each host's masked-encoder baseline grows monotonically with the mask fraction K in every host-donor pair. CLIP paired with the I-JEPA predictor recovers most of the accuracy that masking removed on ImageNet-9 at heavy occlusion, and lifts fine-grained Stanford Dogs from 15.9% to 52.1% (+36 pp). The mechanism is identifiable: the projection pays a fixed cost on visible patches and the predictor provides a growing benefit on masked patches; the benefit dominates the heavy-occlusion regime. At low K on fine-grained classification the projection cost exceeds the benefit, defining the boundary where the linear bridge breaks down. The frozen JEPA predictor functions as a portable operator for occluded feature completion across encoder families, requiring no retraining of either model while fitting matched linear probes per mask fraction.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @joint-embedding predictive architectures 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-jepa-predictor-a…] indexed:0 read:1min 2026-07-21 ·