cd /news/artificial-intelligence/adaptive-depth-in-looped-transformer… · home topics artificial-intelligence article
[ARTICLE · art-71395] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

A study from arXiv (2607.20519v1) finds that adaptive depth in looped Transformers is a joint problem of trajectory formation and exit readout, not just gate learning. Researchers at Ouroboros (Ouro) evaluated 1.4B and 2.6B parameter checkpoints on synthetic tasks and found that fixed-prior depth supervision and post-hoc confidence readouts often outperform learned gates, with measured latency confirming practical inference-time savings.

read1 min views1 publishedJul 24, 2026

arXiv:2607.20519v1 Announce Type: new Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block. Learned halting objectives in looped Transformers typically use a single exit distribution both as the inference-time stopping rule and as the training-time weighting of per-depth losses. This entangles exit selection with trajectory formation: the gate not only chooses which recurrent state to use, but also determines how strongly each intermediate state is supervised. Consequently, poor adaptive-compute performance can arise from the readout, the induced trajectory, or their interaction. We study adaptive depth in looped Transformers through this trajectory--readout lens, across controlled synthetic tasks (modular arithmetic and binary parity) and large-scale Ouro-1.4B and 2.6B checkpoints. We find that fixed-prior depth supervision, which shapes the trajectory without an input-dependent halting policy, produces difficulty-aware trajectories whose intermediate states expose useful stopping signals, and that simple post-hoc confidence readouts often match or outperform learned linear and MLP gates. Fitting gates on frozen trajectories localizes the failure: it appears to stem mainly from the trajectory induced by joint gate training rather than from limited gate expressivity. The same pattern is present in Ouro evaluations, where pretrained ponder gates are competitive but not uniformly Pareto-optimal, and measured latency confirms that the resulting reductions in average exit depth translate into practical inference-time savings. Our systematic diagnostic evaluation reframes adaptive depth in looped Transformers as a joint problem of trajectory formation and exit readout, rather than gate learning alone, highlighting a distinction that prior learned-halting work has often left implicit.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/adaptive-depth-in-lo…] indexed:0 read:1min 2026-07-24 ·