cd /news/artificial-intelligence/whtmix-efficient-stereo-depth-estima… · home topics artificial-intelligence article
[ARTICLE · art-78039] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing

Researchers propose WHTMix, a Walsh-Hadamard token mixer that replaces global self-attention in stereo depth estimation transformers, reducing model compute by a factor of 2.46 and single-image inference latency by a factor of 2.65 while matching baseline end-point error on synthetic driving data. The method achieves log-linear complexity and is particularly effective for high-resolution stereo matching, with a hybrid log-disparity loss further reducing error on distant objects at no extra computational cost.

read1 min views1 publishedJul 29, 2026

arXiv:2607.25234v1 Announce Type: new Abstract: Stereo depth estimation for driving, robotics and augmented reality must run at high resolution under tight latency budgets, yet in transformer-based matchers the global self-attention that aggregates scene context grows quadratically with the number of pixels and comes to dominate runtime. We show that the joint self-attention stage of a stereo transformer, whose role is to spread context across both views, can be replaced by a data-independent Walsh-Hadamard token mixer that mixes tokens globally in the transform domain at log-linear cost, while the data-dependent cross-attention that performs left-right correspondence is retained. On synthetic driving data the mixer matches the attention baseline in end-point error while reducing model compute by a factor of 2.46 and single-image inference latency by a factor of 2.65. A complexity analysis shows the benefit is governed by the ratio of sequence length to channel width, which explains why high-resolution stereo matching is a particularly favorable setting and why classification transformers are not; we confirm this token-to-channel scaling on non-stereo long-sequence benchmarks. Furthermore, we introduce a hybrid log-disparity loss function designed to up-weight small-disparity pixels corresponding to long-range objects. This approach reduces the error on distant objects without incurring any additional computational overhead.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @whtmix 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/whtmix-efficient-ste…] indexed:0 read:1min 2026-07-29 ·