{"slug": "whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing", "title": "WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing", "summary": "Researchers propose WHTMix, a Walsh-Hadamard token mixer that replaces global self-attention in stereo depth estimation transformers, reducing model compute by a factor of 2.46 and single-image inference latency by a factor of 2.65 while matching baseline end-point error on synthetic driving data. The method achieves log-linear complexity and is particularly effective for high-resolution stereo matching, with a hybrid log-disparity loss further reducing error on distant objects at no extra computational cost.", "body_md": "arXiv:2607.25234v1 Announce Type: new\nAbstract: Stereo depth estimation for driving, robotics and augmented reality must run at high resolution under tight latency budgets, yet in transformer-based matchers the global self-attention that aggregates scene context grows quadratically with the number of pixels and comes to dominate runtime. We show that the joint self-attention stage of a stereo transformer, whose role is to spread context across both views, can be replaced by a data-independent Walsh-Hadamard token mixer that mixes tokens globally in the transform domain at log-linear cost, while the data-dependent cross-attention that performs left-right correspondence is retained. On synthetic driving data the mixer matches the attention baseline in end-point error while reducing model compute by a factor of 2.46 and single-image inference latency by a factor of 2.65. A complexity analysis shows the benefit is governed by the ratio of sequence length to channel width, which explains why high-resolution stereo matching is a particularly favorable setting and why classification transformers are not; we confirm this token-to-channel scaling on non-stereo long-sequence benchmarks. Furthermore, we introduce a hybrid log-disparity loss function designed to up-weight small-disparity pixels corresponding to long-range objects. This approach reduces the error on distant objects without incurring any additional computational overhead.", "url": "https://wpnews.pro/news/whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing", "canonical_source": "https://arxiv.org/abs/2607.25234", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 04:22:42.456437+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "machine-learning", "neural-networks", "ai-research"], "entities": ["WHTMix", "Walsh-Hadamard"], "alternates": {"html": "https://wpnews.pro/news/whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing", "markdown": "https://wpnews.pro/news/whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing.md", "text": "https://wpnews.pro/news/whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing.txt", "jsonld": "https://wpnews.pro/news/whtmix-efficient-stereo-depth-estimation-via-walsh-hadamard-token-mixing.jsonld"}}