{"slug": "smat-attention-structured-long-context-sequence-modeling", "title": "SMat-Attention: Structured Long-Context Sequence Modeling", "summary": "Researchers introduced Structured Matrix Attention (SMat-Attention), a long-context sequence modeling method that uses a family of causal masks with structured long-range routing whose row supports have VC-dimension d, according to a paper posted as arXiv:2609.36062v1. The construction recovers the standard causal mask at d=1, and for sequences of length T the hard-routing variant takes O(T^{2-3/d}+T) work despite a dense mask, with constant-time per-token decoding after the distant prefix using O(T^{1-1/d}) cached states. Extensions to Mamba-2 and Gated DeltaNet using learned routing with top-k query reads retain subquadratic prefill and improve recall accuracy over the backbones in several settings.", "body_md": "arXiv:2609.36062v1 Announce Type: new \nAbstract: Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension $d$. In our construction, $d=1$ recovers the standard causal mask, and increasing $d$ permits richer subset-routing patterns. We give chunkwise forward and backward algorithms to enable hardware-efficiency. For sequences of length $T$, the hard-routing construction takes $O(T^{2-3/d}+T)$ work, despite the mask being dense, for our prescribed family. In fixed-horizon streaming, decoding after the distant prefix takes constant time per token using $O(T^{1-1/d})$ cached states. SMat-Attention therefore makes VC-dimension an explicit knob governing access-pattern complexity, prefill cost, and decoding memory. Empirically, subset-routing and rule-assisted multi-key retrieval experiments illustrate the masks' routing expressiveness. Extensions to Mamba-2 and Gated DeltaNet using learned routing with top-$k$ query reads retain subquadratic prefill, improve recall accuracy over the backbones in several settings, and achieve comparable small-scale language-modeling performance.", "url": "https://wpnews.pro/news/smat-attention-structured-long-context-sequence-modeling", "canonical_source": "https://arxiv.org/abs/2609.36062", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:18:13.278383+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["SMat-Attention", "Structured Matrix Attention", "Mamba-2", "Gated DeltaNet", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/smat-attention-structured-long-context-sequence-modeling", "markdown": "https://wpnews.pro/news/smat-attention-structured-long-context-sequence-modeling.md", "text": "https://wpnews.pro/news/smat-attention-structured-long-context-sequence-modeling.txt", "jsonld": "https://wpnews.pro/news/smat-attention-structured-long-context-sequence-modeling.jsonld"}}