{"slug": "decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel", "title": "Decay-Gated O(N) Causal Linear Attention with Fused Triton Kernel", "summary": "A developer has open-sourced a Decay-Gated O(N) Causal Linear Attention architecture with fused Triton/CUDA kernels, aiming to bypass quadratic multi-head attention bottlenecks. The project includes a repository and research paper, and the developer seeks integration into Hugging Face Transformers and benchmark comparisons.", "body_md": "Hey Hugging Face Community,\n\nI’ve engineered and open-sourced a Decay-Gated O(N) Causal Linear Attention architecture featuring fused Triton/CUDA state accumulation kernels designed to bypass quadratic MHA bottlenecks.\n\nKey Specs & Benchmark Metrics:\n\nRepository & Research Paper:\n\nWould love to gather thoughts on integration into Hugging Face Transformers model classes or benchmark comparisons with existing linear models!", "url": "https://wpnews.pro/news/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel", "canonical_source": "https://discuss.huggingface.co/t/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel/178456#post_1", "published_at": "2026-08-04 10:58:34+00:00", "updated_at": "2026-08-04 11:39:39.741289+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["Hugging Face", "Triton", "CUDA"], "alternates": {"html": "https://wpnews.pro/news/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel", "markdown": "https://wpnews.pro/news/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel.md", "text": "https://wpnews.pro/news/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel.txt", "jsonld": "https://wpnews.pro/news/decay-gated-o-n-causal-linear-attention-with-fused-triton-kernel.jsonld"}}