Hey Hugging Face Community,
I’ve engineered and open-sourced a Decay-Gated O(N) Causal Linear Attention architecture featuring fused Triton/CUDA state accumulation kernels designed to bypass quadratic MHA bottlenecks.
Key Specs & Benchmark Metrics:
Repository & Research Paper:
Would love to gather thoughts on integration into Hugging Face Transformers model classes or benchmark comparisons with existing linear models!