{"slug": "alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy", "title": "ALHR-A tree based sparse attention system that retains long context accuracy", "summary": "A developer's ALHR (Adaptive Learnable Hierarchical Routing) sparse attention system, which uses static binary trees and learnable functions to reduce keys read per query, achieved 92.1% top-1 accuracy on the MQAR test at 1024 tokens versus 94.9% for a dense baseline, while reading an average of 30 keys per query against the dense model's 512. The system reports 35.3x KV compression (2.83% read) and 100% cache compression, with peak VRAM of 422 MB scaling linearly versus the dense model's 57 MB scaling quadratically, though the author notes full-scale tests are incomplete and training remains quadratic while inference would be NlogN.", "body_md": "ALHR - Adaptive Learnable Hierarchical Routing, uses static binary trees and learnable functions to minimize the amount of keys to be read. It does use a dense teacher while phase 1 of training however. MQAR TEST AT 1024 TOKENS - Average keys read per query by dense - 512 Keys Average keys read per query by ALHR - 30 keys Top - 1 accuracy of dense - 94.9% Top - 1 accuracy of ALHR - 92.1% KV Compression of dense - 1x(100% read) KV Compression of ALHR - 35.3x(2.83% read) Peak VRAM of dense - 57 MB (Scales quadratically) Peak VRAM of ALHR - 422 MB (scales linearly) Cache compression of ALHR - 100% The true log and Kaggle cell used to run it are in the logs folder in the repo limitations: Full scale tests are still not completed, The training of this model would still be quadratic but the inference would be NlogN (as indicated in the logs in the repo) Would love your opinions\n\nALHR Repository: repo link in comments\n\nComments URL: [https://news.ycombinator.com/item?id=50020238](https://news.ycombinator.com/item?id=50020238)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy", "canonical_source": "https://news.ycombinator.com/item?id=50020238", "published_at": "2026-10-09 13:27:33+00:00", "updated_at": "2026-10-09 13:54:49.464356+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["ALHR", "Adaptive Learnable Hierarchical Routing", "MQAR", "Kaggle"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy", "markdown": "https://wpnews.pro/news/alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy.md", "text": "https://wpnews.pro/news/alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy.txt", "jsonld": "https://wpnews.pro/news/alhr-a-tree-based-sparse-attention-system-that-retains-long-context-accuracy.jsonld"}}