The Local LLM Matrix: Best Models & Quants by VRAM Tier (<=16GB – 256GB+)
A forum thread is crowdsourcing real-world local LLM deployment data across five VRAM tiers — ≤16GB, 24–32GB, 48–64GB, 96–128GB, and 196–256GB+ — asking users to report model, quantization (AutoRound …