cd /news/large-language-models/i-didn-t-know-it-was-impossible-i-ju… · home › topics › large-language-models › article
[ARTICLE · art-148702] src=forum.level1techs.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

I didn't know it was IMPOSSIBLE!... I just needed it." - 400k+ Context on Qwen 3.6 MoE 35B on a Single 16GB GPU (RTX 5060 Ti) at 25 t/s

A user reported running a Qwen 3.6 MoE 35B model at roughly 2.6763 bits per weight (qwen3.6-35b-a3b-12gb-2.6763bpw.gguf) on a single 16GB RTX 5060 Ti at 25 tokens per second with over 400k context, while noting that reliability degrades past 64k tokens when VRAM exceeds 90%. The same user said their Qwen3.8-27B-UD-Q2_K_XL GGUF model uses about 13.2GB before context and hits out-of-memory in llama.cpp beyond 128k tokens, and proposed validating local models by having a cloud model grade a software buildout.

read1 min views1 publishedOct 10, 2026

I can’t imagine this is useful in a autonomous setting. Can it build something accurately? I use 16GB VRAM andsometimes push the usage to 15.2GB in a single model, the weights aren’t that strong, there is more failures over 64k tokens when VRAM in at 90%+ than if I use a model better sized for the card.

I use [Qwen3.8-27B-UD-Q2_K_XL](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) and it cannot surpass 128k without a OOM with llama.cpp. ~13.2GB before context.

For autonomous verfied responses(as in something else checks the work) I can’t trust more than 64K, over that and the quant weights are iffy at best. The repeat loops are funny sometimes, but the quant context pressure seems varied per model weights.

I’d love to see the 16GB card(qwen3.6-35b-a3b-12gb-2.6763bpw.gguf) run a software buildout with a follow up cloud model to judge and grade it. That’s my method of validating local models currently.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen 3.6 moe 35b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-didn-t-know-it-was…] indexed:0 read:1min 2026-10-10 · —