# I didn't know it was IMPOSSIBLE!... I just needed it." - 400k+ Context on Qwen 3.6 MoE 35B on a Single 16GB GPU (RTX 5060 Ti) at 25 t/s

> Source: <https://forum.level1techs.com/t/i-didnt-know-it-was-impossible-i-just-needed-it-400k-context-on-qwen-3-6-moe-35b-on-a-single-16gb-gpu-rtx-5060-ti-at-25-t-s/255275#post_5>
> Published: 2026-10-10 10:38:17+00:00

I can’t imagine this is useful in a autonomous setting. Can it build something accurately? I use 16GB VRAM andsometimes push the usage to 15.2GB in a single model, the weights aren’t that strong, there is more failures over 64k tokens when VRAM in at 90%+ than if I use a model better sized for the card.

I use [Qwen3.8-27B-UD-Q2_K_XL](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) and it cannot surpass 128k without a OOM with llama.cpp. ~13.2GB before context.

For autonomous verfied responses(as in something else checks the work) I can’t trust more than 64K, over that and the quant weights are iffy at best. The repeat loops are funny sometimes, but the quant context pressure seems varied per model weights.

I’d love to see the 16GB card(qwen3.6-35b-a3b-12gb-2.6763bpw.gguf) run a software buildout with a follow up cloud model to judge and grade it. That’s my method of validating local models currently.
