19:09
2026-08-23
forum.level1techs.com
large-language-models
Sanitfy check my Qwen3.8 5090 Results; new to Local AI
A user running Qwen3.8-27B on an NVIDIA RTX 5090 32 GB with LM Studio reports decode speeds ranging from 12.38 to 47.4 tok/s depending on quantization and multi-token prediction settings, with the Q4_โฆ