07:58
2026-08-24
forum.level1techs.com
large-language-models
Sanity check my Qwen3.8 5090 Results; new to Local AI
A user running Qwen3.8-27B on an RTX 5090 with the ninfer runtime reported a 1.72ร speedup over LM Studio/llama.cpp, achieving 124.08 tok/s decode and 8,143.82 tok/s prompt eval with mixed NVFP4/FP8 qโฆ