Performant C/CUDA inference engine for Qwen 3.6 35B on RTX 5090 / Blackwell
A hyper-optimized, zero-dependency C/CUDA inference engine for the Qwen 3.6 35B model on RTX 5090 Blackwell GPUs achieves 13.4k tokens/sec prefill throughput at 2,048 context depth and 270+ tokens/sec decode, outperformi…