17:18
2026-08-18
promptcube3.com
artificial-intelligence
Four RTX 3060s can actually push 100 tok/s prompt processing on
A developer reports achieving 99.4 tok/s prompt processing and 10.1 tok/s text generation with DeepSeek-V4-Flash-0731 (UD-Q4_K_XL GGUF) on four NVIDIA RTX 3060 12GB cards using llama.cpp build b10181,โฆ