i’ve added a 100GB of swap space as a hail-mary, i hope this lets me get past the transient spikes
I got it running !!!, benches will follow, stupid dummy load tester blew up hostram causing OOM’s not seen with the actual effing model
ok so i’m loaded to 82GB of host RAM and around 110GB of VRAM with a GPU kv cache of 500k,
it runs but theres some weirdness:
100 tk/s in concurrent which is …..fine and might be expected considering n-gram streaming from ram, however:
Single user drops to a paltry 4 tk/s, investigating further
turns out i had some debug flags enabled that caused the graphs to get rebuilt per token, I’m at 27.5 tg/s single and 300 tg/s concurrent now
Have you taken a look at tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8-GPTQ on huggingface? The big strength of the R9700 is FP8 matrix multiplication so you want to avoid W4A16, which is best for older Nvidia/AMD hardware.