I picked up an r9700 because everyone in my area wants too much for their used 7900XTX, and I’d rather have a warranty at that point, +8GB vram, so I’m in a similar / adjacent boat. Qwen 3.8 Q5_k_m fits comfortably with 132k Q8_0 context with room to spare but I only get ~22 tokens/s though with Ollama+openwebUI. R9700 ~600GB/s memory makes it not the fastest… The qwen3.6 MoE model gets ~triple that response speed though IIRC. I’m noob so lots to learn. llama.cpp can supposedly provide speed boosts by enabling MTP, but I’ve read mixed opinions on MTP.
I’m also hoping to find some good learning resources for this sort of entry-level local AI, and what sort of tasks it’s best suited for. I thought I’d have it organize a media collection on my filesystem, so I gave it a tree of the directory, a 275Kb txt file. Little did I know that such a tiny file becomes large once its embedded… it failed to work with that much data…
If anyone has questions about r9700 or tips and model recommendations for us, much appreciated.