I think the best model you can realistically run at good enough speed for agentic stuff is Qwen3.8 27b. So the goal here is running that at the best quant + context you can muster.
You didn’t mention how many PCIe slots you have but for the purpose of this recommendation they need to run at at least 8x and I’m going to assume you have two or three. I’m also assuming your PSU can handle whatever you buy or isn’t part of the budget.
For the simplest path here I would say get a used 20gb 3080. It will pair well with your existing card and you should be able to run a Q6-ish quant of Qwen3.8 27b with decent context. IMO this is your best option. Someone mentioned a 3090 but the extra 4gb isn’t worth the premium at the moment. If you want to stretch your budget like this I would instead try selling your current 10gb 3080 and buying a pair of 20gb 3080s. Or maybe even a pair of V100s.
Also mentioned a 32gb V100 is also a reasonable choice but it’s not going to pair super well with your 3080 and being out of support you’re committing to some tinkering and running linux. On the other hand it’s the most VRAM you’re going to get for $700 and even if it bottlenecks your 3080 a bit we’re now talking Q8 full context Qwen3.8 27b.
Sidenote
Never quantize KV cache, it isn’t worth the degradation. Either use less context with stuff like compaction or go to a lower quant model.