Nice you have two RVII’s?
For reference: SummaryI too have finetuned a llama.cpp or two, but for one card, we mashed flashattention into a variant in jan/feb then one of these beat me to it, ended up just using it.
I actively use knguyen298/llama-swap-gfx906 - Docker Image They have built router features into newer llama.cpp, but my harness has the extra model field in API calls and that is easier for me at least. I use ai-infos vLLM for that container.
I’ll spin up this variant later on, the RVII is crunching tokens right now