I purchased a EVO-X3 and I used Claude to replace the calls to the API with calls to this new box. From my website. I started with Windows and got it to work. Using the Claude responses as a way to figure out which models would even pass an acceptance test, then focused on performance. I got Claude to setup Linux which I was surprised how smoothly it went. I have settled on Qwen3.8 Flash-Nect. Did need Claude to work for almost a day to get it to work I use a smaller model for the quick response I need. It does work. Where it does not work so well when lots of request come in. Can take a long time. Interested in adding a GPU card to run the smaller model faster or allow more instances to run. Is this a good solution?
I should have said the larger model does not have this issue only the smaller model. This is ment to be fast and many users using. The one I am usingg is the Qwen3.6 27B 3A expert one. May have got the name wrong.