cd /news/large-language-models/self-hosting-llms-my-journey-so-far-… · home topics large-language-models article
[ARTICLE · art-110735] src=forum.level1techs.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Self-hosting LLMs, my journey so far, and knowledge desired.

A user self-hosting large language models on an AMD Radeon RX 9700 reports achieving only ~22 tokens/s with Qwen 3.8 Q5_K_m, despite fitting 132k Q8_0 context, and seeks tips for improving performance and learning resources for entry-level local AI. The user also found that a 275KB text file became too large once embedded, causing the model to fail on a media organization task.

read1 min views1 publishedAug 25, 2026

I picked up an r9700 because everyone in my area wants too much for their used 7900XTX, and I’d rather have a warranty at that point, +8GB vram, so I’m in a similar / adjacent boat. Qwen 3.8 Q5_k_m fits comfortably with 132k Q8_0 context with room to spare but I only get ~22 tokens/s though with Ollama+openwebUI. R9700 ~600GB/s memory makes it not the fastest… The qwen3.6 MoE model gets ~triple that response speed though IIRC. I’m noob so lots to learn. llama.cpp can supposedly provide speed boosts by enabling MTP, but I’ve read mixed opinions on MTP.

I’m also hoping to find some good learning resources for this sort of entry-level local AI, and what sort of tasks it’s best suited for. I thought I’d have it organize a media collection on my filesystem, so I gave it a tree of the directory, a 275Kb txt file. Little did I know that such a tiny file becomes large once its embedded… it failed to work with that much data…

If anyone has questions about r9700 or tips and model recommendations for us, much appreciated.

── more in #large-language-models 4 stories · sorted by recency
── more on @amd radeon rx 9700 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/self-hosting-llms-my…] indexed:0 read:1min 2026-08-25 ·