{"slug": "self-hosting-llms-my-journey-so-far-and-knowledge-desired", "title": "Self-hosting LLMs, my journey so far, and knowledge desired.", "summary": "A user self-hosting large language models on an AMD Radeon RX 9700 reports achieving only ~22 tokens/s with Qwen 3.8 Q5_K_m, despite fitting 132k Q8_0 context, and seeks tips for improving performance and learning resources for entry-level local AI. The user also found that a 275KB text file became too large once embedded, causing the model to fail on a media organization task.", "body_md": "I picked up an r9700 because everyone in my area wants too much for their used 7900XTX, and I’d rather have a warranty at that point, +8GB vram, so I’m in a similar / adjacent boat. Qwen 3.8 Q5_k_m fits comfortably with 132k Q8_0 context with room to spare but I only get ~22 tokens/s though with Ollama+openwebUI. R9700 ~600GB/s memory makes it not the fastest… The qwen3.6 MoE model gets ~triple that response speed though IIRC. I’m noob so lots to learn. llama.cpp can supposedly provide speed boosts by enabling MTP, but I’ve read mixed opinions on MTP.\n\nI’m also hoping to find some good learning resources for this sort of entry-level local AI, and what sort of tasks it’s best suited for. I thought I’d have it organize a media collection on my filesystem, so I gave it a tree of the directory, a 275Kb txt file. Little did I know that such a tiny file becomes large once its embedded… it failed to work with that much data…\n\nIf anyone has questions about r9700 or tips and model recommendations for us, much appreciated.", "url": "https://wpnews.pro/news/self-hosting-llms-my-journey-so-far-and-knowledge-desired", "canonical_source": "https://forum.level1techs.com/t/self-hosting-llms-my-journey-so-far-and-knowledge-desired/254361#post_7", "published_at": "2026-08-25 18:44:57+00:00", "updated_at": "2026-08-25 19:14:44.754406+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["AMD Radeon RX 9700", "Qwen 3.8", "Ollama", "Open WebUI", "llama.cpp", "MTP"], "alternates": {"html": "https://wpnews.pro/news/self-hosting-llms-my-journey-so-far-and-knowledge-desired", "markdown": "https://wpnews.pro/news/self-hosting-llms-my-journey-so-far-and-knowledge-desired.md", "text": "https://wpnews.pro/news/self-hosting-llms-my-journey-so-far-and-knowledge-desired.txt", "jsonld": "https://wpnews.pro/news/self-hosting-llms-my-journey-so-far-and-knowledge-desired.jsonld"}}