04:44
2026-09-26
octet-stream.net
large-language-models
Getting Deeper into Local Inference
A personal LLM user reports that Qwen3.8-27B running locally via llama.cpp on an RTX 5090 laptop GPU with 24 GB of VRAM now handles Q&A, coding, sysadmin and web research without a cloud provider, delβ¦