How I Fixed vLLM on Strix Halo and Got 3x Better Batch Throughput with Qwen3.5 Continue reading on Towards AI »
source & further reading
pub.towardsai.net — original article
Why You Can’t Give an LLM Direct Write-Access to Your EHR
GPT-5.6 Quietly Broke Our Prompt Cache. One Message Boundary Fixed It.
How to Cut Your AI Agent’s Context Cost With GPT‑6 Prompt Caching