Why self-hosted inference is essential
Red Hat AI warns that enterprises relying on third-party hosted APIs for AI agent inference undermine their own data sovereignty, as every prompt and tool call routes through external datacenters. The…
Red Hat AI warns that enterprises relying on third-party hosted APIs for AI agent inference undermine their own data sovereignty, as every prompt and tool call routes through external datacenters. The…
LocalAI, a self-hosted runtime that runs AI workloads on user-controlled hardware, offers an OpenAI-compatible API supporting text generation, vision, speech, image and video generation, embeddings, a…
RunInfra enables running Kimi-Linear-48B, a distilled version of the full 2.78-trillion-parameter Kimi K3 model, on a single consumer GPU such as the RTX 5090 with 32 GB VRAM, achieving 113.83 tokens …
The accuracy gap between the best open-weight models and GPT-4o has shrunk to under 3% on structured data tasks, according to a developer's hands-on analysis. A fine-tuned Qwen2.5 72B model achieved 9…
A new Chrome extension called Explain This lets users select text on a page, right-click, and get a plain-language explanation from a local LLM running entirely in the browser via WebGPU, with no serv…
DeepSeek V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters, scores 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and achieving the highest score for any open-weight model. Pric…
A developer built a pure-JAX inference engine for Google's Gemma 4 E2B model after discovering that quantization-aware-trained (QAT) checkpoints fail to load in vLLM on TPU due to a loader bug. The en…
Intel's Arc Pro B70 GPU crashes under sustained inference load in production, according to ModDog bot developer who pulled the card from his fleet after six weeks of testing. The 32GB card, priced at …
Microsoft's Azure Kubernetes Service engineering team published a reference implementation on June 29 that separates agent-request routing into three layers: semantic model selection via RouteLLM, gat…
Microsoft released a reference architecture for routing agent traffic on Azure Kubernetes Service that combines the Kubernetes Gateway API Inference Extension, agentgateway, and RouteLLM into one Open…
A developer built a small scheduler in Go to understand vLLM's scheduler for LLM inference, tracing each design decision back to its vLLM equivalent. The scheduler operates in a tick loop with three p…
VLLM's slack-aware preemption policies rescued urgent request attainment at high load in a benchmark on 6× A100 SXM4 80GB GPUs running Llama 70B, where the control policy collapsed to 0% urgent attain…
A new paper, FlowPrefill by Hsieh et al., proposes preempting long LLM inference prefills mid-forward-pass to rescue urgent requests that would otherwise miss their time-to-first-token (TTFT) service-…
LLM inference profitability depends on the trade-off between batch size and GPU count, which determines token latency and cost per million tokens. Applying this model to Kimi K3, which requires at lea…
Moonshot AI released Kimi K3, a 2.78-trillion-parameter mixture-of-experts model with 104.2 billion active parameters, native vision, and a 1,048,576-token context window, claiming it is the first ope…
Together AI introduces ThunderAgent, a system for high-throughput agentic inference that achieves up to 2.5× higher single-node throughput and 2.4× speedup on an 8-node cluster with near-linear scalin…
AgentHound, an open-source offensive security framework for AI agent infrastructure, has been released. The tool performs recon, fingerprinting, credential looting, model inversion, and persistence ac…
Moonshot AI's Kimi Linear attention architecture, introduced in October, now powers the company's 2.8-trillion-parameter Kimi K3 flagship model released in mid-July, marking the first production deplo…
A new hands-on guide teaches programmers how to run large language models on their own hardware, covering tools like Ollama, llama.cpp, vLLM, and Hugging Face's transformers. The guide explains local …
Flash Attention is an algorithm that speeds up training and inference of transformer models by using smart memory management on GPUs. The original version was released in 2022, followed by Flash Atten…