Architecting Secure Prompt Caching
Tinfoil announces cached prompt pricing in its Inference API, a feature that reduces compute for eligible requests by caching recently processed inputs, but the company warns that caching introduces t…
Tinfoil announces cached prompt pricing in its Inference API, a feature that reduces compute for eligible requests by caching recently processed inputs, but the company warns that caching introduces t…
Moonshot AI released Kimi K3, a 2.8-trillion-parameter open Mixture-of-Experts model with native vision and a 1-million-token context window, claiming it as the world's first open 3T-class model. The …
VLLM v0.25.0, released July 11, deletes the original PagedAttention implementation and makes Model Runner V2 the default execution backend for all dense models, delivering a 56% throughput improvement…
NVIDIA released the Nemotron 3 Embed collection of open embedding models, led by an 8B model that ranks #1 on the RTEB leaderboard with a score of 78.5%, alongside efficient 1B variants for production…
A five-engineer team on Sonnet 4.6 saw a $4,800 monthly bill for Claude Code sessions, with only $960 attributed to the model itself and $3,840 coming from inference engineering — the layer between th…
Veta, an AI agent that QA-tests Android apps using a swarm of autonomous sub-agents, runs 100% of its AI inference on AMD GPUs via Fireworks AI on AMD Instinct or self-hosted vLLM on ROCm. The system …
Two open-weight coding models, Kimi K2.7 Code from Moonshot AI and GLM-5.2 from Z.ai (Zhipu AI), launched days apart in June 2026, both designed for agentic coding workflows and supporting vLLM and SG…
Researchers present Fleet, a hierarchical task-based abstraction for multi-die GPUs that introduces Chiplet-tasks to exploit chiplet-level locality and synchronization. On AMD Instinct MI350 with Qwen…
Thinking Machines Lab, Inc. released Inkling, a general-purpose multimodal model with 975 billion total parameters and 41 billion active parameters, on July 15, 2026 under an Apache 2.0 license. The m…
Seven Python frameworks for orchestrating local AI agents are gaining adoption in 2026, according to a technical roundup. Ollama, a lightweight runtime for running open-source LLMs on local hardware, …
Version 1.21.0 of RFC BF introduces a YAML-based providers block that replaces hardcoded LLM providers with a config-declared registry supporting pluggable drivers for Anthropic, OpenAI, Gemini, DeepS…
A developer released AI-CLI, a tiny C terminal assistant that connects user requests to a local LLM and executes returned shell actions directly, supporting over 20 platforms and most LLM engines. The…
NVIDIA released Nemotron-Labs-TwoTower on July 1, achieving 2.42 times faster inference throughput at 98.7% of the baseline model's benchmark quality by adding a second neural network tower trained on…
A user inquired about deploying a multi-GPU environment on Hugging Face Spaces, requesting 8 GPUs with over 40GB VRAM each to run a generative model on 4 GPUs and a vLLM model on the remaining 4 GPUs,…
XGrammar, a grammar-constrained decoding library used by vLLM, SGLang, TensorRT-LLM, and MLC-LLM, masks invalid tokens before the model samples, guaranteeing valid JSON output and eliminating retry lo…
SGLang, an open-source LLM inference engine, introduces RadixAttention for prefix caching and grammar-constrained decoding, achieving 25x throughput improvement over vLLM for structured output tasks. …
Mistral released Leanstral 1.5, a formal verification agent built on Lean 4, on July 2, and in its first public test against 57 open-source repositories it found five bugs that human maintainers had n…
Atuin has open-sourced the Atuin AI server, allowing users to self-host the terminal-focused AI agent that provides agentic tools directly in the shell. The server supports any OpenAI-compatible endpo…
Intel's Arc Pro B70, priced at $949 for 32GB of GDDR6 but now selling above $1,100 due to market shortages, offers up to 85% higher token throughput and 6.2x faster time-to-first-token versus NVIDIA's…
A developer released trollbridge, a policy-enforced HTTP proxy designed to let coding agents operate freely while blocking unsafe requests. The tool uses deterministic allow/deny lists for most traffi…