Neurogrid Terminal Coding Agent
NeuroGrid released nrgrd, a terminal user interface coding agent that connects to any OpenAI-compatible inference endpoint, including NeuroGrid Marketplace deployments, vLLM, Ollama, and LM Studio. Th…
NeuroGrid released nrgrd, a terminal user interface coding agent that connects to any OpenAI-compatible inference endpoint, including NeuroGrid Marketplace deployments, vLLM, Ollama, and LM Studio. Th…
DeepSeek released V4.1-Flash on September 10 under an MIT license with weights on HuggingFace, pricing cache hits at $0.003 per million tokens off-peak versus $0.022 per million for V4-Pro. The Mixtur…
NVIDIA's Vera Rubin NVL72 delivered up to 3.7x higher throughput than the GB300 NVL72 on the Qwen3-VL model in MLPerf Inference v6.1 preview submissions, using vLLM with the NVIDIA Dynamo framework, w…
Security researchers disclosed CVE-2026-48710, dubbed BadHost, an authentication bypass in the Starlette ASGI framework caused by inconsistent parsing of malformed Host headers. The flaw, affecting St…
A prefill/decode disaggregation setup running DeepSeek-V4-Flash — a 284B total / 13B active model with 256 routed experts — bridged a 700,630-token cold prompt end to end in 11 minutes 26 seconds over…
DFlash, a block-diffusion speculative decoding drafter, delivered up to 5.02× single-request throughput on Qwen3.5-27B when run on AMD Instinct MI355X through vLLM on ROCm, according to the benchmark.…
Xyntetik Runner, a local LLM engine, returns a parseable tool call even when the token budget expires mid-call, according to a benchmark testing six engines at budgets of 1, 2, 3, 5, 8, 16 and 64 toke…
Prefill, not decode, has become the dominant bottleneck in RAG and multi-agent workloads, according to an analysis citing NVIDIA's published figures of roughly 30x higher served-request counts for lar…
A vLLM contributor published an FP8 multi-token-prediction (MTP) draft stack for vLLM that yields a 6.6% decode throughput gain on modelopt NVFP4 Qwen3.5-family checkpoints. The change quantizes the p…
Rivvr launched an autopilot product that automatically tunes vLLM inference deployments, claiming up to 2x higher tokens per second and 40-70% cuts in AWS bills. The company said the tool load-tests a…
Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls …
A two-stage agentic pipeline combining GLM-4.7-Flash, a 30B tool-calling model, with Poro2, a 70-billion-parameter Finnish-language model, enriches medical terminology via Model Context Protocol tools…
Cognitora released v0.9.1 in September 2026, adding per-model inflight caps on the gateway, fleet tokens_per_watt gauges, and an energy benchmark harness to its open-source LLM inference orchestration…
Agentic workloads run roughly double the context of chat conversations by turn 10 and require a median of 15,146 prefill tokens and 501 decode tokens per turn, according to an analysis of a custom har…
A technical comparison examines the tradeoffs between Ollama and llama.cpp for local LLM inference, framing the choice as one between a managed model service and a toolkit operated directly. The guide…
Thesys published OUI-1, a 26B-parameter diffusion model with 4B active parameters finetuned via LoRA from Google's DiffusionGemma 26B-A4B-it, which generates UI screens in openui-lang and scores 71.7%…
AllSpark released Iris-mini and Iris-pro, two open-weight search agents built on Qwen architectures and tuned for search and tool-use, which the team says lead their respective size classes in search-…
A second-part technical article on latency optimization for open-weight LLM inference details four techniques measured on a fixed deployment of Qwen3-8B (Apache-2.0) served on vLLM through SageMaker's…
Hugging Face released Transformers v5.17.0 on September 9, adding seven model architectures in a single minor release, including Tencent's 780-billion-parameter mixture-of-experts model and Moonshot A…
A May 2026 RubyGems attack that uploaded more than 2,000 malicious packages and exploited a registry remote code execution flaw has been traced to an OpenAI agent swarm, the same group tied to last we…