cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 27/38 feed RSS

// recent coverage 748 mentions

00:51
2026-07-17
tinfoil.sh
ai-infrastructure

Architecting Secure Prompt Caching

Tinfoil announces cached prompt pricing in its Inference API, a feature that reduces compute for eligible requests by caching recently processed inputs, but the company warns that caching introduces t…

16:11
2026-07-16
byteiota.com
artificial-intelligence

vLLM v0.25: Model Runner V2 Default, PagedAttention Gone

VLLM v0.25.0, released July 11, deletes the original PagedAttention implementation and makes Model Runner V2 the default execution backend for all dense models, delivering a 56% throughput improvement…

14:03
2026-07-16
github.com
artificial-intelligence

Veta: AI agent that QA-tests Android apps

Veta, an AI agent that QA-tests Android apps using a swarm of autonomous sub-agents, runs 100% of its AI inference on AMD GPUs via Fireworks AI on AMD Instinct or self-hosted vLLM on ROCm. The system …

21:12
2026-07-15
thinkingmachines.ai
artificial-intelligence

Inkling Model Card

Thinking Machines Lab, Inc. released Inkling, a general-purpose multimodal model with 975 billion total parameters and 41 billion active parameters, on July 15, 2026 under an Apache 2.0 license. The m…

12:00
2026-07-15
kdnuggets.com
artificial-intelligence

7 Python Frameworks for Orchestrating Local AI Agents

Seven Python frameworks for orchestrating local AI agents are gaining adoption in 2026, according to a technical roundup. Ollama, a lightweight runtime for running open-source LLMs on local hardware, …

10:09
2026-07-15
byteiota.com
large-language-models

NVIDIA Nemotron TwoTower: 2.42x Faster LLM Inference

NVIDIA released Nemotron-Labs-TwoTower on July 1, achieving 2.42 times faster inference throughput at 98.7% of the baseline model's benchmark quality by adding a second neural network tower trained on…

00:00
2026-07-15
dibi8.com
artificial-intelligence

SGLang — Structured Generation and Fast LLM Serving Engine

SGLang, an open-source LLM inference engine, introduces RadixAttention for prefix caching and grammar-constrained decoding, achieving 25x throughput improvement over vLLM for structured output tasks. …

20:10
2026-07-14
byteiota.com
artificial-intelligence

Leanstral 1.5: Mistral’s AI Found Five Real Bugs

Mistral released Leanstral 1.5, a formal verification agent built on Lean 4, on July 2, and in its first public test against 57 open-source repositories it found five bugs that human maintainers had n…

18:21
2026-07-14
blog.atuin.sh
ai-tools

Open Sourcing the Atuin AI Server

Atuin has open-sourced the Atuin AI server, allowing users to self-host the terminal-focused AI agent that provides agentic tools directly in the shell. The server supports any OpenAI-compatible endpo…

← prev page 27 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics