cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 4/38 feed RSS

// recent coverage 743 mentions

08:10
2026-09-22
dev.to
large-language-models

Fine-Tune, Deploy and Use LLM As AI Agent

A developer published a video walkthrough demonstrating an end-to-end pipeline for fine-tuning a large language model and deploying it as an AI agent, using Runpod for GPU rental and serverless infere…

01:45
2026-09-22
github.com
ai-tools

I built an autonomous accounting tool to let AI do my taxes

A developer released Autonomous Accounting, an open-source tool that uses a user-selected vision LLM to convert receipts, invoices and bank statements into a reconciled, categorized ledger plus a down…

00:00
2026-09-22
rocm.blogs.amd.com
large-language-models

Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X

Moonshot AI released the weights for Kimi-K3, a 2.8-trillion-parameter, 1M-context, natively-MXFP4 Mixture-of-Experts model, with AMD Instinct support available on day 0 across three serving framework…

20:59
2026-09-21
g-ftech.com
large-language-models

vLLM Architecture, Memory and Benchmarks Deep Dive

VLLM's PagedAttention and continuous iteration-level batching address the KV cache memory bottleneck that limits LLM inference throughput, according to a technical deep dive on the inference engine's …

20:07
2026-09-21
paradigma.inc
large-language-models

94% on AIME with 1B Params

Paradigma released Limite 1B - Violetto, a 1-billion parameter dense autoregressive transformer trained from scratch on fewer than 300B curated tokens that averages 74.25% on BeyondAIME, ahead of MUSE…

18:33
2026-09-21
dev.to
large-language-models

High-Throughput LLM Inference & Training: A Deep Dive into vLLM

An engineer at g factor detailed how vLLM's PagedAttention and continuous iteration-level batching solve the memory-bandwidth bottleneck in production LLM inference, drawing on benchmarks run on dedic…

18:14
2026-09-21
newsletter.semianalysis.com
ai-infrastructure

Computation and Data Movement for Inference

Mixture of Experts has changed the structure of AI inference serving by altering which tensors are active per token, what must remain close together, which transfers need strong local bandwidth, and h…

13:05
2026-09-21
promptcube3.com
large-language-models

Jev is a total shift in how we use LLMs for automation

Former OpenAI researcher Diogo Almeida launched Jev, a "System One" model that outputs only probability judgments rather than generated text, and developers have built nearly 500 open-source projects …

00:00
2026-09-21
simm.is
ai-infrastructure

A KV cache you can fork

Replikativ released pretrained-rstr, an MIT-licensed inference component that turns a language model's KV cache into a forkable, content-addressed value rather than a process-local optimization. The s…

16:56
2026-09-20
dev.to
ai-infrastructure

How I Debugged a KV-Cache Offloading Bug in vLLM

A developer identified and fixed a KV-cache offloading bug in vLLM that caused incorrect chunking for models with mixed KV-cache groups. The existing implementation assumed a single KV-cache group lay…

04:07
2026-09-20
arxiv.org
ai-safety

Inference-Engine Fingerprinting Attacks Are Practical

A September 17, 2026 arXiv paper shows that a misaligned AI model can fingerprint which inference engine executes it — including vLLM and SGLang — and then use engine-specific exploits to seize contro…

← prev page 4 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics