cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 5/38 feed RSS

// recent coverage 743 mentions

00:00
2026-09-20
signoz.io
ai-infrastructure

vLLM Dashboard: Monitor Token Throughput and Latency

SigNoz published a vLLM monitoring dashboard that tracks token throughput, latency, and KV cache usage for self-hosted vLLM inference servers, requiring SigNoz v0.135.0 or newer and the V2 dashboard s…

00:00
2026-09-20
signoz.io
ai-infrastructure

vLLM Monitoring and Observability with OpenTelemetry

SigNoz published a guide for monitoring self-hosted vLLM inference servers by sending Prometheus metrics and OpenTelemetry request traces to SigNoz. The guide has operators scrape vLLM's /metrics endp…

13:02
2026-09-19
github.com
ai-tools

A local Jev backed by DiffusionGemma

LocalJev, a TypeScript server for Bun 1.2+, implements a Jev-compatible POST /v1/systemone API backed by the DiffusionGemma model diffusiongemma-26B-A4B-it-4bit through an OpenAI-compatible Chat Compl…

02:16
2026-09-19
byteiota.com
ai-agents

Abacus.AI Smaug: Open-Weight Agent Models at 1/100th the Cost

Abacus.AI launched the Smaug line of three open-weight models on September 10, fine-tuned for long-running agentic loops and deployable inside a customer's own VPC at open-source rates. Abacus.AI clai…

19:19
2026-09-18
forum.level1techs.com
ai-infrastructure

Budget Pcie 5.0 16x, 8 Card Baseboard Build

A user who purchased 8 R9700 GPUs is asking for advice on the cheapest way to build a PCIe 5.0 x16 tensor-parallel rig for running TP8 inference on ~200GB models, proposing a single Microchip Switchte…

19:04
2026-09-18
developer.nvidia.com
large-language-models

Benchmarking LLM Inference at Scale with AIPerf

NVIDIA released AIPerf, a ground-up rewrite and designated successor to GenAI-Perf, as a multiprocessed LLM inference benchmarking tool that coordinates worker processes and record-processor services …

21:36
2026-09-17
dev.to
ai-infrastructure

Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

A developer published a step-by-step guide and open-source MCP toolkit for deploying Google's Gemma 4 E2B model to a single AMD Instinct MI300X GPU rented through AMD Developer Cloud at $1.99 per hour…

01:08
2026-09-17
dev.to
ai-tools

When Developers Should NOT Use AI

A developer argues that AI coding assistants should sit at the end of the engineering toolchain, used only when deterministic tools like compilers, tests, debuggers, profilers, and Git cannot answer a…

23:21
2026-09-16
github.com
large-language-models

vLLM: Jev-like mode for the DiffusionGemma model

VLLM contributor mmastrac opened a pull request adding a "Jev-like" structured generation mode for the DiffusionGemma model, retitling it from a work-in-progress draft to "[Core] structured generation…

23:16
2026-09-16
frontierroles.com
ai-research

Research, Tinker, RL Systems — Thinking Machines Lab

Thinking Machines Lab is hiring a Research, Tinker, RL Systems engineer in San Francisco at an annual salary range of $350,000 to $475,000, a figure the job board states sits 73% above the $238,000 me…

← prev page 5 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics