cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 13/23 feed RSS

// recent coverage 443 mentions

14:13
2026-07-13
byteiota.com
artificial-intelligence

Together AI Raises $800M: Open-Source Inference Just Got Serious

Together AI closed an $800 million Series C at an $8.3 billion valuation, reporting $1.15 billion in annual bookings and an inference engine that hits 500 tokens per second on DeepSeek-V3.1. The compa…

06:23
2026-07-13
machinebrief.com
large-language-models

Boosting Language Model Efficiency with AugServe

AugServe, a new framework for optimizing large language model inference, achieves a 4.7x increase in effective throughput compared to vLLM and a 3.3x boost over InferCept, while reducing time-to-first…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators

AMD has integrated an NVFP4 emulation pipeline into vLLM that enables AMD Instinct MI355 accelerators to serve standard NVFP4 quantized checkpoints directly, dequantizing weights to BF16 on-the-fly at…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

QuickReduce INT3 Quantization and Benchmarking on MI355

AMD's QuickReduce library now supports INT3 quantization for all-reduce communication in multi-GPU LLM inference, achieving a 22% reduction in on-wire data volume compared to INT4 on AMD Instinct MI35…

23:00
2026-07-12
github.com
artificial-intelligence

Argocd-AI-Assistant

Said Sef's Argocd-AI-Assistant is an Argo CD UI extension that adds an AI-powered Assistant tab to Kubernetes resource views, enabling natural language queries enriched with manifest, events, and opti…

05:51
2026-07-12
gist.github.com
artificial-intelligence

vllm locally on 5060Ti 16GB x 2

A developer deployed vLLM locally on two NVIDIA RTX 5060 Ti 16GB GPUs using Docker Compose, configuring tensor parallelism, FP8 KV cache, and speculative decoding with MTP. The setup runs an OpenAI-co…

19:12
2026-07-11
byteiota.com
artificial-intelligence

Hugging Face Kernels Are Now Signed Hub Artifacts

Hugging Face announced that custom GPU kernels on its Hub are now signed artifacts governed by a trusted publisher model, requiring a dedicated repository type that replaces the old model-type format.…

09:39
2026-07-11
machinebrief.com
large-language-models

LLM Serving: A Smarter Approach to Prefill and Decode

A new scheduler for large language model serving allows decode nodes to assist prefill phases, cutting P95 time-to-first-token by up to 81% and improving service-level objective attainment by up to 79…

09:28
2026-07-11
pub.towardsai.net
artificial-intelligence

Can AI Model Vendors Track Your Self-Hosted Deployment?

A developer investigating FLUX.1 [dev] for commercial use found its non-commercial license and questioned whether AI model vendors can track self-hosted deployments. The article concludes that technic…

00:00
2026-07-11
ranvier.systems
ai-infrastructure

Route to Where the KV Cache Is, Not Where It Was

Ranvier Systems introduced a load-balancing technique for LLM serving that routes requests based on where the KV cache currently resides rather than historical prefix matches, reducing P99 time-to-fir…

← prev page 13 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics