cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 22/38 feed RSS

// recent coverage 748 mentions

19:54
2026-07-31
devashish.me
artificial-intelligence

Notes from taking third spot at the Dell x NVIDIA AI Hackathon

A team of four developers won third place at the Dell x NVIDIA AI Hackathon by building Squidward, a 100% air-gapped, self-improving IT firewall that runs locally on a Dell GB10 with a Qwen3.6-27B mod…

18:35
2026-07-31
systems.seas.harvard.edu
large-language-models

Bursty arrivals speed up LLM inference

A benchmark study by an independent researcher found that burstier request arrivals speed up LLM inference, contradicting standard intuition. The analysis of vLLM serving shows that higher burstiness …

16:17
2026-07-31
localai.io
artificial-intelligence

Why we write our own C and C++ inference engines

LocalAI's 18 custom C and C++ inference engines, including vllm.cpp and depth-anything.cpp, outperform or match upstream Python-based engines while drastically reducing footprint, with vllm.cpp delive…

08:01
2026-07-31
github.com
ai-infrastructure

vLLM for Baidu Kunlun

Baidu's vLLM Kunlun plugin, open-sourced on Dec 8, 2025, enables vLLM to run on Kunlun XPU hardware, supporting 20+ models including Qwen, Llama, DeepSeek, GLM, and Gemma4, with features like quantiza…

07:16
2026-07-31
discuss.huggingface.co
large-language-models

Need generative model, high-quality description generation

A technical advisory recommends that production systems using raw LLM responses implement a structured content lifecycle, treating model output as a draft rather than the final artifact. The approach,…

06:39
2026-07-31
github.com
artificial-intelligence

HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA

HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and relate…

00:00
2026-07-31
glukhov.org
artificial-intelligence

Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama's simplicity can mask when a local experiment becomes a shared inference service, and vLLM offers better scheduling and observability for production workloads. Migration is warranted when multi…

17:22
2026-07-30
aws.amazon.com
artificial-intelligence

Deploying Kimi K3 on Amazon SageMaker HyperPod and Amazon EKS

On July 27, 2026, Moonshot AI released Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) model, described as the first open-weight system to reach the 3 trillion parameter class. The model ac…

17:22
2026-07-30
aws.amazon.com
artificial-intelligence

Deploying Kimi K3 on AWS

Moonshot AI released Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) model, on July 27, 2026, making it the first open-weight system to reach the 3 trillion parameter class. The model requi…

17:10
2026-07-30
digitalocean.com
artificial-intelligence

Under the Hood: Serving Kimi K3

DigitalOcean launched Kimi K3 on day 0, making it one of the most popular models on the platform and across the market, with the second most likes on Hugging Face and sixth most traffic on OpenCode. T…

15:11
2026-07-30
byteiota.com
mlops

Kubeflow 1.11: MLOps Gets a pip install Moment

Kubeflow 1.11, released at KubeCon Japan, introduces a pip-installable SDK that lets ML practitioners submit distributed training jobs, run hyperparameter searches, and fine-tune Llama 3.2 without wri…

← prev page 22 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics