cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 7/23 feed RSS

// recent coverage 443 mentions

16:17
2026-07-31
localai.io
artificial-intelligence

Why we write our own C and C++ inference engines

LocalAI's 18 custom C and C++ inference engines, including vllm.cpp and depth-anything.cpp, outperform or match upstream Python-based engines while drastically reducing footprint, with vllm.cpp delive…

08:01
2026-07-31
github.com
ai-infrastructure

vLLM for Baidu Kunlun

Baidu's vLLM Kunlun plugin, open-sourced on Dec 8, 2025, enables vLLM to run on Kunlun XPU hardware, supporting 20+ models including Qwen, Llama, DeepSeek, GLM, and Gemma4, with features like quantiza…

07:16
2026-07-31
discuss.huggingface.co
large-language-models

Need generative model, high-quality description generation

A technical advisory recommends that production systems using raw LLM responses implement a structured content lifecycle, treating model output as a draft rather than the final artifact. The approach,…

06:39
2026-07-31
github.com
artificial-intelligence

HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA

HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and relate…

00:00
2026-07-31
glukhov.org
artificial-intelligence

Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama's simplicity can mask when a local experiment becomes a shared inference service, and vLLM offers better scheduling and observability for production workloads. Migration is warranted when multi…

17:22
2026-07-30
aws.amazon.com
artificial-intelligence

Deploying Kimi K3 on AWS

Moonshot AI released Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) model, on July 27, 2026, making it the first open-weight system to reach the 3 trillion parameter class. The model requi…

17:22
2026-07-30
aws.amazon.com
artificial-intelligence

Deploying Kimi K3 on Amazon SageMaker HyperPod and Amazon EKS

On July 27, 2026, Moonshot AI released Kimi K3, a 2.8 trillion parameter Mixture of Experts (MoE) model, described as the first open-weight system to reach the 3 trillion parameter class. The model ac…

17:10
2026-07-30
digitalocean.com
artificial-intelligence

Under the Hood: Serving Kimi K3

DigitalOcean launched Kimi K3 on day 0, making it one of the most popular models on the platform and across the market, with the second most likes on Hugging Face and sixth most traffic on OpenCode. T…

15:11
2026-07-30
byteiota.com
mlops

Kubeflow 1.11: MLOps Gets a pip install Moment

Kubeflow 1.11, released at KubeCon Japan, introduces a pip-installable SDK that lets ML practitioners submit distributed training jobs, run hyperparameter searches, and fine-tune Llama 3.2 without wri…

15:08
2026-07-30
redhat.com
ai-infrastructure

Why self-hosted inference is essential

Red Hat AI warns that enterprises relying on third-party hosted APIs for AI agent inference undermine their own data sovereignty, as every prompt and tool call routes through external datacenters. The…

13:59
2026-07-30
openalternative.co
artificial-intelligence

LocalAI

LocalAI, a self-hosted runtime that runs AI workloads on user-controlled hardware, offers an OpenAI-compatible API supporting text generation, vision, speech, image and video generation, embeddings, a…

13:05
2026-07-30
github.com
artificial-intelligence

Kimi k3 run on RTX 5090

RunInfra enables running Kimi-Linear-48B, a distilled version of the full 2.78-trillion-parameter Kimi K3 model, on a single consumer GPU such as the RTX 5090 with 32 GB VRAM, achieving 113.83 tokens …

11:01
2026-07-30
promptcube3.com
large-language-models

Open-Weight Models Now Match Proprietary Titans

The accuracy gap between the best open-weight models and GPT-4o has shrunk to under 3% on structured data tasks, according to a developer's hands-on analysis. A fine-tuned Qwen2.5 72B model achieved 9…

22:08
2026-07-29
byteiota.com
artificial-intelligence

DeepSeek V4 Pro: 80.6% SWE-Bench at $0.87/M Output

DeepSeek V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters, scores 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and achieving the highest score for any open-weight model. Pric…

← prev page 7 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics