cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 21/38 feed RSS

// recent coverage 748 mentions

10:11
2026-08-05
byteiota.com
artificial-intelligence

Kimi K3 Open Weights: Self-Hosting Reality Check

Moonshot AI released the weights for Kimi K3 on July 27, a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window, scoring third globally behind Claude Fable 5 Max and G…

02:49
2026-08-05
dev.to
large-language-models

Measuring LLM Prefix Caching: The Cache Hit Rate Metric

An engineer's benchmarking guide introduces a cache hit rate metric for measuring prefix caching effectiveness in LLM serving, implemented in the open-source tool llmperf-rs. The metric calculates the…

00:00
2026-08-05
autonomous.ai
artificial-intelligence

On-Premise AI Explained: Benefits, Cost, and Setup

On-premise AI, which involves running AI models on hardware owned by the organization, is gaining traction in 2026 due to data control, predictable costs, and reliable access, according to a guide fro…

22:09
2026-08-04
sourcefeed.dev
artificial-intelligence

Mistral Shrinks Content Moderation to a Single Token

Mistral released Shieldstral, a 3B-parameter Apache-2.0 guard model that moderates text and images by answering a plain-language yes/no question about a policy in a single forward pass, emitting a sin…

17:59
2026-08-04
promptcube3.com
artificial-intelligence

Mistral Shieldstral: 3B Open-Weights Multimodal Moderation

Mistral AI released Shieldstral, a 3B-parameter open-weights multimodal moderation model that outputs toxicity and safety scores across multiple axes, trained on ~600K human-judged examples covering h…

15:17
2026-08-04
dev.to
artificial-intelligence

DeepSeek V4 Flash API Cost: Thinking Mode Corrupts Strict JSON

DeepSeek's V4 Flash 0731 build corrupts integer fields in strict JSON schema outputs when thinking mode is enabled, failing 8 of 13 runs across two request paths, while disabling thinking fixes all ru…

13:10
2026-08-04
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

A single AMD MI300X GPU with 192 GB of HBM3 now serves DeepSeek's 284B-parameter DeepSeek-V4-Flash-0731 checkpoint in mixed FP4+FP8 format, requiring nine patch overlays against a vLLM ROCm nightly pl…

10:35
2026-08-04
promptcube3.com
artificial-intelligence

Production AI Infrastructure

A production AI infrastructure engineer reports that six months of operating an on-prem RAG pipeline for internal document search over 2M legal documents with a Llama-3.1-70B model revealed critical p…

10:00
2026-08-04
github.com
artificial-intelligence

DeepSeek V4 Flash on a Single AMD MI300X

A production configuration for running DeepSeek-V4-Flash-0731 on a single AMD MI300X GPU achieves 168.6 tok/s median single-stream decode and 542 tok/s aggregate across 8 concurrent streams, with the …

07:08
2026-08-04
gainz.fast
artificial-intelligence

Show HN: Gainz.fast – Local Inference, Faster

Carsen Klock launched Gainz.fast, a new site for benchmarking local AI inference speed, reporting that the Laguna XS 2.1 model on an AMD R9700 with llama.cpp HIP achieved 143.3 tokens per second, a 31…

00:00
2026-08-03
autonomous.ai
ai-infrastructure

How to Build a Local AI Server, by Budget Tier

A guide from Autonomous AI outlines how to build a local AI server by budget tier, emphasizing that VRAM is the most critical specification for running AI models locally. The guide provides VRAM requi…

← prev page 21 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics