cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 6/23 feed RSS

// recent coverage 443 mentions

22:09
2026-08-04
sourcefeed.dev
artificial-intelligence

Mistral Shrinks Content Moderation to a Single Token

Mistral released Shieldstral, a 3B-parameter Apache-2.0 guard model that moderates text and images by answering a plain-language yes/no question about a policy in a single forward pass, emitting a sin…

17:59
2026-08-04
promptcube3.com
artificial-intelligence

Mistral Shieldstral: 3B Open-Weights Multimodal Moderation

Mistral AI released Shieldstral, a 3B-parameter open-weights multimodal moderation model that outputs toxicity and safety scores across multiple axes, trained on ~600K human-judged examples covering h…

15:17
2026-08-04
dev.to
artificial-intelligence

DeepSeek V4 Flash API Cost: Thinking Mode Corrupts Strict JSON

DeepSeek's V4 Flash 0731 build corrupts integer fields in strict JSON schema outputs when thinking mode is enabled, failing 8 of 13 runs across two request paths, while disabling thinking fixes all ru…

13:10
2026-08-04
sourcefeed.dev
artificial-intelligence

DeepSeek V4 Flash on One AMD GPU Took Nine Patches

A single AMD MI300X GPU with 192 GB of HBM3 now serves DeepSeek's 284B-parameter DeepSeek-V4-Flash-0731 checkpoint in mixed FP4+FP8 format, requiring nine patch overlays against a vLLM ROCm nightly pl…

10:35
2026-08-04
promptcube3.com
artificial-intelligence

Production AI Infrastructure

A production AI infrastructure engineer reports that six months of operating an on-prem RAG pipeline for internal document search over 2M legal documents with a Llama-3.1-70B model revealed critical p…

10:00
2026-08-04
github.com
artificial-intelligence

DeepSeek V4 Flash on a Single AMD MI300X

A production configuration for running DeepSeek-V4-Flash-0731 on a single AMD MI300X GPU achieves 168.6 tok/s median single-stream decode and 542 tok/s aggregate across 8 concurrent streams, with the …

07:08
2026-08-04
gainz.fast
artificial-intelligence

Show HN: Gainz.fast – Local Inference, Faster

Carsen Klock launched Gainz.fast, a new site for benchmarking local AI inference speed, reporting that the Laguna XS 2.1 model on an AMD R9700 with llama.cpp HIP achieved 143.3 tokens per second, a 31…

19:54
2026-07-31
devashish.me
artificial-intelligence

Notes from taking third spot at the Dell x NVIDIA AI Hackathon

A team of four developers won third place at the Dell x NVIDIA AI Hackathon by building Squidward, a 100% air-gapped, self-improving IT firewall that runs locally on a Dell GB10 with a Qwen3.6-27B mod…

18:35
2026-07-31
systems.seas.harvard.edu
large-language-models

Bursty arrivals speed up LLM inference

A benchmark study by an independent researcher found that burstier request arrivals speed up LLM inference, contradicting standard intuition. The analysis of vLLM serving shows that higher burstiness …

← prev page 6 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics