cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 10/23 feed RSS

// recent coverage 443 mentions

16:05
2026-07-24
promptcube3.com
artificial-intelligence

Claude Code Workflow: Open Weights vs. Closed Models

Open-weights models like Llama and Mistral give developers control over the inference stack, enabling custom quantization, KV cache optimization, and hardware-specific tuning that closed APIs cannot m…

15:05
2026-07-24
promptcube3.com
artificial-intelligence

Open-Weight Models: Why Big Tech is Fighting for Them

Open-weight models, which release trained parameters for public use, prevent a monopoly on AI intelligence by lowering barriers to entry for developers, according to a joint letter from major tech com…

14:50
2026-07-24
promptcube3.com
artificial-intelligence

Claude Code Workflow: Leveraging Open Weights for Local Dev

A developer reports that shifting to a hybrid AI workflow using Anthropic's Claude 3.5 Sonnet for architectural planning and a local Llama 3.1 8B model for unit test generation reduced token spend by …

06:50
2026-07-24
promptcube3.com
artificial-intelligence

Qwen local deployment, AI data analysis guide, GPT

Running Qwen locally via Ollama or vLLM with a local Python environment avoids cloud data exposure and token limits, enabling iterative work on large datasets. Qwen2.5-Coder (7B) on an RTX 3090 genera…

19:02
2026-07-23
promptcube3.com
ai-safety

Gate.cat: Stopping AI Agents from Running rm -rf

Gate.cat, an open-source tool from BGMLAI, intercepts shell commands from AI agents before execution to block dangerous operations like rm -rf, using a fail-closed parser with no LLM call in the veto …

17:58
2026-07-23
promptcube3.com
artificial-intelligence

AI model safety comparison, Qwen local deployment,

A hands-on comparison of AI model safety in local deployment shows Qwen 2 (7B) achieves a false refusal rate of ~4% on stress-test prompts, far lower than Llama 3 (8B) at ~12% and Mistral (7B v0.3) at…

17:00
2026-07-23
runtimewire.com
artificial-intelligence

Mia publishes a $14,000 desktop deployment for GLM-5.2

Mia's AI Lab published an open-source deployment stack on July 23rd that runs Z.ai's 753-billion-parameter GLM-5.2 model across three Nvidia DGX Spark computers, creating a desktop cluster with a 248,…

00:00
2026-07-23
rocm.blogs.amd.com
artificial-intelligence

Serve Kimi-K2.5-MXFP4 on MI355X with ATOM

AMD shows how to serve the pre-quantized amd/Kimi-K2.5-MXFP4 checkpoint on AMD Instinct MI355X GPUs using ATOM, a lightweight vLLM-like framework that integrates AITER kernels and exposes an OpenAI-co…

00:00
2026-07-23
fergusfinn.com
ai-infrastructure

Throughputmaxxing DeepSeek-V4-Flash on Isambard-AI

Doubleword, one of six companies in the first wave of UK Sovereign AI investments, achieved up to 3× the throughput of vLLM for DeepSeek-V4-Flash on a single node of Isambard-AI, the UK's national AI …

12:01
2026-07-22
pub.towardsai.net
large-language-models

The Complete Technical Guide to Running LLMs Locally in 2026

A technical guide to running large language models locally in 2026 provides hardware math, quantization tradeoffs, and benchmarks of five inference engines, with case studies from the author's 16GB Ap…

07:36
2026-07-22
leaddev.com
large-language-models

Your LLM inference benchmark is lying to you

Synthetic LLM inference benchmarks misrepresent production performance because they use fixed prompt lengths, steady request rates, and single-model hardware, while real traffic is bursty and variable…

← prev page 10 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics