cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 25/38 feed RSS

// recent coverage 748 mentions

03:02
2026-07-26
promptcube3.com
artificial-intelligence

OpenAI Models Leaking to Hugging Face: Analysis

OpenAI models are appearing on Hugging Face, offering researchers a rare opportunity to audit proprietary architectures through leaked weights, config files, and inference testing. The leaks expose hi…

19:46
2026-07-25
promptcube3.com
large-language-models

DeepSeek-R1 Local Deployment: My Hardware Struggles

A user reports that deploying the full 671B parameter DeepSeek-R1 model locally requires over 100GB of VRAM and is impractical on consumer hardware, with CUDA out-of-memory errors occurring even at sm…

17:03
2026-07-25
promptcube3.com
artificial-intelligence

Open-Weight AI: Model Wars vs Ecosystem Wars

Open-weight AI models offer freedom but require significant effort to deploy, according to a technical guide that argues the real value lies in deployment pipelines and developer ecosystems rather tha…

14:47
2026-07-25
promptcube3.com
large-language-models

KV Caching: Why Your LLM Inference Costs are Sky-High

KV caching, which stores Key and Value tensors in GPU memory to avoid recalculating attention for every token, is the primary driver of high LLM inference costs because the cache grows linearly with s…

19:03
2026-07-24
promptcube3.com
artificial-intelligence

Open Source AI: Why Closed-Source Lobbying is Failing

Open-source AI is winning over closed-source lobbying due to deployment flexibility, cost, and hardware ecosystem scale, according to a developer analysis. The global infrastructure for running open w…

16:09
2026-07-24
developers.googleblog.com
artificial-intelligence

Run Ray on TPU, Part 2: Ray AI libraries

Ray AI libraries (Serve, Data, Train) now support Google TPU slices through a topology field that reserves a whole ICI-connected slice, preventing multi-host deployment hangs. Ray Serve serves LLMs vi…

16:05
2026-07-24
promptcube3.com
artificial-intelligence

Claude Code Workflow: Open Weights vs. Closed Models

Open-weights models like Llama and Mistral give developers control over the inference stack, enabling custom quantization, KV cache optimization, and hardware-specific tuning that closed APIs cannot m…

15:05
2026-07-24
promptcube3.com
artificial-intelligence

Open-Weight Models: Why Big Tech is Fighting for Them

Open-weight models, which release trained parameters for public use, prevent a monopoly on AI intelligence by lowering barriers to entry for developers, according to a joint letter from major tech com…

14:50
2026-07-24
promptcube3.com
artificial-intelligence

Claude Code Workflow: Leveraging Open Weights for Local Dev

A developer reports that shifting to a hybrid AI workflow using Anthropic's Claude 3.5 Sonnet for architectural planning and a local Llama 3.1 8B model for unit test generation reduced token spend by …

06:50
2026-07-24
promptcube3.com
artificial-intelligence

Qwen local deployment, AI data analysis guide, GPT

Running Qwen locally via Ollama or vLLM with a local Python environment avoids cloud data exposure and token limits, enabling iterative work on large datasets. Qwen2.5-Coder (7B) on an RTX 3090 genera…

19:02
2026-07-23
promptcube3.com
ai-safety

Gate.cat: Stopping AI Agents from Running rm -rf

Gate.cat, an open-source tool from BGMLAI, intercepts shell commands from AI agents before execution to block dangerous operations like rm -rf, using a fail-closed parser with no LLM call in the veto …

17:58
2026-07-23
promptcube3.com
artificial-intelligence

AI model safety comparison, Qwen local deployment,

A hands-on comparison of AI model safety in local deployment shows Qwen 2 (7B) achieves a false refusal rate of ~4% on stress-test prompts, far lower than Llama 3 (8B) at ~12% and Mistral (7B v0.3) at…

17:00
2026-07-23
runtimewire.com
artificial-intelligence

Mia publishes a $14,000 desktop deployment for GLM-5.2

Mia's AI Lab published an open-source deployment stack on July 23rd that runs Z.ai's 753-billion-parameter GLM-5.2 model across three Nvidia DGX Spark computers, creating a desktop cluster with a 248,…

00:00
2026-07-23
rocm.blogs.amd.com
artificial-intelligence

Serve Kimi-K2.5-MXFP4 on MI355X with ATOM

AMD shows how to serve the pre-quantized amd/Kimi-K2.5-MXFP4 checkpoint on AMD Instinct MI355X GPUs using ATOM, a lightweight vLLM-like framework that integrates AITER kernels and exposes an OpenAI-co…

00:00
2026-07-23
fergusfinn.com
ai-infrastructure

Throughputmaxxing DeepSeek-V4-Flash on Isambard-AI

Doubleword, one of six companies in the first wave of UK Sovereign AI investments, achieved up to 3× the throughput of vLLM for DeepSeek-V4-Flash on a single node of Isambard-AI, the UK's national AI …

← prev page 25 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics