cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 1/38 feed RSS

// recent coverage 743 mentions

07:21
2026-10-03
forum.level1techs.com
ai-infrastructure

Quad R9700's on AM4 Shenanigans

A user on the Level1Techs forum reported running four AMD Radeon R9700 GPUs on an AM4 platform to requantize a DeepSeek model to MXFP4 and run it in the vLLM engine, with the first layer of a 48-layer…

00:00
2026-10-03
mindstudio.ai
large-language-models

IQuest-Q1: How to Self-Host the 320B Agentic Coding Model

IQuest released IQuest-Q1, an open-weight 320B-parameter Mixture-of-Experts coding model that activates only 15B parameters per token and supports a 524,288-token context window, with official deploym…

06:11
2026-10-02
byteiota.com
ai-tools

Janus: Run GGUF Models Locally With One Go Binary

Janus, an MIT-licensed single Go binary from the Vibra-Ingenn project, launched as a Show HN on October 1 with 42+ points, wrapping llama.cpp's Vulkan backend to run GGUF models on AMD, Intel, or NVID…

00:00
2026-10-02
mindstudio.ai
ai-infrastructure

How to Cluster Two NVIDIA DGX Sparks for Local LLM Inference

Clustering two NVIDIA DGX Sparks over a 200 GB RoCE RDMA cable and running tensor-parallel inference with vLLM lets the pair serve models in the 150 to 160 GB range, such as DeepSeek V4 Flash, that do…

00:00
2026-10-02
blog.kolen.dev
large-language-models

Extracting metadata with an LLM is data cleaning

A technical blog post argues that non-reproducible LLM metadata extraction should be treated as a data-cleaning and statistical problem rather than a determinism problem, citing He (2025), who obtaine…

18:09
2026-10-01
dev.to
ai-infrastructure

Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs

A developer evaluated speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs, testing draft methods including MTP, EAGLE-3, DFlash, and DSpark that propose candidate tokens verified in a …

00:00
2026-10-01
kraghavan.ca
ai-infrastructure

pd-ratio-coordinator — Does My Own Tool Actually Work?

A developer's Kubernetes operator, pd-ratio-coordinator, failed to compile on its committed KRAG-1 branch, with two real bugs — a DeepCopy type mismatch in register.go and four unqualified Bottleneck …

16:00
2026-09-30
konghq.com
ai-infrastructure

Kong AI Gateway 2.2: Built for what comes next in AI

Kong released AI Gateway 2.2, adding native support for TypeSafe JEV as a provider and format, support for Skills APIs, and a new passthrough mode for custom or non-standard AI interfaces such as vLLM…

00:00
2026-09-30
mindstudio.ai
large-language-models

IQuest-Q1: Inside the 320B MoE Model Built for Agentic Coding

IQuest released IQuest-Q1, an open-weight 320B-parameter Mixture-of-Experts language model with 15B active parameters per token and a 512K-token context window, built for agentic coding and multi-step…

00:00
2026-09-30
mindstudio.ai
large-language-models

How to Deploy IQuest-Q1 with SGLang or vLLM

IQuest released IQuest-Q1, a 320-billion-parameter Mixture-of-Experts model with roughly 15 billion active parameters per token, an 88-layer transformer, 256 experts (8 active), and a 524,288-token co…

00:00
2026-09-30
rocm.blogs.amd.com
ai-infrastructure

AIM: Unified User Experience from Profile Discovery to Deployment

AMD detailed its AMD Inference Microservices (AIMs), standardized Docker-based inference microservices that automatically select a runtime profile from inputs including model, precision, engine, laten…

page 1 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics