cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 31/38 feed RSS

// recent coverage 748 mentions

00:00
2026-06-30
aclanthology.org
artificial-intelligence

CUHKSZ Simultaneous Speech Translation System for IWSLT 2026

The CUHKSZ team submitted a simultaneous speech translation system to IWSLT 2026, built on Qwen3-Omni-30B-A3B with LoRA adaptation, achieving 40.5 BLEU for English→Chinese and 27.7 BLEU for English→Ge…

00:00
2026-06-30
jasonrobert.dev
artificial-intelligence

News Summary for June 30, 2026

Agentic AI systems are maturing from prototypes into production-grade infrastructure, with vLLM's Micro-Agent framework demonstrating that serving-layer orchestration can match or beat frontier models…

00:46
2026-06-28
github.com
artificial-intelligence

AMD Strix Halo RDMA Cluster Setup Guide

AMD Strix Halo cluster setup guide details how to configure a two-node system linked via Intel E810 RoCE v2 for distributed vLLM inference using Tensor Parallelism. The guide covers hardware prerequis…

15:27
2026-06-27
cefboud.com
large-language-models

Distributed LLM Inference with LLM-d

A new open-source tool called llm-d acts as an LLM-aware load balancer for distributed inference, intelligently routing requests across vLLM instances based on KV cache locality and GPU utilization. B…

10:10
2026-06-27
dev.to
ai-agents

DeerFlow 2.0 Review: ByteDance's Open SuperAgent Harness

ByteDance open-sourced DeerFlow 2.0, a long-horizon agent runtime that orchestrates sub-agents, sandboxes, persistent memory, and an extensible skill system. The project reached 74,960 GitHub stars an…

08:06
2026-06-27
github.com
ai-tools

Show HN: Brytlog – AI logger

Developer released Brytlog, an open-source AI logger that replaces raw terminal output with concise AI summaries to save developers time and money. The tool acts as a pre-processor for agentic workflo…

22:35
2026-06-26
cmart.blog
large-language-models

Inference Cards

A new plaintext markup format called Inference Cards aims to standardize how self-hosted LLM performance claims are communicated, requiring details like model variant, quantization, hardware, inferenc…

20:01
2026-06-26
pub.towardsai.net
large-language-models

GOSIM Paris: This Is What Open Source AI Looks Like in 2026

GOSIM Paris 2025, held at Station F on May 5-6, showcased open-source AI developments including LLMs advancing in mathematical reasoning, a call for transparency over speed, and the introduction of Ta…

13:10
2026-06-26
byteiota.com
ai-infrastructure

DGX Spark June 2026: Four Nodes, 700B Models Locally

NVIDIA's June 2026 DGX Spark update introduces automated four-node clustering via Cluster Assistant, enabling local inference of models up to 700B parameters. The update also delivers a 2.6x throughpu…

12:26
2026-06-26
3hcloud.com
ai-agents

How to Set Up and Deploy an OpenClaw AI Agent on a VPS

A new guide walks users through deploying an OpenClaw AI agent on a virtual private server, balancing cost, availability, and privacy. The tutorial covers server configuration, system requirements, an…

10:30
2026-06-26
aazar.me
large-language-models

Stop generating what you already have

A developer reduced LLM extraction latency from 42 seconds to 6 seconds by replacing verbatim text copying with pointer-based extraction and splitting a single large call into multiple parallel calls.…

20:42
2026-06-25
huggingface.co
ai-infrastructure

Run a vLLM Server on HF Jobs in One Command

Hugging Face launched a one-command method to run a vLLM server on its Jobs infrastructure, enabling users to quickly deploy models for testing, evaluation, or batch generation. The feature uses the o…

15:40
2026-06-25
anaconda.com
machine-learning

Why ML/AI Developers and Platform Teams Choose Metaflow

Metaflow, an open-source ML/AI orchestration framework, ranks first in every category of the Cloud Native Computing Foundation's latest Technology Radar report, with 51% of surveyed users highly likel…

12:04
2026-06-25
devclubhouse.com
large-language-models

The Real Cost of the Open-Weight Price Collapse

The launch of Z.ai's GLM 5.2 and DeepSeek V4 Flash has created a 50x price gap between open-weight APIs and closed frontier models, reshaping the build-versus-buy calculus for developers. While open-w…

11:08
2026-06-25
flama.dev
large-language-models

LLM APIs with built-in chatbot in 1 line of code

Flama 2.0 introduces a CLI tool that allows users to download, package, and serve large language models from HuggingFace with a single command, including a built-in chat interface and production-ready…

← prev page 31 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics