cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 15/23 feed RSS

// recent coverage 443 mentions

00:00
2026-07-07
huggingbay.xyz
artificial-intelligence

answerdotai/ModernBERT-base

Answerdotai released ModernBERT-base, a 1.2 GB encoder model under Apache-2.0 license, designed for fill-mask tasks with easy local fit on consumer hardware. The model is available via Hugging Bay wit…

00:00
2026-07-07
huggingbay.xyz
artificial-intelligence

Qwen/Qwen3-Reranker-0.6B

Hugging Face lists Qwen/Qwen3-Reranker-0.6B, a text-ranking model under Apache-2.0 license, with 2.1 million downloads but pending security scan and no hosted files. The model requires file-size revie…

00:00
2026-07-07
huggingbay.xyz
artificial-intelligence

distilbert/distilbert-base-uncased

Hugging Bay lists distilbert/distilbert-base-uncased, a 260 MB DistilBERT encoder under Apache-2.0, as a compact resilience fallback for high-demand AI artifacts. The model has 500 upstream downloads …

00:00
2026-07-07
huggingbay.xyz
large-language-models

Qwen/Qwen3-0.6B

Qwen released Qwen3-0.6B, a small Apache-2.0 licensed language model designed for local experiments, lightweight agents, and edge testing. The model is available on Hugging Bay with external metadata …

19:47
2026-07-06
huggingbay.xyz
large-language-models

mradermacher/sarashina2-70b-GGUF

Hugging Bay has hosted the mradermacher/sarashina2-70b-GGUF model, a 445.1 GB quantized version of the sbintuitions/sarashina2-70b base model under the MIT license, with 2 of 15 files verified and sca…

20:44
2026-07-03
tigera.io
ai-agents

Six AI agent SDKs for enterprise Kubernetes, compared

Six AI agent SDKs—LangGraph, CrewAI, Google ADK, and others—are compared for enterprise Kubernetes deployment, with most being model-agnostic and containerizable for on-premise use, though Anthropic's…

04:19
2026-07-03
huggingbay.xyz
large-language-models

openai/gpt-oss-20b

OpenAI released the GPT-OSS-20B model on Hugging Face under the Apache-2.0 license, a 38.5 GB text-generation model with over 7 million downloads. The model requires large hardware such as multi-GPU o…

09:07
2026-07-01
glukhov.org
large-language-models

Speculative Decoding: 20-50% Faster LLM Inference

Speculative decoding accelerates large language model inference by 20-50% without quality loss, using a draft-verify mechanism that generates multiple tokens per forward pass. The technique amortizes …

03:09
2026-07-01
byteiota.com
large-language-models

MiniMax M3: Open-Weight Model That Beats GPT-5.5 on Coding

MiniMax released M3, a 428-billion-parameter open-weight model, on June 7, achieving 59.0% on SWE-Bench Pro—slightly outperforming GPT-5.5's 58.6%—at $0.30 per million input tokens, making it 16 times…

20:04
2026-06-30
letsdatascience.com
large-language-models

Article Compares Continuous and Static Batching in LLM Inference

A new article compares continuous batching and static batching in LLM inference, explaining how techniques in vLLM and TGI improve throughput and reduce latency. The choice of batching strategy affect…

00:00
2026-06-30
jasonrobert.dev
artificial-intelligence

News Summary for June 30, 2026

Agentic AI systems are maturing from prototypes into production-grade infrastructure, with vLLM's Micro-Agent framework demonstrating that serving-layer orchestration can match or beat frontier models…

00:00
2026-06-30
aclanthology.org
artificial-intelligence

CUHKSZ Simultaneous Speech Translation System for IWSLT 2026

The CUHKSZ team submitted a simultaneous speech translation system to IWSLT 2026, built on Qwen3-Omni-30B-A3B with LoRA adaptation, achieving 40.5 BLEU for English→Chinese and 27.7 BLEU for English→Ge…

← prev page 15 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics