cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 14/23 feed RSS

// recent coverage 443 mentions

22:34
2026-07-09
github.com
artificial-intelligence

Orbit, an Open-Source Toolkit for Retrieval-Based Inference

Orbit, an open-source AI gateway for retrieval-based inference, has been released, enabling self-hosted private RAG, natural-language data access, and tool-calling agents across 37+ model providers. T…

20:53
2026-07-09
discuss.huggingface.co
large-language-models

Distinguish between thinking and responding during generation

A developer building a custom logits processor for generative models seeks to implement two-phase generation that separates reasoning from verbatim text output. The challenge involves detecting the tr…

13:22
2026-07-09
cryptobriefing.com
artificial-intelligence

Ollama raises $65M to bring local AI to 9 million developers

Ollama, an open-source tool for running AI models locally, raised $65 million in a funding round led by Benchmark. The company aims to scale its developer platform to 9 million users, capitalizing on …

23:50
2026-07-08
letsdatascience.com
artificial-intelligence

Hugging Face Speeds Transformers Inference in vLLM

Hugging Face announced that its Transformers modeling backend for vLLM now matches or exceeds native vLLM speed for compatible architectures, reducing the need for separate hand-optimized serving port…

07:10
2026-07-08
byteiota.com
artificial-intelligence

Tencent Hy3: 295B MoE Hits SWE-Bench 78 — Free API Ends July 21

Tencent released Hy3, a 295B Mixture-of-Experts model with Apache 2.0 weights, achieving a 78.0 SWE-bench Verified score and leading open-weight models on tool and search benchmarks. A free API on Ope…

00:00
2026-07-08
huggingface.co
artificial-intelligence

Native-speed vLLM transformers modeling backend

Hugging Face announced that the transformers vLLM modeling backend now matches or exceeds native vLLM throughput for many LLM architectures, allowing model authors to run their transformers implementa…

15:20
2026-07-07
huggingface.co
artificial-intelligence

Hugging Face Models on Foundry Managed Compute

Microsoft Foundry now offers a curated catalog of Hugging Face open-weight models deployable on Foundry Managed Compute, with pre-staged weights in Azure and built-in enterprise security, governance, …

00:42
2026-07-07
letsdatascience.com
large-language-models

Tencent open-sources Hy3 295B MoE model

Tencent released Hy3, an Apache-2.0 open-weight Mixture-of-Experts model with 295B total parameters, 21B active parameters per token, and a 256K context window. The release includes official artifacts…

00:00
2026-07-07
huggingbay.xyz
large-language-models

Qwen/Qwen2.5-1.5B-Instruct

Qwen released Qwen2.5-1.5B-Instruct, an Apache-2.0 licensed 1.5B parameter instruct model for local chat and agent tasks, requiring 8-16 GB RAM/VRAM. The model is available on Hugging Bay with externa…

← prev page 14 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics