cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 29/38 feed RSS

// recent coverage 748 mentions

09:28
2026-07-11
pub.towardsai.net
artificial-intelligence

Can AI Model Vendors Track Your Self-Hosted Deployment?

A developer investigating FLUX.1 [dev] for commercial use found its non-commercial license and questioned whether AI model vendors can track self-hosted deployments. The article concludes that technic…

00:00
2026-07-11
ranvier.systems
ai-infrastructure

Route to Where the KV Cache Is, Not Where It Was

Ranvier Systems introduced a load-balancing technique for LLM serving that routes requests based on where the KV cache currently resides rather than historical prefix matches, reducing P99 time-to-fir…

22:34
2026-07-09
github.com
artificial-intelligence

Orbit, an Open-Source Toolkit for Retrieval-Based Inference

Orbit, an open-source AI gateway for retrieval-based inference, has been released, enabling self-hosted private RAG, natural-language data access, and tool-calling agents across 37+ model providers. T…

20:53
2026-07-09
discuss.huggingface.co
large-language-models

Distinguish between thinking and responding during generation

A developer building a custom logits processor for generative models seeks to implement two-phase generation that separates reasoning from verbatim text output. The challenge involves detecting the tr…

13:22
2026-07-09
cryptobriefing.com
artificial-intelligence

Ollama raises $65M to bring local AI to 9 million developers

Ollama, an open-source tool for running AI models locally, raised $65 million in a funding round led by Benchmark. The company aims to scale its developer platform to 9 million users, capitalizing on …

23:50
2026-07-08
letsdatascience.com
artificial-intelligence

Hugging Face Speeds Transformers Inference in vLLM

Hugging Face announced that its Transformers modeling backend for vLLM now matches or exceeds native vLLM speed for compatible architectures, reducing the need for separate hand-optimized serving port…

07:10
2026-07-08
byteiota.com
artificial-intelligence

Tencent Hy3: 295B MoE Hits SWE-Bench 78 — Free API Ends July 21

Tencent released Hy3, a 295B Mixture-of-Experts model with Apache 2.0 weights, achieving a 78.0 SWE-bench Verified score and leading open-weight models on tool and search benchmarks. A free API on Ope…

00:00
2026-07-08
huggingface.co
artificial-intelligence

Native-speed vLLM transformers modeling backend

Hugging Face announced that the transformers vLLM modeling backend now matches or exceeds native vLLM throughput for many LLM architectures, allowing model authors to run their transformers implementa…

15:20
2026-07-07
huggingface.co
artificial-intelligence

Hugging Face Models on Foundry Managed Compute

Microsoft Foundry now offers a curated catalog of Hugging Face open-weight models deployable on Foundry Managed Compute, with pre-staged weights in Azure and built-in enterprise security, governance, …

← prev page 29 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics