cd/entity/TensorRT-LLM· home entities TensorRT-LLM
grep -l @tensorrt-llm /news/*.json | wc -l → 45

TensorRT-LLM

mentions 45 type Organization page 2/3 feed RSS

// recent coverage 45 mentions

12:08
2026-07-27
sourcefeed.dev
artificial-intelligence

The Real Cost of 'Just Use vLLM,' According to Netflix

Netflix's AI Platform team published a detailed account of its LLM serving platform, revealing that version pinning between NVIDIA Triton Inference Server and vLLM, a Python GIL bottleneck, and KV-cac…

05:18
2026-07-26
kraghavan.ca
large-language-models

Introduction to LLM Inference

A senior engineer with 11 years of distributed systems experience explains the full LLM inference pipeline, from request arrival to text output, detailing the GGUF file structure and the distinction b…

16:05
2026-07-24
promptcube3.com
artificial-intelligence

Claude Code Workflow: Open Weights vs. Closed Models

Open-weights models like Llama and Mistral give developers control over the inference stack, enabling custom quantization, KV cache optimization, and hardware-specific tuning that closed APIs cannot m…

18:07
2026-07-17
netflixtechblog.medium.com
large-language-models

In-House LLM Serving at Netflix

Netflix's AI Platform team built an in-house LLM serving stack, running the full pipeline from model deployment through inference inside its existing production environment. The team selected vLLM as …

14:18
2026-07-09
baremetalrt.ai
artificial-intelligence

BareMetalRT – TensorRT-LLM running natively on Windows (no WSL)

BareMetalRT launches TensorRT-LLM natively on Windows without WSL, enabling heterogeneous tensor parallelism across consumer GPUs over standard networking. Users can run HuggingFace models locally via…

15:20
2026-07-07
huggingface.co
artificial-intelligence

Hugging Face Models on Foundry Managed Compute

Microsoft Foundry now offers a curated catalog of Hugging Face open-weight models deployable on Foundry Managed Compute, with pre-staged weights in Azure and built-in enterprise security, governance, …

09:07
2026-07-01
glukhov.org
large-language-models

Speculative Decoding: 20-50% Faster LLM Inference

Speculative decoding accelerates large language model inference by 20-50% without quality loss, using a draft-verify mechanism that generates multiple tokens per forward pass. The technique amortizes …

13:44
2026-06-30
aimultiple.com
ai-products

DGX Spark vs. Mac Studio and Halo

NVIDIA's DGX Spark, a $4,699 desktop AI supercomputer with 128GB unified memory, launched in 2025, offering one petaflop of FP4 performance. Benchmarks show it excels at prompt processing but lags in …

12:04
2026-06-25
devclubhouse.com
large-language-models

The Real Cost of the Open-Weight Price Collapse

The launch of Z.ai's GLM 5.2 and DeepSeek V4 Flash has created a 50x price gap between open-weight APIs and closed frontier models, reshaping the build-versus-buy calculus for developers. While open-w…

← prev page 2 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics