cd/entity/TensorRT-LLM· home entities TensorRT-LLM
grep -l @tensorrt-llm /news/*.json | wc -l → 45

TensorRT-LLM

mentions 45 type Organization page 1/3 feed RSS

// recent coverage 45 mentions

16:13
2026-08-27
servethehome.com
ai-infrastructure

Oxmiq Labs HBF in AI Compute at Hot Chips 2026

Oxmiq Labs presented high-bandwidth flash (HBF) as a capacity tier for AI inference at Hot Chips 2026, claiming it delivers 8 to 16 times the capacity of HBM at the same cost. The company detailed HBF…

21:26
2026-08-21
promptcube3.com
artificial-intelligence

Nvidia's latest demo proves the inference stack matters more

Nvidia's latest demo shows that optimizing the inference stack, not the model weights, is the key to performance, achieving 4.2× higher throughput at half the latency on identical H100 hardware with L…

12:00
2026-08-20
neon.com
artificial-intelligence

Open-weight models are fast on Neon AI Gateway. Here's why

Neon AI Gateway, powered by Databricks Foundation Model APIs, achieves fast open-weight model inference through optimizations like continuous batching, KV-cache paging, and prompt caching, which boost…

00:20
2026-08-19
inco.ai
artificial-intelligence

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a parallel speculative decoding technique that delivers over 20% more output from every verification pass with around 1% added cycle latency, achieving 2.7–3.4× throughput o…

23:21
2026-08-18
baseten.co
ai-infrastructure

Inference Engineering by Philip Kiely – Digital Download

Philip Kiely's new book, 'Inference Engineering,' is now available as a digital download, offering a comprehensive guide to the technologies and techniques powering AI inference across runtime, infras…

18:49
2026-08-13
acefleet.dev
ai-infrastructure

Scale Your AI Revenue – Not Your Cloud Bill

A new guide outlines strategies for scaling AI revenue while controlling cloud costs, covering hardware accelerators from NVIDIA, Groq, and Cerebras, inference engines like vLLM and TensorRT-LLM, and …

15:42
2026-08-12
dev.to
mlops

How We Cut Inference Cold Starts from Minutes to Seconds

Engineers at an unnamed company cut inference cold start times from minutes to seconds by profiling their startup sequence and finding that 90% of the delay came from moving large files. They reduced …

00:00
2026-08-11
modelplane.ai
artificial-intelligence

Why Day 0 for Nemotron 3.5 Lightning wasn't a scramble

NVIDIA released Nemotron-3.5-Lightning, a 30B mixture-of-experts model with 3B active parameters, on the same day Modelplane, an open-source fleet-level control plane for inference, achieved zero-day …

01:21
2026-08-05
latent.space
artificial-intelligence

[AINews] Megakernels are so dead and so back

Cursor released an open-source megakernel that delivers a 41% increase in tokens per second, potentially saving billions of dollars at scale, according to a post by Stuart Sul, who leads the team behi…

06:39
2026-07-31
github.com
artificial-intelligence

HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA

HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and relate…

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics