cd/entity/FlashInfer· home› entities› FlashInfer
grep -l @flashinfer /news/*.json | wc -l → 17

FlashInfer

mentions 17 type Organization feed RSS

// recent coverage 17 mentions

18:35
2026-10-03
gist.github.com
ai-infrastructure

Dockerfile for running Aleph Alpha Kolibri-1 on DGX Spark

A developer published a Dockerfile that layers Aleph Alpha's proprietary Kolibri-1 inference plugin onto a native SM121 (GB10) vLLM build for the DGX Spark, preserving the Blackwell-compiled kernels. …

22:51
2026-09-29
zhang677.github.io
artificial-intelligence

PTXBench: What about just CUDA-PTX?

PTXBench, a benchmark released August 17, 2026 with code on GitHub, evaluates how well frontier LLMs can write architecture-specific CUDA-PTX for H100 and B200 GPUs, finding GEMM nearly solved while a…

12:30
2026-09-03
wirt.ee
large-language-models

On-Prem LLM Inference

A technical guide details the production deployment of on-premises large language model inference using vLLM, covering an 8-GPU node running GLM-5.2/5.3 (NVFP4 MoE) and a single L40S running gemma-4-2…

07:00
2026-08-24
hiraditya.github.io
artificial-intelligence

A Bug Is a Violation of a Specification

A bug is a violation of a specification, and no specification exists that prefix caching's variable logits violate, according to an analysis of vLLM and SGLang issues. The vLLM PR #34046 adds an opt-i…

21:01
2026-08-18
pub.towardsai.net
artificial-intelligence

What a Kernel Is, and Why Everyone Is Writing New Ones

A kernel is a single function that runs on a GPU, and the steep memory hierarchy—with a 1,600x gap between L2 cache and HBM—makes kernel design crucial for AI performance. FlashAttention, developed by…

14:30
2026-08-18
hiraditya.github.io
artificial-intelligence

The KV Cache Has No ABI

The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…

10:11
2026-08-05
byteiota.com
artificial-intelligence

Kimi K3 Open Weights: Self-Hosting Reality Check

Moonshot AI released the weights for Kimi K3 on July 27, a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window, scoring third globally behind Claude Fable 5 Max and G…

03:01
2026-07-07
lmsys.org
ai-agents

Agent-Assisted SGLang Development: An Initial Exploration

SGLang development is being augmented with agent-assisted workflows that encode procedural engineering knowledge into executable skills, covering LLM serving, GPU kernels, diffusion pipelines, and pro…

21:25
2026-07-02
developer.nvidia.com
ai-safety

Hardware-Rooted AI Security That Won’t Slow You Down

NVIDIA announced that its Confidential Computing technology for Blackwell GPUs achieves up to 98% of the inference performance of non-secure solutions, enabling hardware-rooted AI security without sig…

03:12
2026-05-27
metaworld.me
ai-research

Finding deadlocks in CuTe kernels with SPIN

Researchers at the FlashInfer MLSYS Challenge developed a formal verification method using the SPIN model checker to detect deadlocks in CuTe DSL kernels running on NVIDIA B200 GPUs. The approach, dem…

// co-occurs with top 8 entities
// topics top 6 topics