cd/entity/cuBLAS· home› entities› cuBLAS
grep -l @cublas /news/*.json | wc -l → 18

cuBLAS

mentions 18 type Organization feed RSS

// recent coverage 18 mentions

22:39
2026-10-08
yang-yifan.github.io
ai-infrastructure

How to Make a Compute-Bound Problem Actually Compute-Bound

Anthropic member of technical staff Yifan Yang reported that a CUDA-core fp32 GEMM kernel on an Nvidia V100 GPU reached 10.7 TFLOPS, or 88% of cuBLAS's 12.1 TFLOPS, using only two optimizations over a…

22:51
2026-09-29
zhang677.github.io
artificial-intelligence

PTXBench: What about just CUDA-PTX?

PTXBench, a benchmark released August 17, 2026 with code on GitHub, evaluates how well frontier LLMs can write architecture-specific CUDA-PTX for H100 and B200 GPUs, finding GEMM nearly solved while a…

17:19
2026-09-13
github.com
ai-infrastructure

OpenGEMM: Open-source B200 GEMM kernels

OpenGEMM, an open-source project from developer aramesh10, released CUDA GEMM kernels for NVIDIA's B200 GPU (sm_100a), installable via pip and requiring PyTorch 2.8+ and CUDA 12.9+. The library emits …

07:00
2026-08-25
hiraditya.github.io
artificial-intelligence

The Illusion of Determinism in Disaggregated Inference

SGLang and vLLM, the two major LLM serving engines, have attempted to enforce deterministic inference but face fundamental challenges due to floating-point non-associativity in GEMM kernels, which cau…

07:00
2026-08-24
hiraditya.github.io
artificial-intelligence

A Bug Is a Violation of a Specification

A bug is a violation of a specification, and no specification exists that prefix caching's variable logits violate, according to an analysis of vLLM and SGLang issues. The vLLM PR #34046 adds an opt-i…

00:02
2026-08-10
martinkristiansen.com
machine-learning

Optimizing a GPT-2-Class Transformer on a GPU

A developer's optimization campaign on an RTX 3080 Ti cut a GPT-2-small-class transformer's forward pass from 78.2ms to 1.60ms, a 49× speedup, beating torch.compile's 1.72ms and reaching 136,000 token…

19:00
2026-07-21
hiraditya.github.io
machine-learning

How torch.compile Actually Works

Torch.compile is not a traditional compiler but a system that intercepts Python bytecode at runtime, extracts compilable regions through a multi-stage pipeline, and stitches them back with eager Pytho…

16:53
2026-07-13
tornadovm.org
developer-tools

What If Java Apps Could Access CUDA Ecosystem Gracefully

TornadoVM now natively integrates NVIDIA CUDA libraries cuBLAS, cuFFT, and cuDNN, allowing Java applications to call GPU-accelerated linear algebra and deep learning primitives directly from JIT-compi…

06:33
2026-06-17
arxiv.org
machine-learning

Fearless Concurrency on the GPU

Researchers introduced cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust that extends Rust's ownership discipline to GPU kernels. On the NVIDIA B200 GPU, cuTile Rust ac…

// co-occurs with top 8 entities
// topics top 6 topics