cd/entity/CUTLASS· home entities CUTLASS
grep -l @cutlass /news/*.json | wc -l → 13

CUTLASS

mentions 13 type Organization feed RSS

// recent coverage 13 mentions

21:01
2026-08-18
pub.towardsai.net
artificial-intelligence

What a Kernel Is, and Why Everyone Is Writing New Ones

A kernel is a single function that runs on a GPU, and the steep memory hierarchy—with a 1,600x gap between L2 cache and HBM—makes kernel design crucial for AI performance. FlashAttention, developed by…

21:06
2026-07-29
github.com
artificial-intelligence

PyCuTe: Reference implementation and examples of the CuTe Layout

NVIDIA researcher Cris Cecka released PyCuTe, a pure-Python reference implementation of the CuTe layout algebra used in CUTLASS 3.x and the CuTe DSL, enabling learning, prototyping, and test-vector ge…

04:54
2026-07-11
machinebrief.com
large-language-models

Accelerating SwiGLU: A breakthrough for Large Language Models

Researchers introduced two CUTLASS-based kernels that accelerate the SwiGLU activation function in large language models by up to 2.47x on NVIDIA H100 GPUs, shifting workloads from memory-bound to com…

16:36
2026-07-07
int21.ai
ai-agents

Stop Waiting for a Bigger Context Window

INT21 has built SwarmOS, a cloud-native platform for multi-agent AI systems, arguing that orchestrating specialized agents is more effective than relying on larger context windows. The company demonst…

17:00
2026-06-10
pytorch.org
large-language-models

Portable vLLM Model Inference Kernels in Helion

Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments demonstrated that Helion provides a productive PyTorch-nat…

12:53
2026-05-29
zartbot.github.io
ai-chips

Dissecting the SM_120 Microarchitecture

NVIDIA's Blackwell consumer GPU (GB203/SM_120) features a unified TensorCore pipeline where all 12 non-FP64 precision formats share identical 29-cycle latency and 23-cycle throughput, reducing precisi…

03:12
2026-05-27
metaworld.me
ai-research

Finding deadlocks in CuTe kernels with SPIN

Researchers at the FlashInfer MLSYS Challenge developed a formal verification method using the SPIN model checker to detect deadlocks in CuTe DSL kernels running on NVIDIA B200 GPUs. The approach, dem…

// co-occurs with top 8 entities
// topics top 6 topics