cd/entity/FlashAttention-4· home entities FlashAttention-4
grep -l @flashattention-4 /news/*.json | wc -l → 9

FlashAttention-4

mentions 9 type Organization feed RSS

// recent coverage 9 mentions

17:59
2026-09-07
research.colfax-intl.com
machine-learning

Optimization diaries: S/P ping-pong for FlashAttention-4 decode

A new optimization for FlashAttention-4 (FA4) decoding on NVIDIA Blackwell GPUs overlaps softmax and matrix multiplication operations by using spare tensor memory (TMEM) in a ping-pong scheme, achievi…

12:30
2026-09-02
iaroslavelistratov.github.io
artificial-intelligence

B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams

A new technical blog post by Iaroslav Elistratov presents a visual guide to building a B200 attention kernel from scratch in CUDA and PTX, achieving 94.4% of FlashAttention-4 performance on 4K, 8K, an…

00:00
2026-07-05
jasonrobert.dev
artificial-intelligence

News Summary for July 5, 2026

AI agent infrastructure is maturing but facing reliability challenges, as FlashAttention-4 achieves 71% utilization on NVIDIA's Blackwell B200 GPUs while agentic AI systems grapple with architecture q…

15:23
2026-06-19
research.colfax-intl.com
large-language-models

FlashAttention-4: Algorithm and Kernel Pipelining Co-Design

Researchers from Princeton University, Together AI, Meta, Colfax Research, NVIDIA, and Georgia Tech introduced FlashAttention-4, an algorithm and kernel co-design that optimizes attention for Blackwel…

23:22
2026-06-11
modal.com
large-language-models

Making FlashAttention-4 faster for inference

Modal AI engineers Charles Frye and David Wang optimized FlashAttention-4 for large language model inference, focusing on decode-heavy workloads dominated by memory bandwidth-limited token generation.…

// co-occurs with top 8 entities
// topics top 6 topics