cd/entity/FlashAttention· home› entities› FlashAttention
grep -l @flashattention /news/*.json | wc -l → 36

FlashAttention

mentions 36 type Organization page 1/2 feed RSS

// recent coverage 36 mentions

05:48
2026-09-30
arxiv.org
artificial-intelligence

Compiling Triton kernels without the Triton compiler

A September 29, 2026 arXiv paper reports that an LLM agent translating Triton kernels directly into NVIDIA PTX, a process the authors call "AI lowering," achieved 0.83x to 3.34x the performance of aut…

14:01
2026-09-14
chizkidd.github.io
artificial-intelligence

Understanding FlashAttention Pt 1: Personal Notes

A technical handbook on FlashAttention explains that the algorithm achieves wall-clock speedups by reducing data movement between GPU high-bandwidth memory (HBM) and on-chip SRAM rather than by approx…

18:26
2026-09-08
mapathak-commits.github.io
large-language-models

What Happens Inside an LLM

Manas Pathak's primer 'What Happens Inside an LLM' explains the internal workings of a large language model during a forward pass, detailing how tokens are converted into vectors, processed through tr…

20:31
2026-09-07
sslog.dpdns.org
developer-tools

Programming an attention kernel in Triton

A developer documented writing GPU kernels in Triton to understand PyTorch operations, starting with vector add and fused ReLU/dropout, highlighting the benefit of fused kernels in reducing memory rou…

21:01
2026-08-18
pub.towardsai.net
artificial-intelligence

What a Kernel Is, and Why Everyone Is Writing New Ones

A kernel is a single function that runs on a GPU, and the steep memory hierarchy—with a 1,600x gap between L2 cache and HBM—makes kernel design crucial for AI performance. FlashAttention, developed by…

14:30
2026-08-18
hiraditya.github.io
artificial-intelligence

The KV Cache Has No ABI

The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…

07:33
2026-08-10
htor.inf.ethz.ch
machine-learning

Let's Stop Calling Everything "Linear Attention"

A new critique argues that the term 'linear attention' is overused and ambiguous, covering at least six different resource claims that range from constant computation per token to merely reduced KV ca…

06:39
2026-07-31
github.com
artificial-intelligence

HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA

HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and relate…

21:29
2026-07-30
wheels.astral.sh
developer-tools

Astral GPU Indexes

Astral has launched pre-built GPU-enabled Python wheels for the PyTorch ecosystem, supporting packages like FlashAttention across multiple CUDA versions (11.8 through 13.2) and PyTorch versions. The A…

17:46
2026-07-23
promptcube3.com
artificial-intelligence

Mage-Flow: 4B Params vs 32B Giants

Microsoft's Mage-Flow image generation model, with only 4 billion parameters, outperforms larger models including FLUX.2-dev (32B) and Qwen-Image (20B) on GenEval benchmarks, scoring 0.88 against thei…

19:13
2026-07-22
dev.to
large-language-models

RoPE: How 2D Rotations Solved Transformer Long-Context

Rotary Position Embedding (RoPE), introduced by Su et al. in 2021, solves the Transformer long-context problem by rotating Query and Key vectors in 2D sub-planes, making attention depend only on relat…

16:08
2026-07-21
byteiota.com
machine-learning

PyTorch 2.13: FlexAttention on Apple Silicon Is 12x Faster

PyTorch 2.13, released July 8, brings FlexAttention to Apple Silicon with up to 12x speedup on sparse attention patterns, such as a 32,768-token sequence with a 256-token sliding window (35ms vs 431ms…

10:24
2026-07-13
machinebrief.com
artificial-intelligence

STEEL: Revolutionizing Energy-Efficient AI on Laptops

Researchers have introduced STEEL, the first open-source implementation of FlashAttention optimized for XDNA-like neural processing units (NPUs), achieving significant energy and speed advantages for …

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics