cd/entity/FlashAttention-3· home› entities› FlashAttention-3
grep -l @flashattention-3 /news/*.json | wc -l → 4

FlashAttention-3

mentions 4 type Organization feed RSS

// recent coverage 4 mentions

14:42
2026-08-19
promptcube3.com
ai-infrastructure

Cerebras CS-4 rack density pushes wafer-scale cooling to new

Cerebras Systems' CS-4 rack achieves 85% model flops utilization (MFU) on a 400B parameter training run across 16 systems, compared to 30-40% for traditional GPU clusters, thanks to the WSE-3's on-waf…

00:00
2026-07-15
dibi8.com
artificial-intelligence

SGLang — Structured Generation and Fast LLM Serving Engine

SGLang, an open-source LLM inference engine, introduces RadixAttention for prefix caching and grammar-constrained decoding, achieving 25x throughput improvement over vLLM for structured output tasks. …

12:32
2026-05-30
maltebuettner.eu
large-language-models

DocumentAI Visual Benchmark - GPT 5.5, Gemini 3.5, Qwen...

A new benchmark evaluating DocumentAI models on bounding box accuracy shows GPT-5.5 and Gemini 3.5 leading with 67.7% and 67.5% scores respectively, while Qwen, Kimi, and Mistral trail significantly. …

00:00
2026-05-14
maltebuettner.eu
large-language-models

documentai bbox benchmark

Malte Buettner benchmarked bounding box accuracy for Document AI models using pages from the FlashAttention-3 paper, testing Qwen, Kimi, and Mistral via OpenRouter. The evaluation scored models on cov…

// co-occurs with top 8 entities
// topics top 6 topics