cd/entity/SRAM· home› entities› SRAM
grep -l @sram /news/*.json | wc -l → 10

SRAM

mentions 10 type Organization feed RSS

// recent coverage 10 mentions

14:01
2026-09-14
chizkidd.github.io
artificial-intelligence

Understanding FlashAttention Pt 1: Personal Notes

A technical handbook on FlashAttention explains that the algorithm achieves wall-clock speedups by reducing data movement between GPU high-bandwidth memory (HBM) and on-chip SRAM rather than by approx…

02:32
2026-09-10
keplercompute.com
ai-chips

Kepler Compute- New memory startup for AI

Kepler Compute, a new memory startup, announced it is developing leadership memory beyond HBM and SRAM that it says can be built without dependence on EUV lithography, with sampling planned for 2026 a…

17:48
2026-08-14
siliconcodesign.com
ai-infrastructure

A Deep Dive into SRAM: The Staging Ground of LLM Inference

SRAM serves as the staging ground for KV Cache values written from HBM in LLM inference, and its size directly determines the maximum local data available to local compute, according to a technical de…

19:28
2026-07-26
spectrum.ieee.org
artificial-intelligence

Optical Memory Link Could Boost AI in Robotics

Cornell Tech researchers have developed an optical memory link that uses QR-code-like light patterns to directly edit static random-access memory (SRAM) on AI processors, eliminating power-hungry anal…

15:03
2026-06-27
devclubhouse.com
ai-chips

IBM's 0.7nm Breakthrough and the Future of AI Compute

IBM Research announced the world's first sub-1 nanometer chip technology at the 0.7nm node on June 25, 2026, using a vertically stacked 3D nanosheet architecture called 'nanostack' to overcome quantum…

14:01
2026-06-26
pub.towardsai.net
large-language-models

Flash Attention Mechanics: How Tiled Attention Fits in SRAM

A new technique called Flash Attention uses tiled attention to fit the N×N attention matrix into SRAM, reducing memory reads/writes and speeding up self-attention in transformers.…

18:57
2026-06-16
injuly.in
large-language-models

Inference cost at scale with napkin math

A technical analysis calculates the dollar cost per user for serving large language models at scale using napkin math, breaking down GPU resources, matrix multiplication costs, and attention mechanism…

// co-occurs with top 8 entities
// topics top 6 topics