cd/entity/Transformer· home› entities› Transformer
grep -l @transformer /news/*.json | wc -l → 124

Transformer

mentions 124 type Organization page 4/7 feed RSS

// recent coverage 124 mentions

02:58
2026-07-16
discuss.huggingface.co
machine-learning

My latest ablation run: integrating Engram onto two backbones

A 200-step, ~1.7B-parameter ablation run in OLMo-core comparing Engram on a standard attention Transformer versus a 3 GDN-layer + 1 attention-layer hybrid found that Transformer + Engram reached sligh…

13:10
2026-07-14
machinebrief.com
machine-learning

DAG-FM: The New Heavyweight in Causal Discovery

DAG-FM, a new model using Transformers and a Mixture-of-Leaf-Experts mechanism, outperforms traditional algorithms and other foundation models in causal discovery from observational data, achieving hi…

07:39
2026-07-11
machinebrief.com
artificial-intelligence

Transformer Efficiency: A Closer Look at KV Cache Compression

New research characterizes the intrinsic compressibility of KV caches in Transformer models, proposing a principled algorithm for efficient inference. The study uses minimax risk to determine when acc…

03:26
2026-07-10
sbondaryev.dev
artificial-intelligence

How a Transformer Plays Tic-Tac-Toe

An interactive guide demonstrates how Transformer models use attention, embeddings, and positional encoding to predict moves in a fading Tic-Tac-Toe game, making the architecture's inner workings acce…

15:31
2026-07-07
blog.bytebytego.com
large-language-models

ChatGPT vs Gemini vs Claude: How They Differ

OpenAI's ChatGPT, Google's Gemini, and Anthropic's Claude differ in behavior due to distinct architectural choices made by their developers, including decisions on model density, training methods, and…

10:54
2026-07-07
john463212.substack.com
large-language-models

Building a 350M Transformer from Scratch in PyTorch

A developer built a 350-million-parameter Transformer model from scratch using PyTorch, detailing the architecture, training process, and performance benchmarks. The project demonstrates how to implem…

07:42
2026-07-07
blog.stackademic.com
large-language-models

How Transformers Actually Work — No Math, Just the Mental Model

Transformers, the architecture behind GPT and all large language models, are explained using a code review analogy instead of dense mathematics. The article argues that traditional sequential reading …

04:00
2026-07-07
letsdatascience.com
machine-learning

GestaltMML improves rare disease diagnosis accuracy

Researchers revised the GestaltMML paper on arXiv, describing a Transformer-based multimodal model that combines frontal facial images, demographics, and clinical text to improve rare genetic disease …

08:28
2026-07-01
promptcube3.com
artificial-intelligence

AI\'s Memory Crunch Hits Indian Smartphones

Indian smartphone manufacturers are pushing 12GB or 16GB of RAM as the new standard for mid-range devices to support on-device AI features, which consume significant memory through model weights and K…

18:30
2026-06-30
hello-fri-end.github.io
large-language-models

Why do transformers have outliers?

Transformer models develop outlier channels—feature dimensions with unusually large values in weights and activations—due to the softmax normalization in attention layers, which forces tokens to assig…

14:16
2026-06-30
ranpara.net
large-language-models

Words Are a Byproduct of Consciousness. For LLMs, It's Backwards

A writer argues that large language models (LLMs) generate words without underlying consciousness, reversing the human process where ideas precede language. The piece traces the evolution of human com…

08:04
2026-06-28
bharad.dev
large-language-models

A Transformer Becomes an LLM

A transformer architecture becomes a large language model through training on trillions of tokens, using residual connections to enable deep networks, tokenization to split text into manageable units,…

← prev page 4 / 7 next →
// co-occurs with top 8 entities
// topics top 6 topics