cd/entity/Transformer· home entities Transformer
grep -l @transformer /news/*.json | wc -l → 92

Transformer

mentions 92 type Organization page 3/5 feed RSS

// recent coverage 92 mentions

07:42
2026-07-07
blog.stackademic.com
large-language-models

How Transformers Actually Work — No Math, Just the Mental Model

Transformers, the architecture behind GPT and all large language models, are explained using a code review analogy instead of dense mathematics. The article argues that traditional sequential reading …

04:00
2026-07-07
letsdatascience.com
machine-learning

GestaltMML improves rare disease diagnosis accuracy

Researchers revised the GestaltMML paper on arXiv, describing a Transformer-based multimodal model that combines frontal facial images, demographics, and clinical text to improve rare genetic disease …

08:28
2026-07-01
promptcube3.com
artificial-intelligence

AI\'s Memory Crunch Hits Indian Smartphones

Indian smartphone manufacturers are pushing 12GB or 16GB of RAM as the new standard for mid-range devices to support on-device AI features, which consume significant memory through model weights and K…

18:30
2026-06-30
hello-fri-end.github.io
large-language-models

Why do transformers have outliers?

Transformer models develop outlier channels—feature dimensions with unusually large values in weights and activations—due to the softmax normalization in attention layers, which forces tokens to assig…

14:16
2026-06-30
ranpara.net
large-language-models

Words Are a Byproduct of Consciousness. For LLMs, It's Backwards

A writer argues that large language models (LLMs) generate words without underlying consciousness, reversing the human process where ideas precede language. The piece traces the evolution of human com…

08:04
2026-06-28
bharad.dev
large-language-models

A Transformer Becomes an LLM

A transformer architecture becomes a large language model through training on trillions of tokens, using residual connections to enable deep networks, tokenization to split text into manageable units,…

11:56
2026-06-27
arxiv.org
large-language-models

Tapered Language Models

Researchers introduced Tapered Language Models (TLMs), an architectural principle that allocates more parameter capacity to earlier layers and less to later layers under a fixed budget. Across four ar…

03:10
2026-06-20
devclubhouse.com
artificial-intelligence

Why Noam Shazeer’s Move to OpenAI Reshapes the Frontier

Noam Shazeer, co-author of the seminal Transformer paper and a key engineer at Google, has left for OpenAI less than two years after Google spent $2.7 billion to reacquire him from Character.AI. The m…

15:54
2026-06-19
people.idsia.ch
artificial-intelligence

Munich 1991: The Roots of the Current AI Boom

The foundations of today's AI boom, including the first Transformer variant, unsupervised pre-training, neural network distillation, and deep residual learning, were all published in 1991 by Jürgen Sc…

14:31
2026-06-18
pub.towardsai.net
machine-learning

Nested Learning: One Memory, Many Clocks

The Continuum Memory System introduces a novel architecture that replaces the Transformer's two-rate memory with a spectrum of blocks, each updating on its own clock, potentially improving efficiency …

← prev page 3 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics