cd/entity/Transformer· home› entities› Transformer
grep -l @transformer /news/*.json | wc -l → 124

Transformer

mentions 124 type Organization page 3/7 feed RSS

// recent coverage 124 mentions

18:01
2026-08-12
dev.to
machine-learning

How the Transformer Paper Came About

The Transformer architecture, introduced in the 2017 paper 'Attention Is All You Need,' was designed primarily to reduce training time by enabling parallelization, not to improve translation quality. …

12:01
2026-08-05
pub.towardsai.net
artificial-intelligence

Understanding Transformers by Building One from Scratch

A developer built a Transformer model from scratch to translate English to Sanskrit, training a custom Byte Pair Encoding tokenizer with about 5,000 tokens and a model with roughly 22 million paramete…

04:00
2026-08-05
arxiv.org
artificial-intelligence

Self-Organising Digital Circuits

Researchers introduced Self-Organising Digital Circuits, a new architecture using a topology-masked Transformer to configure Lookup Tables of Boolean gates, achieving near-perfect recovery (>99.99% ac…

21:43
2026-07-31
blog.us.fixstars.com
artificial-intelligence

How to Unlock AI Performance

Llama 3 70B, a dense Transformer with about 70 billion parameters, requires roughly 140 billion FLOPs to generate a single token and about 140 trillion FLOPs for a 1,000-token answer, while training t…

18:07
2026-07-27
alexzhang13.github.io
artificial-intelligence

Language model harnesses are compositional generalizers

A new study from researchers argues that language model harnesses, specifically Recursive Language Models (RLMs), can achieve compositional generalization by shaping each call to the underlying Transf…

01:00
2026-07-27
dev.to
artificial-intelligence

The Evolution of AI, Explained in Stages

A developer explains the evolution of AI through distinct stages over 70 years, from rule-based systems to modern large language models and AI agents. The post highlights how each stage removed limita…

16:46
2026-07-25
promptcube3.com
artificial-intelligence

Building a Tiny LM from Scratch in Node.js

A developer built a tiny language model from scratch in Node.js using a custom autograd engine and explicit neuron implementations, achieving correct outputs on basic logic tests after pre-training an…

10:02
2026-07-25
promptcube3.com
artificial-intelligence

Understanding Transformers: A Visual Deep Dive

A visual deep dive into Transformer architectures uses a 3D animated breakdown with a 'publishing workshop' analogy to explain tokenization, embeddings, positional encoding, the QKV engine, and networ…

20:54
2026-07-22
manish.sh
artificial-intelligence

Character.ai: The Technical Story of What Keeps Users Hooked

Character.ai, co-founded by former Google engineers Noam Shazeer and Daniel De Freitas, keeps users hooked through a combination of technical optimizations including its Kaiju dense transformer models…

04:00
2026-07-22
machinebrief.com
artificial-intelligence

Dual Attention Residuals

Researchers propose Dual Attention Residuals (DAR), a method that brings multi-stream interaction into historical retrieval for Transformer models through reciprocal cross-stream addressing. Across de…

13:56
2026-07-21
arxiv.org
artificial-intelligence

Loop the Loopies

Researchers have introduced the Loopie series, two Mixture-of-Experts models (20B parameters with 2B active and 6B with 0.6B active) that outperform vanilla Transformer baselines trained with the same…

← prev page 3 / 7 next →
// co-occurs with top 8 entities
// topics top 6 topics