cd/entity/Transformer· home entities Transformer
grep -l @transformer /news/*.json | wc -l → 92

Transformer

mentions 92 type Organization page 2/5 feed RSS

// recent coverage 92 mentions

18:07
2026-07-27
alexzhang13.github.io
artificial-intelligence

Language model harnesses are compositional generalizers

A new study from researchers argues that language model harnesses, specifically Recursive Language Models (RLMs), can achieve compositional generalization by shaping each call to the underlying Transf…

01:00
2026-07-27
dev.to
artificial-intelligence

The Evolution of AI, Explained in Stages

A developer explains the evolution of AI through distinct stages over 70 years, from rule-based systems to modern large language models and AI agents. The post highlights how each stage removed limita…

16:46
2026-07-25
promptcube3.com
artificial-intelligence

Building a Tiny LM from Scratch in Node.js

A developer built a tiny language model from scratch in Node.js using a custom autograd engine and explicit neuron implementations, achieving correct outputs on basic logic tests after pre-training an…

10:02
2026-07-25
promptcube3.com
artificial-intelligence

Understanding Transformers: A Visual Deep Dive

A visual deep dive into Transformer architectures uses a 3D animated breakdown with a 'publishing workshop' analogy to explain tokenization, embeddings, positional encoding, the QKV engine, and networ…

20:54
2026-07-22
manish.sh
artificial-intelligence

Character.ai: The Technical Story of What Keeps Users Hooked

Character.ai, co-founded by former Google engineers Noam Shazeer and Daniel De Freitas, keeps users hooked through a combination of technical optimizations including its Kaiju dense transformer models…

04:00
2026-07-22
machinebrief.com
artificial-intelligence

Dual Attention Residuals

Researchers propose Dual Attention Residuals (DAR), a method that brings multi-stream interaction into historical retrieval for Transformer models through reciprocal cross-stream addressing. Across de…

13:56
2026-07-21
arxiv.org
artificial-intelligence

Loop the Loopies

Researchers have introduced the Loopie series, two Mixture-of-Experts models (20B parameters with 2B active and 6B with 0.6B active) that outperform vanilla Transformer baselines trained with the same…

02:58
2026-07-16
discuss.huggingface.co
machine-learning

My latest ablation run: integrating Engram onto two backbones

A 200-step, ~1.7B-parameter ablation run in OLMo-core comparing Engram on a standard attention Transformer versus a 3 GDN-layer + 1 attention-layer hybrid found that Transformer + Engram reached sligh…

13:10
2026-07-14
machinebrief.com
machine-learning

DAG-FM: The New Heavyweight in Causal Discovery

DAG-FM, a new model using Transformers and a Mixture-of-Leaf-Experts mechanism, outperforms traditional algorithms and other foundation models in causal discovery from observational data, achieving hi…

07:39
2026-07-11
machinebrief.com
artificial-intelligence

Transformer Efficiency: A Closer Look at KV Cache Compression

New research characterizes the intrinsic compressibility of KV caches in Transformer models, proposing a principled algorithm for efficient inference. The study uses minimax risk to determine when acc…

03:26
2026-07-10
sbondaryev.dev
artificial-intelligence

How a Transformer Plays Tic-Tac-Toe

An interactive guide demonstrates how Transformer models use attention, embeddings, and positional encoding to predict moves in a fading Tic-Tac-Toe game, making the architecture's inner workings acce…

15:31
2026-07-07
blog.bytebytego.com
large-language-models

ChatGPT vs Gemini vs Claude: How They Differ

OpenAI's ChatGPT, Google's Gemini, and Anthropic's Claude differ in behavior due to distinct architectural choices made by their developers, including decisions on model density, training methods, and…

10:54
2026-07-07
john463212.substack.com
large-language-models

Building a 350M Transformer from Scratch in PyTorch

A developer built a 350-million-parameter Transformer model from scratch using PyTorch, detailing the architecture, training process, and performance benchmarks. The project demonstrates how to implem…

← prev page 2 / 5 next →
// co-occurs with top 8 entities
// topics top 6 topics