cd/entity/GPT-2· home› entities› GPT-2
grep -l @gpt-2 /news/*.json | wc -l → 147

GPT-2

mentions 147 type Organization page 2/8 feed RSS

// recent coverage 147 mentions

16:10
2026-09-08
dwarkesh.com
artificial-intelligence

Pretraining progress is mostly coming from data

A study by Epoch AI finds that from 2019 to 2025, data improvements contributed 12.0x compute efficiency gains compared to 3.7x from model improvements, meaning 3.24x more gains came from data at a 1e…

00:00
2026-09-08
mindstudio.ai
artificial-intelligence

What Is an AI Agent Harness? The Scaffolding Explained

An agent harness is the code and structure around a language model that converts raw text prediction into goal-directed behavior, and it has driven larger performance gains than model upgrades, accord…

17:25
2026-09-03
gilesthomas.com
artificial-intelligence

Putting my JAX-trained models on the Hugging Face Hub

Developer gpjt has uploaded PyTorch-compatible versions of all his JAX-trained GPT-2 models to the Hugging Face Hub, including models from his blog series on writing an LLM from scratch and Chinchilla…

08:19
2026-09-03
github.com
large-language-models

Show HN: PicoLM v1.0-rc1

PicoLM v1.0-rc1, an LLM inference engine written in C99, has been released, supporting llama-2, GPT-2, Qwen 3.6/3.8(+MoE), and Gemma-3n models. The engine features CPU SIMD acceleration, CUDA/HIP supp…

00:00
2026-09-03
ben3d.ca
large-language-models

Running LLMs in the Browser with Three.js

Ben Houston's Three-LLM library runs LLMs entirely in the browser via Three.js and WebGPU, supporting GPT-2, Llama-style, Gemma 3, Phi, and Qwen3.5 architectures, with checkpoints from 3M to 1.3B para…

09:24
2026-09-02
healeycodes.com
large-language-models

What Makes LLM Tokenization Slow?

LLM tokenization, while a small part of overall latency, sits on the hot path and can occur multiple times per request, according to a technical analysis of GPT-2's reference encoder. The analysis sho…

20:55
2026-09-01
tokencontributions.substack.com
natural-language-processing

Small pre-tokenization bugs with a big multilingual price

A developer's analysis of pre-tokenization regexes shows that GPT-2's word-splitting pattern omitted Unicode's Mark category, a bug inherited by GPT-4, Llama 3, Qwen 3, and GLM-4/5, forcing BPE to tok…

20:21
2026-08-26
runtimewire.com
artificial-intelligence

Latitude opens Voyage, betting AI roleplay needs rules after all

Latitude opened its AI roleplaying platform Voyage to the public on Wednesday, removing the waitlist for the open beta available on iOS, Android, and the web. The platform separates world state manage…

14:00
2026-08-25
dev.to
machine-learning

A Better FP4 Gradient Quantizer That Training Couldn't Notice

A developer found a scale-selection rule for NVFP4 gradient quantization that reduces mean squared error by 14% on real gradient tensors compared to the state-of-the-art MS-EDEN estimator, but trainin…

13:50
2026-08-25
futurism.com
large-language-models

Even Babies Are Still Way Better at Learning Than AI Models

Even infants are far more efficient language learners than today's AI chatbots, according to experts interviewed by MIT Technology Review. Stanford cognitive scientist Michael C. Frank noted that todd…

04:00
2026-08-24
machinebrief.com
machine-learning

RODE: A Radial-Orthogonal Decoupled Engine for Optimization

Researchers introduced RODE, a matrix-aware optimizer that decouples radial and directional updates, outperforming Muon variants across language modeling and image classification tasks. At 1.5B scale,…

21:41
2026-08-23
github.com
large-language-models

Implementation of GPT-2 in pure CMake

AlpinDale released gpt2.cmake, an implementation of OpenAI's GPT-2 language model written entirely in CMake, using Q16.16 fixed-point integer arithmetic to run inference without floating-point operati…

22:01
2026-08-22
pub.towardsai.net
artificial-intelligence

One Formula to Map the Positional Encoding Landscape

A new survey of positional encoding methods in Transformers argues that the field is best understood not as a chronological progression but as answers to a single question: where position information …

09:10
2026-08-21
honnibal.dev
artificial-intelligence

I don't think AI is a bubble

In a blog post, an unnamed author argues that AI performance will not plateau, countering the common belief that progress is driven by brute-force scale with diminishing returns. The author contends t…

← prev page 2 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics