Pretraining progress is mostly coming from data
A study by Epoch AI finds that from 2019 to 2025, data improvements contributed 12.0x compute efficiency gains compared to 3.7x from model improvements, meaning 3.24x more gains came from data at a 1e…
A study by Epoch AI finds that from 2019 to 2025, data improvements contributed 12.0x compute efficiency gains compared to 3.7x from model improvements, meaning 3.24x more gains came from data at a 1e…
Researchers introduced AURORA-LM, a continuous-latent diffusion language model that achieves the strongest performance among evaluated continuous and diffusion-based language models on OpenWebText fre…
A new study introduces CaRE, a compute-aware evaluation framework for masked diffusion language models (MDLMs), revealing that current evaluation standards conflate algorithmic improvements with hidde…
A new model compression method called requential coding compresses a generative model by coding a training process built from self-generated data, achieving code length equal to the cumulative teacher…
Researchers introduce token time continuous diffusion (TTCD), a new diffusion language model that operates in continuous space and incorporates per-token times, allowing some tokens to proceed from no…
A 124M-parameter GPT-2 model trained from scratch on OpenWebText data using a custom deep learning library achieved a validation loss of 2.764 nats and a perplexity of 15.87 after 56,000 steps (27.5B …