03:37
2026-08-04
dejan.ai
machine-learning
XM on text diffusion: a small real gain, and a bigger effect we were not looking for
A 115.7M-parameter masked diffusion language model trained with Explorative Modeling (XM) at K=3 for 10,000 steps on one RTX 4090 matched the compute of a 20,000-step baseline but showed a 0.19 nats h…