cd/entity/FineWeb-Edu· home entities FineWeb-Edu
grep -l @fineweb-edu /news/*.json | wc -l → 12

FineWeb-Edu

mentions 12 type Organization feed RSS

// recent coverage 12 mentions

14:00
2026-08-25
dev.to
machine-learning

A Better FP4 Gradient Quantizer That Training Couldn't Notice

A developer found a scale-selection rule for NVFP4 gradient quantization that reduces mean squared error by 14% on real gradient tensors compared to the state-of-the-art MS-EDEN estimator, but trainin…

17:09
2026-08-16
sourcefeed.dev
artificial-intelligence

Fine-Tuning Can't Teach What Pretraining Never Saw

Researchers at the Max Planck Institute for Intelligent Systems, ELLIS Tübingen, and ETH Zürich built LittleLearner, a 5B-parameter language model pretrained exclusively on 88 billion tokens of U.S. e…

09:08
2026-08-16
sourcefeed.dev
artificial-intelligence

Pretraining Sets a Ceiling Post-Training Can't Break

A new controlled study from the LittleLearner project, led by researchers including Ryan Cotterell of ETH Zürich and Wieland Brendel of the Max Planck Institute for Intelligent Systems, found that pos…

07:37
2026-08-16
littlelearner-ll.github.io
large-language-models

What happens when an LLM never sees material beyond fifth grade?

Researchers at the University of Zurich released LittleLearner, a family of language models (0.6B, 1.3B, and 5B parameters) trained from scratch on LittleCurriculum, an 88B-token corpus filtered to U.…

04:00
2026-07-17
machinebrief.com
artificial-intelligence

Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers

A 135M-class depth-recurrent transformer trained on FineWeb-Edu converges to a per-token fixed point, with mean successive-output KL divergence falling from 3.9e-1 at the second loop to 8.5e-6 by the …

21:58
2026-07-12
github.com
artificial-intelligence

I trained a 113M-parameter earthquake LLM from absolute scratch

A developer trained a 113M-parameter language model for earthquake science from scratch using open-access papers, Wikipedia, and FineWeb-Edu on two NVIDIA A30 GPUs. The model achieved a 35% reduction …

15:50
2026-06-03
research.nvidia.com
machine-learning

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Researchers introduced Gated DeltaNet-2, a linear attention model that decouples the erase and write operations in recurrent state updates using separate channel-wise gates. The model outperforms Mamb…

00:00
2026-05-25
hash.dev
machine-learning

Deriving slugs from embeddings: vec2slug

Researchers at HASH built a 25M-parameter transformer decoder, vec2slug, that generates URL slugs directly from vector embeddings, achieving 14-19x faster and 85x cheaper performance than a Haiku-clas…

// co-occurs with top 8 entities
// topics top 6 topics