cd/entity/Nvidia-Nemotron-Pretraining-Code-v2Β· homeβ€Ί entitiesβ€Ί Nvidia-Nemotron-Pretraining-Code-v2
grep -l @nvidia-nemotron-pretraining-code-v2 /news/*.json | wc -l β†’ 1

Nvidia-Nemotron-Pretraining-Code-v2

mentions 1 type Organization feed RSS

// recent coverage 1 mentions

12:00
2026-07-13
research.ibm.com
artificial-intelligence

This could be the largest synthetic code dataset yet

IBM open-sourced CodeAlchemy, a synthetic dataset of nearly 1 trillion tokens across 15 programming languages, which may be the largest synthetic code dataset to date. The dataset includes 1.3 million…

// co-occurs with top 7 entities
// topics top 5 topics