>10x More Efficient Pretraining Magic reported that its pretraining recipe is now more than 10x more compute-efficient than leading open-weight base models, matching DeepSeek V4 Pro Base with roughly 50x fewer FLOPs — about half of GPT-3's pretraining compute, or around $0.5M on GB200. Magic continued scaling 10x (about $4M) and said it meaningfully outperformed all publicly available open base models on perplexity evals, while noting that training a model this capable would cost over $100M under DeepSeek V4 Pro's recipe. The company said it believes pretraining, agentic RL, and long-context are sufficient to build superhuman coding agents and automate AI R&D. 10x More Efficient Pretraining Research update on compute-efficient pretraining and scaling to trillion-parameter models. Frontier pretraining is said to be a big-lab-only game. We don’t have 100k chips yet, so there’s only one way: algorithmic efficiency. After compounding for … a while …, our pretraining recipe is now 10x more compute-efficient than that of leading open-weight base models. We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. We continued scaling 10x ~$4M and meaningfully outperformed all publicly available open base models on perplexity evals. By the scaling laws in Figure 1, training a model this capable would cost $100M under DeepSeek V4 Pro’s recipe and this is ignoring how much data exists . Of course, we won’t stop scaling there. We believe pretraining, agentic RL, and long-context are sufficient to build superhuman coding agents and automate AI R&D. We started with long-context https://magic.dev/blog/100m-token-context-windows . Today’s blog post is about pretraining. We measured bits-per-byte loss a metric that normalizes out differences in tokenizers on heldout data and fit a scaling law https://arxiv.org/abs/2203.15556 to project how much compute is needed to reach a given level of capability. Better training compute efficiency means stronger models at all budgets. We evaluated the latest available open-weight base models