cd /news/artificial-intelligence/why-ilya-sutskever-says-pre-training… · home topics artificial-intelligence article
[ARTICLE · art-80614] src=mindstudio.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Why Ilya Sutskever Says Pre-Training Is Over and Research Is Back

Ilya Sutskever, co-founder of Safe Superintelligence, argued at NeurIPS in December 2024 that the scaling era of AI pre-training is ending because high-quality human-generated data is finite, not because compute has plateaued. He predicts the field is entering a new research-driven phase where generalization and brain-inspired learning will be key to advancing toward AGI. Sutskever's claims are backed by Safe Superintelligence's reported $5 billion equity deal with Nvidia, which signals a shift toward scaling novel research rather than larger Transformers.

read8 min views1 publishedJul 30, 2026
Why Ilya Sutskever Says Pre-Training Is Over and Research Is Back
Image: Mindstudio (auto-discovered)

Ilya Sutskever argues AI's scaling era is giving way to research again. Here's what that means for generalization, brain-inspired learning, and AGI.

What did Ilya Sutskever actually say about pre-training? #

Ilya Sutskever’s core claim is that the scaling recipe behind GPT-3 and GPT-4 (more data, more compute, bigger training runs) is running out of runway, not because compute stopped improving but because high-quality human-generated data hasn’t kept pace. He put it bluntly at NeurIPS in December 2024: there is effectively only one internet. That means the fuel supply for pre-training is finite in a way that raw compute isn’t. His argument was never that AI progress stops. It was that the field needs a different recipe to keep advancing, and finding that recipe requires research, not just bigger clusters.

TL;DR #

  • Sutskever’s argument is that pre-training scaling worked because compute and data grew together, but usable high-quality text data is now the bottleneck, not hardware. - He splits recent AI history into two phases: an age of researchers(roughly 2012 to 2020) when architectures were experimental, and an** age of scaling**(2020 to 2025) when one recipe just got bigger. - His Dwarkesh Patel interview in November 2025 argued the field is swinging back into a research-driven phase, except now with much larger computers behind it. - The biggest technical gap he points to is generalization: humans learn from few examples and transfer that understanding broadly, while AI systems ace benchmarks and still trip on situations that seem obvious to people. - He describes superintelligence less as a finished all-knowing system and more as a continual learner that adapts quickly inside new environments after deployment. - Safe Superintelligence’s deal with Nvidia, reportedly including a $5 billion equity investment and roughly a tenfold compute increase, signals that SSI believes its internal research has crossed a threshold worth scaling. - Nvidia’s own announcement ties Sutskever’s research to overlooked mechanisms of human brain function, hinting SSI is chasing something beyond making Transformers bigger.

Remy doesn't write the code. It manages the agents who do. #

Remy runs the project. The specialists do the work. You work with the PM, not the implementers.

Why does Sutskever think the “scaling era” is ending? #

Sutskever’s case rests on a simple resource mismatch. Training runs like GPT-3 and GPT-4 improved predictably as labs threw more parameters, more compute, and more data at the same basic architecture. That worked because all three inputs could grow together. But the internet’s supply of high-quality, human-generated text isn’t expanding at the same rate compute is. Once you’ve trained on most of what’s available, adding more GPUs doesn’t create new data out of nothing.

Sutskever has been careful to clarify that this isn’t a claim that scaling stopped mattering. Making current models larger can still produce gains. His point is narrower and, in some ways, more unsettling for the industry: the recipe that got AI this far probably can’t be the same recipe that gets it to the next major leap. Something else has to be figured out first, and that something is a research problem, not an infrastructure problem.

What is the “age of research” he’s describing? #

In the Dwarkesh interview, Sutskever framed recent AI history in two stages. From roughly 2012 to 2020, the field was in an age of researchers, when different groups tried different architectures, training objectives, and ideas without a single dominant approach. Then, once transformer-based scaling showed it reliably improved with more data and compute, the field entered an age of scaling, roughly 2020 to 2025, where the incentive was simply to build bigger versions of a known-working system.

His prediction is that AI is moving back into a research-heavy period, but with a twist: this time the researchers doing the experimenting have access to far larger computers than the researchers of the early 2010s ever did. That combination, small teams doing genuine exploratory research backed by industrial-scale compute, is the model he’s built Safe Superintelligence around.

Why does Sutskever think generalization is the missing piece? #

The gap he keeps returning to is generalization. Humans can encounter a handful of examples of something new and apply that understanding in a completely different context almost immediately. Current AI systems, by contrast, can dominate difficult academic benchmarks while failing at tasks that would be trivial for a person. That inconsistency suggests these systems aren’t learning the way humans do. They’re pattern-matching extremely well within the distribution of their training data, but the underlying efficiency of human learning, the ability to extract a general rule from very little evidence, hasn’t been replicated.

Sutskever’s description of what superintelligence might actually look like follows from this. Rather than picturing a static, finished system that instantly knows everything, he describes something closer to a powerful continual learner: a system that can walk into an unfamiliar environment, pick up what it needs quickly, and keep improving after it’s already been deployed. That’s a meaningfully different target than “a bigger version of today’s chatbot.” It implies the missing ingredient isn’t scale but a mechanism, something closer to how biological learning works.

Is Safe Superintelligence’s Nvidia deal proof that Sutskever found something? #

Everyone else built a construction worker.

We built the contractor.

One file at a time.

UI, API, database, deploy.

Safe Superintelligence was built specifically to avoid the distractions Sutskever sees pulling at bigger AI labs: inference costs, product engineering, sales, marketing, and pricing strategy. Since its founding in 2024, after Sutskever left OpenAI where he had been co-founder and chief scientist, SSI has said almost nothing publicly about its technical direction. Its stated goal has been singular: build safe superintelligence, not an assistant product or an API business.

That’s what makes the reported terms of its new partnership with Nvidia notable. Nvidia is reportedly investing around $5 billion in SSI, and the arrangement is expected to increase SSI’s available compute by roughly ten times over the next year using Nvidia’s next-generation systems. Nvidia’s own announcement states it was given access to SSI’s closely guarded research before entering the deal, and describes SSI as having spent two years pursuing a research direction aimed at powerful, reliably aligned AI.

None of this confirms SSI has solved generalization or built anything resembling superintelligence. But it is a strong external signal, arguably the strongest so far, that something inside the company moved from “we’re still searching” to “we found a direction worth scaling.” Sutskever’s own language shifted accordingly: from insisting the old recipe wasn’t enough, to saying the research has reached a point where scaling it makes sense.

What might SSI actually be building? #

The clearest hint so far comes from Nvidia’s announcement itself, which ties SSI’s research to overlooked aspects of how the human brain functions. That phrasing suggests SSI isn’t simply trying to make Transformer models larger. It points toward mechanisms that let humans learn quickly, generalize across contexts, and adapt continuously, the exact properties Sutskever flagged as missing from current systems in his Dwarkesh interview.

If that’s the direction, the sequence Sutskever laid out starts to make sense: the old scaling recipe hit its ceiling, the field needed to go back to research, SSI spent roughly two years searching for a new mechanism, and it now believes it has found something that behaves well enough to justify a tenfold compute increase. Whether that mechanism actually closes the generalization gap is unknown from the outside. What’s known is that Nvidia looked at SSI’s internal work closely enough to commit billions of dollars behind it.

Frequently Asked Questions #

Did Ilya Sutskever say scaling is completely over?

No. He specifically clarified that scaling current systems can still produce improvements. His argument is that the pre-training recipe of GPT-3 and GPT-4 era models, more data and compute on the same architecture, can’t keep producing the same magnitude of gains indefinitely because high-quality data is limited, not because compute stopped mattering.

What does Sutskever mean by “the age of research”?

He’s describing a return to the kind of experimental, architecture-searching work that characterized AI roughly between 2012 and 2020, before scaling one proven recipe became the dominant strategy from 2020 to 2025. The difference this time is that research-stage experimentation now happens with access to much larger computers.

What is generalization, and why does it matter to Sutskever?

Generalization is the ability to learn a concept from limited examples and apply it correctly in new, unrelated situations, something humans do naturally and efficiently. Sutskever argues current AI systems can excel at hard benchmarks yet fail at simple tasks outside their training distribution, and he sees closing that gap as central to reaching more general intelligence.

What has Safe Superintelligence actually announced?

Other agents ship a demo. Remy ships an app. #

Real backend. Real database. Real auth. Real plumbing. Remy has it all.

SSI announced a strategic partnership with Nvidia that reportedly includes a roughly $5 billion equity investment and an increase in SSI’s compute of about ten times over the next 12 months using Nvidia’s next-generation systems. Nvidia stated it was given access to SSI’s research before the deal, but neither company has disclosed the specific technical breakthrough behind it.

Does this mean SSI has built superintelligence?

There’s no evidence of that. What the deal suggests is that SSI’s internal research crossed a point where its team believes scaling it up is worthwhile, and that Nvidia found the underlying work credible enough to back with a major investment. The specifics of what SSI discovered remain undisclosed.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @ilya sutskever 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/why-ilya-sutskever-s…] indexed:0 read:8min 2026-07-30 ·