cd /news/artificial-intelligence/when-human-knowledge-has-been-exhaus… · home topics artificial-intelligence article
[ARTICLE · art-88664] src=news.northeastern.edu ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

When human knowledge has been exhausted, where will AI get its data?

Graduate student Harsh Raj of Northeastern University's Khoury College of Computer Sciences is working on co-ops at Scale AI and Bespoke Labs to find data sources that could train large language models to surpass human knowledge, potentially matching Nobel Prize or Fields Medal winners. Raj creates synthetic data and analyzes LLM limitations, guided by professor David Bau, who secured a $9 million National Science Foundation grant for the National Deep Inference Fabric project.

read5 min views1 publishedAug 7, 2026
When human knowledge has been exhausted, where will AI get its data?
Image: News (auto-discovered)

Graduate student Harsh Raj explores cutting-edge LLMs in search of data that could one day help them think like a Nobel Prize winner.

As AI large language models, or LLMs, grow in power and sophistication, where will their architects turn when, someday — as experts predict — algorithms outgrow the limits of general human knowledge and begin craving information possessed by only the world’s most elite thinkers and creators?

Worrying as it might seem, this quest for the next frontier of data is already underway, spearheaded by research engineers in the technology industry, according to Harsh Raj, a graduate student in Northeastern University’s Khoury College of Computer Sciences.

Raj is one such intrepid explorer. As part of his two, high-profile co-ops, Raj has plumbed the depths of the most cutting-edge LLMs in search of data that could one day help them think like a winner of a Nobel Prize or Fields Medal — the highest prize for someone in mathematics — he said.

LLMs are a type of artificial intelligence trained on massive amounts of data so that they can understand and create human-like writing. They can process text, summarize lengthy documents or write code in accordance with specific prompts fed to them by users.

“It’s very hard to collect data which trains the model to be better than humans,** **because there are very few humans who can create that data,” Raj said. “You want to make it better than a Fields medalist or a Nobel Prize winner. How do you collect that?”

That question has become a focus for Raj. Now a research engineer wrapping up experiential learning at Scale AI, Raj has worked to understand the limitations and failings of LLMs in order to offer the most viable paths toward super-human knowledge capabilities. His specialized background in innovative AI projects has made him an attractive candidate for technology companies during his graduate studies.

Though Raj was courted by several technology companies hoping to leverage his skills, he chose to pursue his first co-op with Bespoke Labs, a Bay Area company that creates learning environments in which AI agents can be trained.

While there, he worked alongside other researchers to understand the fail points of models hosted on cloud computing infrastructure and offer specific solutions, he said. He also created “synthetic” data which, unlike human-created writing or code, is generated artificially before being reused as training material for the LLMs, Raj said. Creating synthetic data that matches the quality of human-made data is a tall task, Raj said. Computing power can be acquired relatively easily, but finding the data to support the project is uncharted territory, he added.

“Harsh is a self-starter, very enterprising,” said David Bau, Northeastern professor of machine learning and a prominent figure in the field of artificial intelligence. “He went on to his new opportunities on his own, and the success he found in his co-op is all his doing.”

Raj was inspired to enroll at Northeastern in large part because of Bau, he said. In addition to his work as a software engineer during the early days of Microsoft and Google, Bau in 2024 secured a $9 million grant from the U.S. National Science Foundation to launch National Deep Inference Fabric, a research computing infrastructure project at Northeastern to examine the mysteries of large-scale AI systems. Raj trusted that Bau’s expertise would offer him the best opportunity to expand his knowledge and further his career.

And that’s what happened. The Bau Lab, located on Northeastern’s Boston campus, provided Raj guidance in analyzing and interpreting LLMs’ “black box,” or the internal math and reasoning that lead from a prompt to a response. Raj quickly found ways to advance the lab’s research and improve the team’s understanding of this inner reasoning.

“Normally the internal monologue of an LLM is an unstructured ‘chain of thought,’” Bau said. “While working in our lab, Harsh built ways of training models to be able to follow specific instructions to make their internal thoughts easier to process and understand.”

Raj is now working in a co-op position at Scale AI, a company that provides high-quality data and evaluations to AI labs, governments and Fortune 500 companies. His role there has him creating and publishing scientific research papers on the failure taxonomies of LLMs, or the myriad reasons an LLM fails to deliver a user’s expected result. To do so, he said, he investigates the flaws in newer models and suggests targeted improvements by experimenting with the models, reading scientific literature, attending conferences and exchanging ideas with industry experts.

Raj said attempting to map out the next frontier of AI data is as fascinating as it is challenging. As AI laboratories exhaust human knowledge, teams are looking to target the collective knowledge of entire teams of people so that LLMs can not only perform the work of a competent coder, but a full start-up staff, he said.

“You want to create data which is automating science, automating discovery,” said Raj, who dismissed concerns over AI superseding the limits of human knowledge. “There are very few people on earth who actually have experience in that … so it’s very hard to create data.”

Raj said he plans to continue this exploration after his graduation using the skills picked up from his pioneering work during his Northeastern co-ops.

“I like to think that working at the edge of scientific knowledge about LLMs helped give Harsh a bit of confidence,” Bau said, “that he could hold his own and make contributions in fundamental research.”

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @harsh raj 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-human-knowledge…] indexed:0 read:5min 2026-08-07 ·