Via samsara.com
The London-based nonprofit wants to combine human and AI judgment to keep increasingly powerful models in check
Two former Google DeepMind researchers have left one of the world’s most powerful AI labs to tackle what they see as the field’s most pressing unsolved problem: figuring out who, or what, is actually qualified to judge whether an AI system is doing its job well.
Rishub Jain and Josh Jacob launched Sampura Research on August 25, a London-based nonprofit focused on developing better evaluation systems for AI. The organization has secured $11M in initial funding from Coefficient Giving, with $7M allocated for the first year and an additional $4M pledged for future use.
The judge problem #
The nonprofit calls this “Human-AI Complementarity for Scalable Oversight,” which boils down to a practical question: how do you build evaluation frameworks that scale alongside the technology they’re supposed to monitor?
Sampura’s proposed solution involves developing what the organization calls “judges,” systems composed of human evaluators, AI components, or some combination of both. These judges would be deployed across the AI lifecycle, from training to evaluation to real-world deployment, catching problems that purely automated or purely human oversight would miss.
The name itself offers a clue about the founders’ philosophy. “Sampura” derives from the Sanskrit word sampūraka, meaning “complementary.” The framing is deliberate: this isn’t about humans supervising AI or AI replacing human judgment, but about finding the specific combination where each compensates for the other’s blind spots.
From AlphaFold to alignment #
Both founders built their credentials inside DeepMind, contributing to high-profile projects including AlphaFold, the protein-structure prediction system that represented one of AI’s most celebrated scientific breakthroughs. They also worked on AI safety programs within Google’s research division.
Sampura’s work builds on research initiatives at DeepMind dating back to 2023, including the SPAR and MARS fellowship programs. Those earlier efforts explored similar questions about how to maintain meaningful human oversight as AI systems grow more autonomous and capable.
The nonprofit is currently recruiting founding technical staff with AI research experience.
Why evaluation matters more than it sounds #
One of the most persistent problems in AI development is something called reward hacking, where a model learns to game its evaluation metrics rather than actually performing well at its intended task. If a language model is rewarded for sounding confident, it learns to sound confident even when it’s wrong.
Sampura’s focus on hybrid evaluation attempts to address this by introducing evaluation frameworks that are harder for AI systems to game. A human evaluator catches different failure modes than an automated one, and vice versa.
The $11M from Coefficient Giving represents a meaningful bet on this approach. While it’s modest compared to the billions flowing into frontier model development, it’s substantial for a nonprofit research organization focused on a specific technical problem.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our