DeepMind alumni are building an AI agent that actually Faraday, an AI agent built by Inherent, a startup founded by DeepMind alumni, has outperformed models from Anthropic and OpenAI in replicating scientific research findings, according to the Inherent team. The agent is designed to autonomously handle environment setup, data acquisition, code execution, and parameter tuning, positioning it as a specialized 'lab assistant' rather than a general-purpose chatbot. DeepMind alumni are building an AI agent that actually Faraday isn't just another LLM wrapper designed to summarize text. It is positioned as an AI "teammate" specifically engineered for scientific replication. The core claim from the Inherent team is that Faraday has already outperformed models from industry giants like Anthropic and OpenAI when it comes to the specific task of replicating research findings. This is a significant distinction because while Claude /en/tags/claude/ or GPT-4 are incredible at reasoning and coding, they often lack the specialized workflow required to navigate the messy, iterative process of scientific experimentation. Why replication is the ultimate benchmark In the current AI landscape, we see a lot of "intelligence" being measured by how well a model can pass a bar exam or write a poem. But for the scientific community, the real metric of utility is whether an agent can execute a complex, multi-step technical workflow. Replication requires: Environment Setup: Configuring specific software dependencies, hardware requirements, and OS environments. Data Acquisition: Finding, cleaning, and structuring the exact datasets mentioned in a paper. Code Execution and Debugging: Running experimental code, encountering errors, and having the autonomy to fix them without human hand-holding. Parameter Tuning: Adjusting variables to match the reported results in the original study. If Faraday can handle these steps autonomously, it moves the needle from "AI as a chatbot" to "AI as a lab assistant." The competitive edge over OpenAI and Anthropic It is telling that Inherent is targeting the "agentic" side of science rather than just general reasoning. While OpenAI and Anthropic are pouring billions into making models smarter and more conversational, Inherent is betting on specialized workflows. General-purpose models often struggle with "hallucinating" code that looks correct but fails in a real-world terminal, or they lose track of long-running experimental processes. By focusing on the specific constraints of scientific research, Faraday seems to be utilizing a more robust AI workflow that integrates better with real-world tools. This isn't just about having a bigger parameter count; it's about how the agent interacts with a terminal, a file system, and a scientific codebase. The implications for a rapid deployment of research are huge. If we can automate the verification of existing literature, we can accelerate the pace of discovery by ensuring the "ground truth" is established much faster. It turns the research process into something much more streamlined, where the AI handles the tedious reconstruction and the human scientists focus on the high-level hypothesis and novel directions. It will be interesting to see how this specialized approach holds up as the bigger players try to integrate more agentic capabilities into their mainstream models. LLM watermarking isn't about visible text or hidden ads 21h ago /en/news/7418/ Carlsen is suing OpenAI over copyright issues with NEINhorn 1d ago /en/news/7399/ Anthropic might list AI backlash as a major risk in their IPO 1d ago /en/news/7337/ Is the AI rally a genuine productivity boom or a 2d ago /en/news/7282/ Anthropic quietly rewrites its enterprise data retention rules 2d ago /en/news/7199/ The consciousness debate feels like a distraction tactic 2d ago /en/news/7184/ Next Windows on ARM runs x86 apps poorly? Here's how I fixed mine → /en/news/7513/ an AI side-hustle playbook https://tanyan888.com/ , with plenty of directly applicable cases.