{"slug": "can-llms-work-in-the-wet-lab", "title": "Can LLMs work in the wet lab?", "summary": "Benchling released BenchBench-Protocol, a benchmark built from thousands of real-world experiments, showing that Anthropic's Opus 5 leads at 59.2%, followed by OpenAI's GPT 5.6 at 47.1%, and open-source Kimi K3 at 45.7%, revealing gaps in practical lab reasoning. The benchmark aims to help model developers improve LLMs for wet lab biology and assist scientists in selecting models.", "body_md": "# Can LLMs work in the wet lab?\n\nToday, we’re publishing the results from **BenchBench-Protocol**, a benchmark built from thousands of real-world experiments that tests whether models can troubleshoot and optimize wet lab protocols like expert scientists. Read the paper [here](https://assets.ctfassets.net/kzeezny59h5p/1spGaR3QP9B9CBfuXHMrLl/0ba5a1de599daf1e57faabc62564eb88/BenchBench-Protocol.pdf).\n\nLLMs know an extraordinary amount of biology. But knowing biology and *doing* biology are different things. A textbook can tell you how CRISPR works. It can’t teach you to troubleshoot a CRISPR protocol when cells die after transfection.\n\nWe’ve spent the last year bringing AI to scientists across academia and industry. While there’s universal excitement for AI, the gap is clear. AI needs to get better at reasoning like an expert scientist in the lab. Today, LLMs are often evaluated on general reasoning and scientific knowledge, not the problems scientists actually face in the lab. That [seam](https://www.benchling.com/blog/ai-scientist-that-deserves-the-name) between the digital and the physical worlds is where we are focused.\n\nWe recently built a team focused on evaluating and improving LLMs in wet lab biology. We are publishing the first in a series of benchmarks so model developers have a goal to improve on and scientists can choose models that best fit their work.\n\n**BenchBench-Protocol\n**\n\nDebugging and optimizing protocols is a challenge scientists face every day. Scientists start from a published protocol for a new technique they’re trying to use. They get it to work in their lab through repeated trial and error: adjusting it for different instruments, swapping reagents for ones they already have, tuning incubation times and concentrations.\n\nWe wanted to measure how well LLMs reason through these problems. We collaborated with scientific experts to study thousands of real-world experiments. We identified the optimizations scientists made to protocols to get them to work, had them reviewed by multiple experts, and then turned them into tasks for an LLM.\n\nOpus 5 performs the best on BenchBench-Protocol at 59.2%, followed by GPT 5.6 at 47.1%. Kimi K3 performs surprisingly well among open-source models at 45.7%.\n\nThe failure modes point to gaps in the practical reasoning that makes an expert successful in the lab. Models give answers that sound sensible on paper but miss how experiments actually behave: assuming a measurement translates directly to a result without considering calibration, treating a sample as pure despite visible evidence it isn’t, or missing subtle changes in physical technique that can ruin an experiment.\n\nThese gaps are unsurprising given what LLMs have been trained on. To get better, models need to learn not just from the clean results in published papers, but from real experiments and their messy outcomes.\n\n**Supporting the ecosystem**\n\nReinforcement learning has made models dramatically better in other domains like coding: give them thousands of tasks, tell them whether their answers work, and they learn better ways to reason. We believe the same thing can happen in biology.\n\nCareful engineering can turn messy scientific data into tasks for models to attempt, learn from, and be evaluated against. We’ve learned a lot about how to do this well: finding tasks that capture real scientific judgment, separating idiosyncratic decisions from sound scientific reasoning, and developing rubrics that can reliably tell the difference. We’re now scaling this work up to increasingly difficult problems, from sequence design and recommending the next experiment all the way up to critical drug program decisions.\n\nRigorous evaluations have become essential to how model labs improve frontier models. We expect they’ll become increasingly important to biopharma too, as companies look to understand how well AI can reason about their own science.\n\nIf you’re interested in evaluating and improving AI for your own company, we’re happy to help. Get in touch [here](mailto:ai@benchling.com).", "url": "https://wpnews.pro/news/can-llms-work-in-the-wet-lab", "canonical_source": "https://www.benchling.com/blog/can-llms-work-in-the-wet-lab", "published_at": "2026-08-15 00:58:27+00:00", "updated_at": "2026-08-15 01:11:18.358301+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-tools"], "entities": ["Benchling", "BenchBench-Protocol", "Anthropic", "Opus 5", "OpenAI", "GPT 5.6", "Kimi K3"], "alternates": {"html": "https://wpnews.pro/news/can-llms-work-in-the-wet-lab", "markdown": "https://wpnews.pro/news/can-llms-work-in-the-wet-lab.md", "text": "https://wpnews.pro/news/can-llms-work-in-the-wet-lab.txt", "jsonld": "https://wpnews.pro/news/can-llms-work-in-the-wet-lab.jsonld"}}