Marvin tops CompBioBench with 99/100, updated model roster Marvin, an AI agent from FutureHouse, scored 99/100 on Genentech's CompBioBench computational biology benchmark, with Marvin Lite scoring 95/100, outperforming GPT-5.6 Sol xhigh entries. The results highlight rapid progress in agentic biology but underscore persistent challenges in nuanced scientific reasoning. Marvin now supports GPT-6 Astra and other models, expanding provider options. Marvin top scored 99/100 on CompBioBench https://huggingface.co/spaces/Genentech/compbiobench-leaderboard-v1 , putting it at the top of Genentech’s computational biology benchmark. Meanwhile, “Marvin Lite” scored 95/100 while being limited to 1 iteration only. Both are several points higher than the next best GPT-5.6 Sol xhigh entries on the leaderboard. We have two main takeaways from these submissions. First, it’s truly remarkable how quickly the field has reached this point. CompBioBench was introduced in April with 100 questions requiring agents to work with biological datasets and specialized research tools. Less than six months later, general-purpose coding agents are already scoring in the 90s, even without a dedicated scientific harness. Frontier models have become capable of handling substantial computational biology workflows, which is encouraging for researchers who spend much of their time performing these analyses. On the other hand, agentic biology is NOT a solved problem, as demonstrated by the difficulty of the long tail in an otherwise saturating benchmark. The remaining failures increasingly sit in the part of the distribution where examples are scarce and a familiar method can be wrong for reasons that only become apparent with the right biological context. These cases are difficult precisely because the patterns that usually lead to the right answer can point confidently in the wrong direction when nuance is required. As we discussed in our post on vibe science /blog/2026/07/when-scientific-output-outruns-scientific-judgment/ , an analysis can look convincing while depending on assumptions that do not hold for the particular problem. Marvin’s skills draw on experts who know where those methods break down and which seemingly minor details warrant a different approach. That said, there will still be situations where reasonable scientists disagree, so Marvin makes those decisions available for review. Its logic trail records why an approach was chosen, while Threads connects hypotheses to the experiments and findings that support or challenge them. In reviewing Marvin’s work after a single iteration “lite mode” vs the full run, we found cases where Marvin had already considered the alternative that later earned credit with more deliberation, but rejected it in favor of a more conservative or scientifically defensible interpretation. The original analyses preserved enough context for us to understand that choice and revisit the assumption behind it. New model support in Marvin Marvin now supports GPT-6 Astra We’ve also recently added Gemini 3.8 Flash and Gemini 3.5 Flash-Lite, Kimi K3, DeepSeek V4 Pro and Flash, and GLM-5.3. The Qwen 3.8 family is available through OpenRouter, joining existing options such as MiniMax-M3. In addition to direct provider support, we’ve broadened Together’s model selection and added Fireworks as another provider, alongside our existing direct connections and OpenRouter support. This gives you more choice over which models you use for your research and where you access them. We’ll keep adding new models as they’re released, so you can use the ones you prefer with Marvin.