Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools
A new benchmark called ABLE, introduced in arXiv paper 2609.05818v1, evaluates how well LLM agents use biological AI models such as ProteinMPNN and AlphaFold3 in dual-use protein design workflows. Testing 15 frontier mod…