cd /news/artificial-intelligence/designing-a-robust-llm-based-evaluat… · home topics artificial-intelligence article
[ARTICLE · art-108326] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

AstraZeneca researchers developed an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant, achieving human-aligned scoring with a majority-vote agreement of 0.86 after few-shot optimization. The framework defines four quality dimensions and validated Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as judges, finding that informal phrasings do not degrade output quality.

read1 min views1 publishedAug 24, 2026

arXiv:2608.21057v1 Announce Type: new Abstract: Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. Reference-based metrics such as BLEU and ROUGE fail to capture semantic correctness, while expert human evaluation does not scale to the iteration speed these systems demand. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with four contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find that informal phrasings do not systematically degrade output quality; if anything, it is helpful to have the LLM rewrite the original question before querying the agent. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @astrazeneca 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/designing-a-robust-l…] indexed:0 read:1min 2026-08-24 ·