cd /news/artificial-intelligence/refinebench-evaluating-refinement-ca… · home topics artificial-intelligence article
[ARTICLE · art-91602] src=research.nvidia.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

RefineBench: Evaluating Refinement Capability of Language Models via Checklists

Researchers introduced RefineBench, a benchmark of 1,000 challenging problems across 11 domains, to evaluate language models' ability to refine their own responses. In self-refinement, even frontier models like Gemini 2.5 Pro and GPT-5 scored only 31.3% and 29.1%, with most models failing to improve consistently across iterations. However, with guided refinement, both proprietary and large open-weight models (>70B) achieved near-perfect results within five turns, highlighting the need for breakthroughs in self-refinement.

read1 min views1 publishedAug 11, 2026

Can language models (LMs) self-refine their own responses? This question is increasingly relevant as a wide range of real-world user interactions involve refinement requests. However, prior studies have largely tested LMs' refinement abilities on verifiable tasks such as competition math or symbolic reasoning with simplified scaffolds, whereas users often pose open-ended queries and provide varying degrees of feedback on what they desire. The recent advent of reasoning models that exhibit self-reflection patterns in their chains-of-thought further motivates this question. To analyze this, we introduce RefineBench, a benchmark of 1,000 challenging problems across 11 domains paired with a checklist-based evaluation framework. We evaluate two refinement modes: (1) guided refinement, where an LM is provided natural language feedback, and (2) self-refinement, where LMs attempt to improve without guidance. In the self-refinement setting, even frontier LMs such as Gemini 2.5 Pro and GPT-5 achieve modest baseline scores of 31.3% and 29.1%, respectively, and most models fail to consistently improve across iterations (e.g., Gemini-2.5-Pro gains only +1.8%, while DeepSeek-R1 declines by -0.1%). By contrast, in guided refinement, both proprietary LMs and large open-weight LMs (>70B) can leverage targeted feedback to refine responses to near-perfect levels within five turns. These findings suggest that frontier LMs require breakthroughs to self-refine their incorrect responses, and that RefineBench provides a valuable testbed for tracking progress.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @refinebench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/refinebench-evaluati…] indexed:0 read:1min 2026-08-11 ·