{"slug": "trace-bench-task-driven-roleplay-agentic-checklist-evaluation", "title": "TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation", "summary": "Researchers introduced TRACE Bench, a task-driven agentic checklist evaluation framework for roleplay models that decomposes role profiles into fixed checklists and uses a User Agent to update checklist states from model responses, enabling scores to trace back to specific checklist items and dialogue evidence. In coverage cross-validation against the MiniMax Role-play Benchmark, TRACE Bench reached 99.91% coverage in fewer turns, compared to 73.74% coverage from released free-chat transcripts. Across 26 models, TRACE Bench provides overall rankings, capability breakdowns, and checklist traces, and supports Closed-Loop Benchmark Evolution to improve future evaluations.", "body_md": "arXiv:2608.11236v1 Announce Type: new\nAbstract: Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and which dialogue evidence supports the judgment. We propose TRACE Bench, a task-driven agentic checklist evaluation framework. It decomposes each role profile offline into a fixed checklist, then uses a User Agent to converse naturally with the target roleplay model while privately updating checklist states from model responses. Scores therefore trace back to checklist items and supporting dialogue turns rather than a black-box holistic impression. For coverage cross-validation, we audit released M2 free-dialogue transcripts from the MiniMax Role-play Benchmark against the same role-derived checklist. The released free-chat transcripts cover only 73.74% of key role-profile points, whereas TRACE Bench reaches 99.91% coverage in fewer turns. Robustness experiments show stable rankings under repeated runs and User Agent replacement. Across 26 models, TRACE Bench reports overall rankings together with capability breakdowns and checklist traces. It also supports Closed-Loop Benchmark Evolution, distilling verification methods proven effective in failed traces so later evaluations can more reliably elicit and examine observed failure modes.", "url": "https://wpnews.pro/news/trace-bench-task-driven-roleplay-agentic-checklist-evaluation", "canonical_source": "https://arxiv.org/abs/2608.11236", "published_at": "2026-08-13 04:00:00+00:00", "updated_at": "2026-08-13 04:11:10.197149+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["TRACE Bench", "MiniMax Role-play Benchmark", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/trace-bench-task-driven-roleplay-agentic-checklist-evaluation", "markdown": "https://wpnews.pro/news/trace-bench-task-driven-roleplay-agentic-checklist-evaluation.md", "text": "https://wpnews.pro/news/trace-bench-task-driven-roleplay-agentic-checklist-evaluation.txt", "jsonld": "https://wpnews.pro/news/trace-bench-task-driven-roleplay-agentic-checklist-evaluation.jsonld"}}