Your Agent Aced the Task. Will It Do It Again?
IBM Research introduced consistency guidelines in its altk-evolve toolkit, built on a new Consistency Analyzer, to address a 24.4-point consistency gap in LLM agent reliability. A ReAct agent using GPT-4.1 on AppWorld's …