KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
Researchers introduced KC-Bench, a dynamic interactive benchmark with 238 manually screened tasks for evaluating how LLM agents handle knowledge conflicts across world-knowledge, input inconsistencies…