04:00
2026-09-30
machinebrief.com
large-language-models
CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence
A new benchmark called CypherTurn, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena, shows the best of 15 evaluated models reaches only 64.7% execution…