00:00
2026-08-06
machinelearning.apple.com
large-language-models
DeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Researchers introduced DeepAmbigQA, a dataset of 3,600 questions requiring multi-hop reasoning with half containing explicit name ambiguity, to benchmark LLM answer completeness. Tests showed that eveโฆ