WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Researchers introduced WearableQA, a benchmark of 4,084 10-option multiple-choice questions designed to test whether AI systems can reason over a real user's longitudinal wearable data. The benchmark targets health reasoning over continuous physiological and behavioral signals from wearable sensors, an area the authors say existing benchmarks rarely evaluate. Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user's longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-ch