From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning NVIDIA and academic collaborators introduced TAR (Traffic Anomaly Reasoning) and TAR-Bench, datasets with 44,040 chain-of-thought training annotations across 10 tasks for 3,670 CCTV videos (~26 hours) from eight public datasets, plus 960 human-curated test annotations for 80 held-out clips from 17 public YouTube videos. Eleven vision-language models evaluated on TAR-Bench showed that strong question-answering accuracy does not reliably predict temporal or scene reasoning ability, while multi-task fine-tuning on TAR improved aggregate score by 21.4 points over zero-shot baseline. TAR and TAR-Bench serve as official training and in-domain evaluation data for AI City Challenge 2026 Track 3. arXiv:2608.10317v1 Announce Type: new Abstract: We present TAR Traffic Anomaly Reasoning and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection. TAR contains 44,040 chain-of-thought training annotations across 10 tasks for 3,670 CCTV videos $\sim$26 hours from eight public datasets. Its evaluation component, TAR-Bench, contains 960 human-curated test annotations for 80 held-out clips trimmed from 17 public YouTube videos. TAR's training annotations are produced with MAVEN, which consolidates multi-scale video evidence into structured event descriptions before generating question-answer pairs and reasoning traces. On TAR-Bench, eleven vision-language models reveal that strong question-answering accuracy does not reliably predict temporal or scene reasoning ability. Multi-task fine-tuning on TAR yields consistent gains, with the full 10-task model improving aggregate score by 21.4 points over its zero-shot baseline. TAR and TAR-Bench provide the official training and in-domain evaluation data for AI City Challenge 2026 Track 3. The dataset is available at https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning