ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory Researchers introduced ICM-Bench, the first benchmark for evaluating identity-centric reasoning in multimodal agents over long video memories, containing 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults. Gemini 3.1 Pro achieved the highest overall accuracy of 74.0%, but its score dropped to 60.3% on questions requiring long-term identity profiles, revealing that current systems are less reliable when evidence must be accumulated around a stable person. arXiv:2609.04438v1 Announce Type: new Abstract: Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associated objects, events, and social relations to consistent identities over time. Existing long-video and multimodal-agent benchmarks measure broad memory question answering, but they do not isolate the ability to maintain recurring person identities and reason over their cross-time relations. We introduce ICM-Bench Identity-Centric Memory Benchmark , which, to the best of our knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents. The benchmark contains 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. A theme-configurable pipeline generates the video collection and associates each question with its target identities and traceable supporting evidence. We compare direct caption-memory baselines, memory-augmented agents, and graph-retrieval systems. Gemini 3.1 Pro achieves the highest overall accuracy of 74.0%, yet its score falls to 60.3% on questions that require long-term identity profiles. The results show that current systems recover many event-level memories but remain less reliable when evidence must be accumulated around a stable person.