cd /news/artificial-intelligence/icm-bench-person-level-identity-reas… · home topics artificial-intelligence article
[ARTICLE · art-121883] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

Researchers introduced ICM-Bench, the first benchmark for evaluating identity-centric reasoning in multimodal agents over long video memories, containing 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults. Gemini 3.1 Pro achieved the highest overall accuracy of 74.0%, but its score dropped to 60.3% on questions requiring long-term identity profiles, revealing that current systems are less reliable when evidence must be accumulated around a stable person.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04438v1 Announce Type: new Abstract: Long-horizon multimodal agents should remember not only what happened but also who participated. This capability depends on linking recurring faces, voices, names, person-associated objects, events, and social relations to consistent identities over time. Existing long-video and multimodal-agent benchmarks measure broad memory question answering, but they do not isolate the ability to maintain recurring person identities and reason over their cross-time relations. We introduce ICM-Bench (Identity-Centric Memory Benchmark), which, to the best of our knowledge, is the first benchmark specifically designed to evaluate identity-centric reasoning over long video memories in multimodal agents. The benchmark contains 839 synthetic clips spanning 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. A theme-configurable pipeline generates the video collection and associates each question with its target identities and traceable supporting evidence. We compare direct caption-memory baselines, memory-augmented agents, and graph-retrieval systems. Gemini 3.1 Pro achieves the highest overall accuracy of 74.0%, yet its score falls to 60.3% on questions that require long-term identity profiles. The results show that current systems recover many event-level memories but remain less reliable when evidence must be accumulated around a stable person.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @icm-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/icm-bench-person-lev…] indexed:0 read:1min 2026-09-07 ·