From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization A study by researchers including Myeongseob Ko, posted on arXiv on March 19, 2026 and revised May 29, 2026, shows that LLM-based agents can reconstruct real-world identities from scattered non-identifying cues, achieving 79.2% identity reconstruction in the Netflix Prize deanonymization setting versus 56.0% for a classical matching baseline. The findings indicate that privacy evaluations for agentic systems should measure inferred identities, not just accessed or disclosed information. Computer Science Artificial Intelligence Submitted on 19 Mar 2026 v1 https://arxiv.org/abs/2603.18382v1 , last revised 29 May 2026 this version, v2 Title:From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents View PDF /pdf/2603.18382 HTML experimental https://arxiv.org/html/2603.18382v2 Abstract:Anonymization is often assumed to protect privacy once explicit identifiers are removed, because re-identification has historically required specialized expertise, tailored algorithms, and manual corroboration. We show that LLM-based agents weaken this barrier: by combining scattered, individually non-identifying cues with public evidence, they reconstruct real-world identities, sometimes even during benign tasks. We evaluate this risk across three settings -- classical linkage incidents, a controlled benchmark \emph{InferLink} that varies fingerprint type, task framing, and attacker knowledge, and open-ended human--AI interaction traces. In the sparsest regime of the Netflix Prize deanonymization setting, agents reconstruct 79.2\% of identities, against 56.0\% for a classical matching baseline; on \emph{InferLink}, they link individuals even without an explicit re-identification request, and more often once one is given. In redacted human--AI interaction traces, agents further resolve anonymized profiles to specific individuals by corroborating contextual cues with public evidence. These findings suggest that privacy evaluations for agentic systems should measure not only what information is accessed or disclosed, but also what identities can be inferred. Submission history From: Myeongseob Ko view email /show-email/bdcf36ad/2603.18382 Thu, 19 Mar 2026 00:59:26 UTC 7,535 KB v1 /abs/2603.18382v1 v2 Fri, 29 May 2026 16:18:22 UTC 7,527 KB References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .