Researchers from Harvard and MIT, with Stanford collaborators, have built an AI system called MatrAIx that simulates roughly one persona for every human on Earth, and companies can now run products past it before a single real user sees them.
The paper, titled "MatrAIx: Simulating the World with 8.3 Billion Persona Agents," hit arXiv on August 4, 2026. Its core is a database called Persona 8B: 8.3 billion synthetic persona records, each one scored across 1,290 categorical dimensions, from basic demographics to finer behavioral traits. Lead contact Xiaomin Li built it with Yuexing Hao, Marinka Zitnik and James Zou, whose affiliations span Harvard, MIT and Stanford.
Here's the number that matters most: in validation trials, the declared behavior of a persona was expressed or correctly suppressed 91.5% of the time, based on 366 of 400 trials. That's the researchers' pitch. A synthetic population that mostly stays in character, at a scale no human panel could ever match.
You don't get access to all 8.3 billion records. The team released a filtered coreset of about 1 million personas, 599,847 grounded in real human data and 400,000 fully synthetic, according to the arXiv paper. For the human-grounded slice, the researchers say they stripped names and contact details before release, keeping only extracted attributes and descriptions. Whether that de-identification holds up is another matter. Re-identification attempts on a dataset this granular, with over a thousand dimensions per person, are a real risk the paper doesn't fully settle.
Alongside the dataset, MatrAIx ships a Playground with four environments: Survey, Chatbot, Web and App. In each one, AI agents wearing personas interact with a target product. The system logs what happens. Run a concept test through Survey and you get structured and open-ended feedback across a synthetic population instead of a recruited panel. Run a support bot through Chatbot and you get scores on task completion, helpfulness, safety and multi-turn reliability. Run a prototype through Web or App and you get usability, navigation and latency data.
That's the pitch to market research and consulting firms, and it's a real threat to their business model. Traditional panels are slow and expensive. Recruiting a few hundred people for a concept test, paying them, scheduling sessions, running the study, can take weeks and cost tens of thousands of dollars. A synthetic population that behaves like real users, tested overnight against 1,290 dimensions of variation, undercuts that on both speed and price. Consulting firms that sell human-panel research as a core offering now have a Harvard-and-MIT-branded competitor sitting on arXiv and GitHub for anyone to try.
Frankly, the more interesting fight isn't speed. It's whether synthetic answers are the right answers. The ACM's Interactions blog has already flagged what it calls the "synthetic persona fallacy": large language models don't have knowledge, wants, or a stable preference system the way a real customer does, so a persona that stays in character 91.5% of the time isn't the same as a persona that predicts real behavior 91.5% of the time. Those are different claims, and MatrAIx's own validation numbers measure the first one, not the second.
The privacy question nobody's fully answered #
Modeling roughly one persona per living person, across 1,290 behavioral dimensions, raises a consent question the paper addresses only partially. The human-grounded records come from real people's data, de-identified before release. But de-identification at this resolution is a harder problem than stripping a name and an email address. Enough behavioral dimensions, correlated the right way, can narrow a record back down to one person even without a name attached. The researchers don't claim to have solved that; they describe a mitigation, not a guarantee.
There's also the question of what this replaces. If a synthetic panel can stand in for real users during early testing, companies may run fewer studies with actual people, meaning fewer opportunities to catch the things a model can't predict: genuine confusion, a product that feels wrong for reasons nobody articulated in a survey box. GWI's own writing on synthetic personas makes a similar point: these tools are useful for narrowing options fast, not for replacing the moment you watch a real person struggle with your product.
MatrAIx is free to try, and the code sits on GitHub for anyone to pull. Behind it are two of the most credentialed AI labs in the country. That combination will get it into a lot of product teams' testing pipelines fast, probably faster than the open questions about privacy and predictive accuracy get resolved.
Also read: JPMorgan Raises S&P 500 Target to 8,000 Saying AI Spending Is Finally Paying Off • Meta Releases Muse Glimmer, an Open-Weight AI Model That Runs on a Laptop • SpaceX Is Closing In On A $6 Billion Deal For Israeli AI Startup Decart