cd /news/artificial-intelligence/ai-powered-simulation-and-human-beha… · home › topics › artificial-intelligence › article
[ARTICLE · art-142179] src=blog.expectedparrot.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

AI-Powered Simulation and Human Behavioral Research Are Almost Certainly Complements

Aaru's AI model predicts survey outcomes at roughly 8 total variation distance (TVD), a level of accuracy comparable to other frontier models and equivalent to about 50 human responses, according to a benchmark posted by Expected Parrot. The analysis argues that cheap, imperfect AI screening complements rather than replaces human behavioral research, since a screen that lifts success probability from 5% to 10% would halve the human screening needed to reach a 20%-success "promising" idea and increase total human research demand.

by read4 min views1 publishedSep 30, 2026
AI-Powered Simulation and Human Behavioral Research Are Almost Certainly Complements
Image: Blog (auto-discovered)

There is a tremendous amount of discussion about “silicon sampling” and “digital twins.” The framing is almost always:

a) it works and humans are no longer necessary (investment pitches) or

b) it doesn’t “work” and is useless (LinkedIn hot takes or PNAS papers) or

c) it sorta works but you always need to “check” with real humans.

In the interest of brevity I think a) is absurd and, for reasons I should blog about, a logical impossibility. But I also don’t think b) is credible: there is clearly enough regularity in human responses that even with “new” scenarios, language models can do far better than chance and predicting responses because they’ve learned so much about our world.

In a commendably detailed blog post, Aaru showed their model doing well at predicting survey outcomes unlikely to be in the training data. I should note, however, that its performance is ballpark comparable to other frontier models (which I did here), at least on predicting “marginals”:

It’s good…but this ~8 TVD is about the performance you would get from about -50 humans responses, give or take. Not enough to feel confident in many if not most real business decisions where a survey could have been valuable. And this is clearly not he space of all possible questions. So the c) “always need to check” seemingly applies, which defeats the purpose and benefit: if I’m just going to run a real study anyway, what’s the point? Where are the savings?

“Lump of studies” fallacy #

A key warning in labor economics is not to fall victim to the “lump of labor” fallacy: the idea that there is some finite amount of work to be done and, say, better tools or immigration would mechanically lead to fewer hours per worker. I propose that c) thinking is an analogous fallacy about research: if some studies are done with AI instead, then there are fewer human studies to be done. Or, if there are not fewer human studies (because of c—everything must be checked), then there is a false economy and AI is a waste.

But this is missing the key so-what, that the studies we do are not fixed, but rather are done because of the benefit we expect to receive from running them. And that benefit changes when we can do a cheap AI screen.

The value of a cheap, imperfect screen #

Suppose I’m trying to decide whether to launch a product, perhaps the quintessential market research task. Like all real products, I have a vast space of potential ways to design the product (size, features, color, packaging, flavors, price, etc.). Most of the ideas in this space are “bad” in the sense that if launched in that configuration, it would fail. Let’s make it concrete with some numbers: suppose the baseline probability of success is 5% (I think our CPG friends would love those odds). If I do some market research with humans, I learn more—it tells me which ideas are “promising” which have a 20% chance of success. To get one “promising” idea, I have to screen with human research 4 ideas, and even more if I sometimes screen out good ideas.

Now suppose I have an AI technology that will let me get to 10% success probability ideas essentially for free without screening out any good ideas. Note that this is far less effective than the “real” human research. If human research further selects a pool with a 20% success rate, I now need to screen only 2 ideas with humans, on average, to find one “promising” idea. I’ve doubled the productivity of my human market research and now I want to do even more of it, all else equal. I’m doing more human research than before because of AI. It’s sort of like the c) point but actually good and welcome!

I had ChatGPT make this little diagram to make the point: with a cheap AI screen that excludes bad ideas, your human research dollars just go further:

Does AI ever replace human evaluation? #

Well, if you could simply had an AI that could tell you if the idea was good a) from above, then yes. But that’s not even remotely plausible. What instead can happen is that as AI gets better, eventually the returns to human evaluation falls. Below is a little one-pager model that lays out the idea a bit more formally. But I think we’re very far away from that point and for the kind of decision-focused research I’m talking about here, I see nothing but complementarities.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @aaru 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-powered-simulatio…] indexed:0 read:4min 2026-09-30 · —