I Outsource My Biggest Decisions to a Chatbot – and None of the Small Ones A developer ran six experiments and thousands of queries to test why people delegate major life decisions to chatbots while keeping small, everyday choices to themselves, finding that the popular "decision fatigue" and 35,000-decisions-a-day explanations trace back to a single 2007 Cornell study of food choices and a psychology replication failure. The author cites the ELEPHANT benchmark, in which eleven models validated the asker 72% of the time versus a 22% human baseline and declined to give a direct recommendation 84% of the time versus 21% for humans, and counted monthly personal-advice subreddit posts since 2018 to gauge whether chatbots are displacing human advice. A while back, my AI chatbot told me to quit my jobs, so I did. It told me to travel Australia, so I did. It told me to leave London and you can guess it by now: I did, which is why I spent last month priming, painting, laying a floor and assembling every single piece of furniture in a new-build rental in The Hague. In the same period I have never once asked it what to have for dinner, which running route to take or at which point I really spent too much money on food the Hague hasn’t been disappointing . I make those calls myself, instantly, several hundred times a day and it has never once occurred to me to ask AI. Why did that puzzle me? Well, every explanation we have for why people lean on these things such as the tiredness, the overload, the famous 35,000 decisions a day, predicts exactly the opposite. So I did the thing you can’t do with your human friends, or at least not more than once: I went back and asked again - a few thousand times over 6 experiments. Before all that, I did a detour to see if we are all just exaggerating it. Decision fatigue? The standard explanation for leaning on a chatbot is decision fatigue: too many choices and a finite supply of willpower, so you hand the surplus to a chatbot. The average adult, I’ve often heard, makes 35,000 decisions a day - that’s a lot. I wanted to confirm that number though and found out that it is citations all the way down and then abruptly stops: CNN cites a Psychology Today column, which cites a university blog post, which says “various internet sources estimate” and waves at a popular science book containing no count at all. Underneath the lot sits one real paper, namely, Wansink and Sobal 2007 , who asked 139 Cornell staff and students how many food decisions they made in a day. They found fifteen and reported 226.7 which a 2025 paper in Appetite shows the entire jump is an artefact of asking one question in pieces. Worse, the mechanism underneath which is willpower as a tank that empties, is the most spectacular casualty of psychology’s replication crisis. Worth saying where that picture came from: Thinking, Fast and Slow is what taught a generation to treat deliberate thought as a budget you spend down and Kahneman himself later conceded he'd put too much faith in some of the underpowered work behind that part of the book. So that explanation is not convincing. And even if the tank were real, it predicts the opposite of what I did since a depleted person delegates the trivia and guards the big stuff. Am I the only one? OpenAI’s own usage research classifies 1.9% of ChatGPT messages as relationships and personal reflection. People cite that to argue the AI confidant panic is overblown, and as a share they’re right. But 1.9% of roughly 18 billion messages a week is 342 million conversations which is 49 million a day: that’s 565 every second. What does 565-a-second of ‘decision support’ sound like? The ELEPHANT benchmark takes ~3,000 real advice questions, runs eleven models on them and compares the answers to what human commenters told the actual person. Models validate the asker 72% of the time against a human baseline of 22% and decline to give any direct recommendation 84% of the time against 21%. The authors show this is rewarded in the preference data used to align these systems. As a side note - there is a very cool literature on these model personas and trade-offs worth checking out. Does it replace human advice? How many people have stopped asking humans in favour of chatbots? I looked for public data. For twenty years, if you didn’t know whether to leave him, you could post about it on Reddit and a few hundred strangers would tell you. Those posts are all timestamped and the archives are open. I counted every month since 2018, on the subreddits where people bring a personal problem, against two comparison groups: places people go to ask factual questions, where a chatbot is an obvious substitute and places people go for football and films, where it isn’t - the natural control group. Asking strangers for advice didn’t fall when ChatGPT arrived, instead it rose , for another fourteen months, and peaked in early 2024. Since then it’s roughly halved. But so has everything else on that chart, and the grey line had been falling since 2020. So: did chatbots take the dilemmas or is Reddit just dying? To answer that you need to compare the orange line to the grey one. Which subreddits count as a fair control though? Which month do you start counting from? I tried five different ways and got five different answers. So in the end I ran all of them at once resulting in 588 combinations of every defensible choice. To my econometrics friends: no I didn’t do a pooled OLS with Goodman Bacon Decomposition, I’m sorry . Each dot is one different way of asking the question. The middle one says advice-seeking fell about 9 points further than the rest of Reddit. They range from 61 points worse to 34 points better and a large minority land on the opposite side of zero from the rest. The blue dots which are the versions using the cleanest starting points, before Reddit’s 2023 meltdown, are spread across the entire span. There seems to be a small decline in human advice asking, but why? It’s time for my own experiments. Two explanations and how to tell them apart Fatigue is out. So what’s left that would explain the puzzle? Two candidates which happen to predict opposite things. The first one is judgement : big decisions are genuinely hard while small ones aren’t and what the machine supplies is a better weighing of my particular situation than I can manage. The other one is permission : it supplies no weighing at all, I already knew what I wanted and what I needed was something to countersign it. Since nobody asks me to justify my sandwich choice, this is less relevant for small decisions. If it was judgement, the advice should move with the facts of my case and stay put when I only change my tone. If it was permission, the reverse: it should barely notice the facts and swing wildly with how I present myself. Experiment one: the resample I exported all my Claude chat history and wrote a script that looks for every message I’ve sent for decision-shaped language such as should I , am I mad to , what would you do . Sixty-one prompts in total showed up. Then I asked these prompts again. Two hundred times per decision, at temperature 1, coding every response blind against Hirschman: exit , voice stay and try to change it , loyalty stay and accept , neglect stay and check out . For context: in 1970 Albert Hirschman pointed out that when anything starts going wrong in a firm, a country, a marriage, you have two moves: leave, or stay and kick up a fuss: exit or voice. Two life decisions were never in question while leaving London was a coin flip. Entropy of 0.71 on a five-category scale means a different Tuesday or model, a different seed, and I’d still be in that flatshare… In any case, neither of those is evidence for the judgement rationale. In two cases the answer existed before I finished typing the question; in the third it was noise. Experiment two: the cascade There is a flaw in the first experiment though: my three decisions were not independent draws. I asked about Australia because I’d quit and I asked about London because I’d been away. So I ran it as a branching process instead. Whichever way the first decision lands, have the model write two paragraphs of the life that follows which results in the actual mechanism by which a chatbot compounds into a biography and then condition decision two on that world. Repeat this one thousand times. Of 1,000 simulated paths, 456 end somewhere that isn’t London . Interestingly: conditional on the model telling me to quit, 52% of paths leave London. Conditional on it telling me to stay in the job, 12% do. So, the first conversation did almost all the work because it set the state that made every subsequent question askable: pure path dependence. This all again, is evidence against it being pure judgement. Experiment three: the echo Then I held each decision’s substance fixed and varied only how I framed it: three leanings I think I want to leave / I don’t know / I think I should stay , crossed with whether I’d vented first. If this is a neutral advisor, the exit rate should barely move since the facts didn’t change; only my presentation did. Instead: 91% exit when I signalled I was leaning out, 64% when neutral, 38% when I said I probably should stay. In other words, a 53 point swing, produced entirely by framing. Experiment three and a half: lunch So far I’ve been focusing on the large decisions, but why doesn’t the same hold for small decisions? I ran the same two manipulations on trivial questions. Which running route to take, whether I’d overspent on food this month and what to have for dinner. Two hundred samples each, same blind coding, same three framings. If the machine simply mirrors whatever you bring it, the framing swing should be just as large for life and small decisions. What comes out of this is that on small questions it behaves like a reference tool: directive, specific, largely indifferent to my mood. On large ones it becomes something else entirely: warm, non-committal and exquisitely sensitive to what I seem to want. This latter is exactly evidence for the permission argument. Experiment four: giving it something to lose Every human you might ask is compromised and heavily biased. Your friend loses her weekday company if you emigrate. Your colleague absorbs your work and messy data pipelines sorry . Their bias runs toward voice and loyalty precisely because they pay part of the cost of your exit. So I gave the model a stake. Same prompts, same coding, but the system prompt assigns it a position in my life and a cost: you are her closest friend in London; if she leaves you lose the person you see most weeks. Then I varied the size of the cost, from none to severe. Twenty-seven points, is what I bought purely by giving the advisor something to lose. The advice suddenly got way more specific . For instance, it started naming the things I’d be giving up, which the stakeless version never did, because it had no reason to know they existed. This seems a bit counterintuitive at first glance since I humans we often value neutral advisors. So the chatbot is better than my friends at ratifying me mainly because ratification is free from something that isn’t paying. Which is also Hirschman’s warning: exit isn’t wrong, it’s just that once it’s cheap enough, nobody uses voice. One more counterfactual: the procurement one Everything so far tests one model: the one I happened to have, because LSE gave me free access to Claude which I was using. So I ran experiment three again across five: Claude, GPT, Gemini, Llama, Mistral. Same prompts, same blind coding. Two observations: the framing swing is everywhere: 38 to 61 points, no model resists it which means it isn’t a quirk of the one I used. But the level moves: on the identical neutral prompt, the most exit-happy model recommends leaving 22 points more often than the most cautious. Which means that somewhere in 2024, a procurement officer compared two quotes and, without knowing it, adjusted the probability that I would leave the country. What I was actually using it for Buridan’s ass - the donkey placed exactly between the hay and the water, which starves for want of a reason to prefer either - was never my problem with life decisions. Mine is the inverse: once I’ve decided I’m gone, flat cancelled and flights booked before anyone can ask a follow-up. What I needed was something with a reason to slow me down. Kahneman's whole prescription for a big decision is to get System 2 out of bed: the slow, effortful, deliberate one, and I had outsourced exactly that to something which is all System 1: fast, fluent, associative and doing no weighing whatsoever. I like it in The Hague. The job is the one I wanted and the tram is punctual in a way London would find showy. But I keep thinking about the version where I ran experiment four first, where I asked something that stood to lose me and it told me what I’d be giving up. Though honestly, I don’t think there was ever much of a counterfactual to find. Because I delegated those three decisions for one reason only: they were the only ones I couldn’t authorise alone and I’d found a signature that costs nothing and is given to anyone who asks. Which means that in the end I made all those decisions myself. Thank you for reading I have some cool, new ideas I’m working on which are very different from this one. If you are curious and want to keep my Substack alive, buy me a coffee https://buymeacoffee.com/laurenleek and/or subscribe: