Harvard Study Shows AI Undermines Leaders’ Judgment in Innovation Evaluation A Harvard field experiment with 288 experienced evaluators assessing 48 submissions to an MIT global social impact innovation challenge found that AI assistance can erode leadership judgment, with evaluators who received LLM assistance with narrative explanations showing different agreement patterns with a four-expert panel, indicating automation bias. The study suggests organizations deploying large language models for idea screening may introduce hidden biases and should require evaluators to document their own reasoning before seeing AI results. August 19, 2026 , Inside AI — Harvard researchers have uncovered a troubling pattern in how artificial intelligence is reshaping leadership judgment. A new field experiment suggests that when experienced evaluators lean on large language models, their decision-making quality can erode in subtle but significant ways. The study, which examined innovation assessment, found that AI assistance does not automatically improve human judgment. In some cases, it quietly undermines it. The findings raise urgent questions for executives who increasingly rely on AI tools to evaluate ideas, candidates, and strategic proposals. The researchers recruited 288 experienced evaluators to assess 48 submissions to an MIT global social impact innovation challenge . They tested three conditions: human-only evaluations, LLM evaluations with a narrative explanation, and LLM evaluations with no further context. The evaluators' decisions were then compared against those made by four experts affiliated with the innovation challenge. The experiment's design was deliberate. By pitting human judgment against AI-assisted judgment in a real-world evaluation setting, the researchers aimed to isolate how LLM outputs influence expert decision-making. The stakes were not hypothetical: innovation challenges like MIT's often determine which social ventures receive funding and support. AI's Quiet Erosion of Expert Confidence The most striking finding was not that AI made worse decisions than humans. It was that AI changed how humans made decisions. Evaluators who received LLM assistance with narrative explanations showed different patterns of agreement with the expert panel, suggesting that the AI's reasoning shaped their own judgment more than they may have realized. This phenomenon, known as automation bias, has been documented in other fields such as aviation and medicine. When humans trust automated systems too readily, they may override their own expertise or fail to critically evaluate the AI's output. The Harvard study suggests this bias now extends to strategic innovation assessment. Notably, the researchers tested a condition where LLM evaluations came with no further context. This design choice allowed them to separate the influence of the AI's raw score from the influence of its narrative explanation. The results indicate that the format of AI output matters as much as the output itself. The implications for leadership are direct. If experienced evaluators can be swayed by AI reasoning, then organizations that deploy LLMs for idea screening, grant review, or product evaluation may be introducing hidden biases. The very tools meant to improve judgment could be dulling it. What Leaders Can Do Differently The study does not suggest abandoning AI. Instead, it points to the need for structured oversight. Leaders should treat LLM outputs as one input among many, not as a default recommendation. Requiring evaluators to document their own reasoning before seeing AI results could help preserve independent judgment. Training also matters. Evaluators who understand how LLMs generate text, including their tendency to produce confident but flawed reasoning, are better equipped to challenge AI outputs. Organizations that invest in AI literacy for decision-makers may see more reliable outcomes. The Harvard findings align with broader concerns in the AI research community. Studies on algorithmic decision support have repeatedly shown that human-AI teams do not always outperform humans alone. The key variable is not the AI's accuracy but how humans integrate its advice. For innovation leaders, the lesson is clear: AI can accelerate evaluation, but it cannot replace the disciplined judgment that comes from domain expertise. The challenge is to use AI without surrendering the critical thinking that makes human evaluators valuable in the first place.