{"slug": "harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation", "title": "Harvard Study Shows AI Undermines Leaders’ Judgment in Innovation Evaluation", "summary": "A Harvard field experiment with 288 experienced evaluators assessing 48 submissions to an MIT global social impact innovation challenge found that AI assistance can erode leadership judgment, with evaluators who received LLM assistance with narrative explanations showing different agreement patterns with a four-expert panel, indicating automation bias. The study suggests organizations deploying large language models for idea screening may introduce hidden biases and should require evaluators to document their own reasoning before seeing AI results.", "body_md": "**August 19, 2026**, (Inside AI) — **Harvard researchers** have uncovered a troubling pattern in how artificial intelligence is reshaping leadership judgment. A new field experiment suggests that when experienced evaluators lean on large language models, their decision-making quality can erode in subtle but significant ways.\n\nThe study, which examined innovation assessment, found that AI assistance does not automatically improve human judgment. In some cases, it quietly undermines it. The findings raise urgent questions for executives who increasingly rely on AI tools to evaluate ideas, candidates, and strategic proposals.\n\nThe researchers recruited **288 experienced evaluators** to assess **48 submissions** to an **MIT global social impact innovation challenge**. They tested three conditions: human-only evaluations, LLM evaluations with a narrative explanation, and LLM evaluations with no further context. The evaluators' decisions were then compared against those made by **four experts** affiliated with the innovation challenge.\n\nThe experiment's design was deliberate. By pitting human judgment against AI-assisted judgment in a real-world evaluation setting, the researchers aimed to isolate how LLM outputs influence expert decision-making. The stakes were not hypothetical: innovation challenges like MIT's often determine which social ventures receive funding and support.\n\n## AI's Quiet Erosion of Expert Confidence\n\nThe most striking finding was not that AI made worse decisions than humans. It was that AI changed how humans made decisions. Evaluators who received LLM assistance with narrative explanations showed different patterns of agreement with the expert panel, suggesting that the AI's reasoning shaped their own judgment more than they may have realized.\n\nThis phenomenon, known as automation bias, has been documented in other fields such as aviation and medicine. When humans trust automated systems too readily, they may override their own expertise or fail to critically evaluate the AI's output. The Harvard study suggests this bias now extends to strategic innovation assessment.\n\nNotably, the researchers tested a condition where LLM evaluations came with no further context. This design choice allowed them to separate the influence of the AI's raw score from the influence of its narrative explanation. The results indicate that the format of AI output matters as much as the output itself.\n\nThe implications for leadership are direct. If experienced evaluators can be swayed by AI reasoning, then organizations that deploy LLMs for idea screening, grant review, or product evaluation may be introducing hidden biases. The very tools meant to improve judgment could be dulling it.\n\n## What Leaders Can Do Differently\n\nThe study does not suggest abandoning AI. Instead, it points to the need for structured oversight. Leaders should treat LLM outputs as one input among many, not as a default recommendation. Requiring evaluators to document their own reasoning before seeing AI results could help preserve independent judgment.\n\nTraining also matters. Evaluators who understand how LLMs generate text, including their tendency to produce confident but flawed reasoning, are better equipped to challenge AI outputs. Organizations that invest in AI literacy for decision-makers may see more reliable outcomes.\n\nThe Harvard findings align with broader concerns in the AI research community. Studies on algorithmic decision support have repeatedly shown that human-AI teams do not always outperform humans alone. The key variable is not the AI's accuracy but how humans integrate its advice.\n\nFor innovation leaders, the lesson is clear: AI can accelerate evaluation, but it cannot replace the disciplined judgment that comes from domain expertise. The challenge is to use AI without surrendering the critical thinking that makes human evaluators valuable in the first place.", "url": "https://wpnews.pro/news/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation", "canonical_source": "https://insideai.news/news/machine-learning/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation/8209/", "published_at": "2026-08-19 13:12:33+00:00", "updated_at": "2026-08-19 13:42:25.702096+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-ethics"], "entities": ["Harvard", "MIT", "Inside AI"], "alternates": {"html": "https://wpnews.pro/news/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation", "markdown": "https://wpnews.pro/news/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation.md", "text": "https://wpnews.pro/news/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation.txt", "jsonld": "https://wpnews.pro/news/harvard-study-shows-ai-undermines-leaders-judgment-in-innovation-evaluation.jsonld"}}