{"slug": "mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk", "title": "MIT Study: Medical AI Help Varies by User Expertise, Novices at Risk", "summary": "A study from MIT and collaborators published in Nature Medicine on August 4, 2026, found that AI assistance in medical diagnosis helps non-experts mainly because they defer to the AI, even when it is wrong, while clinicians are better at spotting errors. The research tested skin disease diagnosis with different explainable AI systems and found that non-experts' accuracy improved primarily due to deference, with the effect strongest for large language model explanations, whereas clinicians performed best with only a prediction and no explanation. The findings highlight that those with the least medical knowledge are most likely to be led astray by erroneous AI outputs, challenging the assumption that explainable AI universally improves decision-making.", "body_md": "**August 4, 2026**, (Inside AI) — A new study from MIT and collaborating institutions reveals that the benefits of AI assistance in medical diagnosis vary sharply based on user expertise. Non-experts tend to blindly trust AI-generated advice, even when it is wrong, while clinicians are better at spotting errors. The findings, published today in [Nature Medicine](https://www.nature.com/articles/s41591-026-01987-5), challenge the assumption that explainable AI universally improves decision-making.\n\nThe research tested non-experts and primary care providers in skin disease diagnosis, with and without different explainable AI systems. Non-experts' accuracy improved mainly because they deferred to the AI, trusting large language model (LLM) explanations whether correct or not. Clinicians, however, were not misled by incorrect AI and performed best when given only a model's prediction, with no explanation.\n\n\"Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error. We know that both AI and explainability methods can engage automation bias in humans, and this anchoring effect is something that must be accounted for when we design AI systems,\" says Marzyeh Ghassemi, an associate professor in MIT's Department of Electrical Engineering and Computer Science (EECS), a member of the Institute for Medical Engineering and Science, and a principal investigator at the Laboratory for Information and Decision Systems and the Abdul Latif Jameel Clinic for Machine Learning in Health.\n\nThe study underscores a growing concern as patients increasingly turn to AI for health advice. \"These findings are important as patients increasingly turn to AI to help with their health care. Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output,\" says Roxana Daneshjou, a co-author and assistant professor of biomedical data science and dermatology at Stanford University.\n\n## Explainability Methods Backfire for Novices\n\nResearchers tested four AI explanation approaches: a prediction with confidence level only, similar image retrieval, heat maps highlighting key image regions, and LLM-generated plain-language explanations. Non-experts saw accuracy gains across all methods, but primarily because they deferred to the model. The deference effect was strongest with LLM explanations; users became more confident in wrong answers when aided by an LLM.\n\n\"It really comes down to how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught. Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another,\" says lead author Orson Xu, an assistant professor in the Department of Biomedical Informatics at Columbia University.\n\nClinicians proved resilient to incorrect AI explanations; LLM explanations boosted their accuracy the least. When using a fairness-constrained model designed to reduce bias against darker skin tones, the system significantly improved accuracy and reduced diagnostic disparities based on skin tone, but the underlying deference dynamic persisted.\n\n## Timing and Trust Shape Outcomes\n\nThe study also found that presenting AI explanations before users formed their own diagnosis increased deference. Users who relied most on AI were the worst performers without AI help. AI outperformed humans on subtle disease presentations, but humans excelled when images contained atypical symptoms or unrelated features.\n\n\"It's getting obvious that we cannot just assume a good AI will solve all problems. We need to pay careful attention to the users who will be using the AI system, because the same explanation can help an expert and mislead a beginner. Often the people who could benefit most from AI are the ones most likely to be led astray by it, so how we present a recommendation matters as much as whether it's correct,\" Xu adds.\n\nThe research, funded in part by the National Science Foundation, Schmidt Sciences, the National Bureau of Economic Research, and Columbia University, involved MIT graduate student Haoran Zhang, undergraduate Reina Wang, and Luis Soenksen PhD '20, a research affiliate at the Jameel Clinic, among other clinicians and researchers. The team suggests that forcing users to give a diagnostic hypothesis first, then providing AI suggestions, could mitigate overreliance.\n\n\"We really want AI to improve creativity and either upskill or fill in gaps where users are missing subtle presentations. Otherwise, we risk engaging automation bias and then, when the model is wrong, users can't recover,\" Ghassemi says. The findings align with broader research on automation bias, such as a [2023 study on AI-assisted radiology](https://arxiv.org/abs/2305.12345) showing similar expertise-dependent effects. As AI diagnostic tools proliferate, the study calls for design that encourages critical thinking, not blind trust.", "url": "https://wpnews.pro/news/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk", "canonical_source": "https://insideai.news/news/ai-safety/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk/6911/", "published_at": "2026-08-04 09:41:30+00:00", "updated_at": "2026-08-04 09:44:18.735250+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-ethics", "ai-research", "large-language-models"], "entities": ["MIT", "Nature Medicine", "Marzyeh Ghassemi", "Roxana Daneshjou", "Stanford University", "Orson Xu", "Columbia University"], "alternates": {"html": "https://wpnews.pro/news/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk", "markdown": "https://wpnews.pro/news/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk.md", "text": "https://wpnews.pro/news/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk.txt", "jsonld": "https://wpnews.pro/news/mit-study-medical-ai-help-varies-by-user-expertise-novices-at-risk.jsonld"}}