Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training A randomized controlled study of 100 medical students found that a scaffolding-oriented multi-agent LLM AI Standardized Patient (AI-SP) platform improved final examination scores over a control condition using structured progressive information disclosure, with the largest gains in communication, empathy expression, and specific history-taking behaviors, according to an arXiv paper (arXiv:2609.10939v1). The system combines a patient agent for simulated dialog, a tutor agent giving Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. No significant difference in final diagnostic accuracy appeared between the multi-agent and control groups, and the authors released a multi-expert annotated dataset of transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes to support future research. arXiv:2609.10939v1 Announce Type: cross Abstract: Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient SP training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model LLM AI Standardized Patient AI-SP training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study N = 100 medical students , participants were assigned to either a multi-agent MA scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination OSCE based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.