A mother asks a chatbot whether acetaminophen is safe for her baby. It replies yes. What it may not explain is that infant and children’s versions can have different concentrations and require different doses. A true statement has just set up a potentially dangerous mistake.
That is the risk when people turn to chatbots with medical questions. The systems answer in seconds and can sound calm and authoritative. But they may invent a fact, play down an emergency or omit a warning. A user may delay care, take the wrong dose or follow a treatment that will not help.
“If it’s 90% right, what do you do?” said Marjorie Freedman, a principal scientist and associate director of the AI division at the USC Information Sciences Institute, or ISI, in USC Viterbi School of Engineering and the USC Stevens School for Computing and Artificial Intelligence. “That 10% can be really misleading” — and potentially harmful, she said.
The flaws can be hard to spot, especially in a long, polished response, making it difficult to know how good chatbots are. ALICE aims to close that safety gap. The project is developing better ways to test whether chatbot answers to medical questions are accurate, complete and safe to act on.
ALICE Looks Beyond a Right-or-Wrong Score #
Freedman leads ALICE, short for Assessing LLM Integrity for Clinical Engagement. Her collaborator Jonathan May is a research associate professor at the Thomas Lord Department of Computer Science, also within USC Viterbi and USC Stevens, and a principal scientist and the director of the artificial intelligence division at ISI.
ALICE is building tools to test answers that chatbots give to medical questions. Its goal is to find false statements and missing information while also asking whether an answer is clear, useful and appropriately urgent for patients, parents and health professionals.
Such a system could show where a chatbot needs improvement.
In addition to building the tools that could automate chatbot evaluation, the research team is creating the human-annotated data sets to validate the automated evaluators. The researchers studied hundreds of responses, some about 500 words long, from three chatbots. Their work explains why you can’t just have one person check an answer once and call it reliable.
In one study, the first people to review answers often missed real mistakes — mistakes that other reviewers, medical experts, and professional fact-checkers later caught. But even these experts didn’t always agree with each other about whether something was actually wrong. When a chatbot was added as an extra reviewer, it found more possible errors than the humans had but also missed some mistakes that the humans caught.
“There is a great amount of subtlety in trying to determine the objective truth to a simple statement,” May said.
Catching the Facts That Never Appear #
A second study focused on omissions. Medical experts reviewed answers to real patient questions and marked important information that was absent. Researchers then tested three ways to spot those gaps automatically.
The strongest approach caught more than 90% of the real omissions in one set of tests. But it also raised many false alarms. For now, that makes it better suited to flagging answers for expert review than replacing a doctor’s judgment.
“It’s already difficult to say whether a fact is true or not, but there’s an infinity of facts that are unsaid,” May said. Deciding which missing facts matter in a particular situation is “a really hard problem for us.”
Automated tools can help experts focus on the riskiest passages, Freedman said, but people must resolve uncertain cases. In addition to accuracy and completeness, tests should also reflect what patients and other healthcare consumers need: plain explanations, useful next steps and a clear signal when care is urgent.
Freedman said the team is preparing to share its lessons, data and methods so other researchers can continue working on the problem after the award ends. She also wants to dig deeper into what makes an answer complete, a quality that remains difficult to define.
Today, people may judge a chatbot by how polished, clear or sympathetic its answer sounds. None of those qualities proves the medical advice is reliable. May wants independent testing to give users a stronger basis for deciding which systems deserve their confidence.
“We would like to be able to provide a similar kind of certification to chatbots in high-risk contexts, like medicine and medical questions,” May said, “so that consumers can feel that they can trust the outputs.”
*The **Advanced Research Projects Agency for Health *is funding the two-year project with a $6.2 million award that runs through Sept. 17, 2026.
This research was, in part, funded by the Advanced Research Projects Agency for Health (ARPA-H). The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the United States Government.
Published on August 26th, 2026
Last updated on August 26th, 2026