15:16
2026-07-29
amazon.science
artificial-intelligence
A new benchmark for evaluating patient-facing health AI agents
OpenAI released PatientAgentBench, a clinician-vetted benchmark for evaluating patient-facing AI agents across six safety and workflow dimensions, finding that even capable foundation models fail to mโฆ