Advancing brain tumor research with privacy-first AI Google Cloud is collaborating with MLCommons through the MedPerf initiative to use Confidential Computing for privacy-first benchmarking of medical AI models, enabling validation on diverse real-world patient data without exposing model code or patient data. The technology, which runs on Google Cloud's A3 machines with NVIDIA H100 GPUs, is already being used in the Federated Tumor Segmentation (FeTS) initiative to validate brain tumor AI models on private MRI data, addressing performance gaps where a model might be 95% accurate at one site but only 63% at another. The intersection of medicine and AI has led to remarkable innovations. However, developers now face the thorny challenge of building robust medical AI tools that have been tested and evaluated on diverse, real-world patient data while also protecting patient privacy. At Google Cloud, our approach combines strategic collaboration with Confidential Computing. To help protect both patient privacy and AI models during validation, we’re collaborating with MLCommons https://mlcommons.org/ through the MedPerf initiative https://www.medperf.org/ . First announced at Google Cloud Next earlier this year, this partnership uses Confidential Computing https://cloud.google.com/security/products/confidential-computing to establish a secure clean room for benchmarking AI models in real-world settings. MLCommons, a global community with over 125 members across tech and academia, launched MedPerf https://mlcommons.org/medical/ in 2023 to standardize the evaluation of medical AI. MedPerf, an open-source platform for benchmarking AI models, has advanced clinical research using federated evaluation https://en.wikipedia.org/wiki/Federated learning to test models. By using Google Cloud Confidential Space https://docs.cloud.google.com/confidential-computing/confidential-space/docs/confidential-space-overview , proprietary AI models can be evaluated inside hardware-isolated Trusted Execution Environments TEEs . This special virtual machine encrypts memory in-use and hardens the operating system, so none of the parties — the hospital or research institution, other participants, or Google — can see model code or patient data while it's evaluated. Medical AI benchmarking is compute-heavy, so the Confidential VM extends beyond the CPU to the GPU. To protect model weights and patient data even during GPU-accelerated inference, MedPerf runs on Google Cloud's A3 machine series with NVIDIA H100 GPUs, which pairs Intel TDX technology on the CPU with NVIDIA Confidential Computing https://www.nvidia.com/en-us/data-center/solutions/confidential-computing/ on the GPU. Before any patient data is released into the workload, the system provides cryptographic proof that only the approved code is running on genuine Confidential Computing hardware and that the environment has been properly hardened. This technology is already driving critical research through the Federated Tumor Segmentation https://www.med.upenn.edu/cbica/fets/ FeTS initiative. Brain tumors, such as glioblastomas, are rare, making it difficult for any single hospital to collect enough data for high-accuracy AI training. Compounding the problem, a model that performs perfectly in one hospital can struggle in another due to differences in patient demographics, data acquisition techniques, and even in equipment. Working with visionary researchers like Indiana University’s Dr. Spyridon Bakas https://medicine.iu.edu/faculty/64865/bakas-spyridon , Northwestern University’s Dr. Yury Velichko https://www.feinberg.northwestern.edu/faculty-profiles/az/profile.html?xid=30630 , and the University of Alberta, Canada’s Dr. Amber Simpson https://apps.ualberta.ca/directory/person/asimpso4 , MedPerf on Google Cloud is validating AI models on private brain MRI data from around the world, and identifying potential performance gaps. For example, a model might be 95% accurate at one site but only 63% accurate at another. Our collaborative approach demonstrates that when an AI tool reaches a clinician, it has been proven to work across a truly representative patient population. The impact of this collaboration is best summarized by those on the front lines of clinical research. "My experience testing federated learning on Google Cloud has shown that the future of medical AI lies in secure, scalable, and collaborative cloud environments," said Dr. Yury Velichko, associate professor, Radiology, Northwestern University. "Moving beyond the controlled lab setting to test these workflows in a production-ready infrastructure provided a unique opportunity to evaluate the performance and security of federated learning in real-world clinical applications.” "Medical AI holds enormous promise for patients around the world, but that promise can only be realized if clinicians, researchers, and regulators can trust the benchmarks we use to evaluate it,” said Alexandros Karargyris, MedPerf lead, MLCommons. By bringing MedPerf onto Google Cloud's Confidential Computing infrastructure, we have taken a major step toward a future where AI models can be rigorously tested on real patient data — without compromising privacy, intellectual property, or benchmark integrity.” The collaboration between MLCommons and Google Cloud represents a fundamental shift toward privacy by design in healthcare AI. By making it easier to securely share and evaluate data and models, we are clearing the path for faster, safer, and more equitable medical breakthroughs. Research institutions and healthcare model developers interested in using the MedPerf platform on Google Cloud should contact medical@mlcommons.org mailto:medical@mlcommons.org or your Google Cloud account team.