Google DeepMind pilots double-blind AI evaluations without sharing prompts or model weights Google DeepMind announced on August 27 that it is piloting what it calls the world's first double-blind AI evaluation system, designed to let external evaluators test frontier models without exposing their private test prompts or the company's model weights. The lab, co-founded by Demis Hassabis, has not disclosed participating models, evaluators, domains, security architecture, or results, leaving its industry-first claim independently unverified. Google DeepMind pilots double-blind AI evaluations without sharing prompts or model weights Google DeepMind says its pilot keeps external evaluation prompts and model weights private, though its first-of-its-kind claim remains independently unverified. By RuntimeWire Staff /author/runtimewire-staff ยท Published Primary source: Google DeepMind https://x.com/GoogleDeepMind/status/2092961763553677387 Why it matters Closed-model audits can force evaluators to expose private tests or developers to share valuable model weights. A credible double-blind process could protect both, but Google DeepMind has not disclosed enough about its pilot to assess its security, results or evaluator independence. Google DeepMind https://deepmind.google/?ref=runtimewire , the AI lab co-founded by Demis Hassabis https://www.ucl.ac.uk/about/search-faces-ucl/smart-tech-smarter-minds-demis-hassabis-ai-and-neuroscience-trailblazer?utm source=openai&ref=runtimewire , said on August 27 that it is piloting a double-blind evaluation system intended to let outsiders test frontier models without exposing either side's sensitive material. Google DeepMind's announcement on X https://x.com/GoogleDeepMind/status/2092961763553677387?ref=runtimewire Under the arrangement described in the company's announcement https://deepmind.google/blog/piloting-the-worlds-first-double-blind-ai-evaluations/?ref=runtimewire , external evaluators do not receive model weights and Google DeepMind does not see their private test prompts. Google DeepMind called the pilot an industry first. That claim has not been independently verified. The announcement, technical report https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/piloting-the-worlds-first-double-blind-ai-evaluations/double-blind-evaluations-technical-report.pdf?ref=runtimewire and other materials available for this report do not identify the participating models, external evaluators, evaluation domains, technical security architecture or results. They also do not provide enough detail to assess who controls the evaluation environment or whether the process is independent of the model owner. Hassabis co-founded DeepMind in 2010 after founding video-game company Elixir Studios and training as a cognitive neuroscientist. He built the lab around combining neuroscience, machine learning and computing to study intelligence and apply it to scientific problems. In a separate leadership change this month, Hassabis handed over his day-to-day operational responsibilities https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/?ref=runtimewire and became Google DeepMind's chair and Alphabet's chief scientist. Koray Kavukcuoglu became senior vice president of Google DeepMind, overseeing Gemini model development, frontier AI research, and the Gemini app and developer teams. The disclosure stops at the trust model External model testing creates a conflict over who must surrender sensitive material. Sending confidential prompts through a model provider's ordinary systems can expose a benchmark and make later tests easier to anticipate. Transferring model weights to an evaluator creates intellectual-property and security risks for the developer. Google DeepMind says its secure environment keeps test prompts and model weights hidden from the opposing party. That description establishes the intended confidentiality boundary, while leaving its implementation undisclosed. Google DeepMind has not named the hardware, software, access controls or verification services involved. It has also withheld the identities of the evaluators and models, along with any evaluation scores. Those omissions limit what can be concluded from the pilot. Double-blind handling could reduce prompt leakage and benchmark contamination, but the label alone does not establish that evaluators can verify the model being tested, inspect the approved computation or prevent the model owner from accessing logs and outputs. Evaluator independence also depends on governance and control of the infrastructure, neither of which Google DeepMind detailed in the announcement. OpenMined already tested one version of the idea A documented precursor came from OpenMined, the privacy-preserving software organization created by Andrew Trask https://openmined.org/blog/author/andrew-trask/?ref=runtimewire . OpenMined has spent years developing systems that allow organizations to approve computations over data and models they cannot directly inspect. In November 2024, OpenMined, Anthropic and the UK AI Safety Institute ran a secure-enclave experiment for AI evaluation https://openmined.org/blog/secure-enclaves-for-ai-evaluation/?ref=runtimewire . That proof of concept used open-source GPT-2 as a proxy for Anthropic's non-public models and a five-row CAMEL-bio sample as a stand-in for confidential UK AI Safety Institute data. OpenMined's documented setup placed the proxies in a confidential container on Microsoft Azure using AMD SEV-SNP and an NVIDIA H100. Its PySyft framework controlled how code and data entered the environment, while Microsoft Azure Remote Attestation and NVIDIA Remote Attestation supplied the attestation services. The parties reviewed the planned computation and checked the environment before authorizing the run. Bounded results were then released to an approved recipient, and the temporary environment was shut down. That 2024 experiment used simulated assets and a model with a public architecture. OpenMined described it as, to its knowledge, the first practical test of NVIDIA H100 secure enclaves for AI evaluation across two organizations. Its claim is separate from Google DeepMind's broader description of its current pilot as the first double-blind evaluation for frontier AI. The earlier work shows how one founder-led organization approached the confidentiality problem. It cannot establish how Google DeepMind's pilot works because Google DeepMind has not disclosed whether it uses a similar architecture, the same software or any of the same partners. The missing details determine independence For an external evaluation to carry weight, an evaluator needs confidence that the tested system is the model the developer claims it is, that the approved tests ran without interference and that sensitive prompts remained inaccessible. Readers also need enough methodology and results to interpret the findings. Google DeepMind has disclosed none of those operational details for this pilot. There is no public basis in the available materials to connect it to a particular Gemini model, Google Cloud service, hardware configuration, benchmark or national AI safety institute. Nor is there enough information to determine who verifies the environment or controls the release of results. The pilot arrives while Google DeepMind is shipping models across several domains. RuntimeWire reported this month that it released WeatherNext weights for cyclone forecasting /article/google-deepmind-open-sources-weathernext-cyclone-models and deployed Gemini Robotics 2 across Apptronik's Apollo 2 configurations /article/google-deepmind-gemini-robotics-2-apptronik-apollo-2 . A faster release schedule puts pressure on safety evaluations to keep pace without forcing every review into a negotiation over source code, model weights and benchmark ownership. Google DeepMind's announcement defines a useful objective for that process: keep the model and tests confidential from the opposing party. Assessing whether the pilot meets that objective will require Google DeepMind to disclose its architecture, participants, verification process and results.