AI benchmarks have a trust problem and Google wants to fix it Google DeepMind is piloting a double-blind evaluation of a frontier AI model for the first time, using cryptographic protection through Confidential Space to prevent Google from seeing test questions and evaluators from seeing model weights. The pilot with the Singapore AI Safety Institute uses a Gemini Flash Lite model and could establish a new standard for tamper-proof AI benchmarks. Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights. The pilot project with the Singapore AI Safety Institute uses a Gemini Flash Lite and could set a new standard for tamper-proof AI benchmarks. The article AI benchmarks have a trust problem and Google wants to fix it https://the-decoder.com/ai-benchmarks-have-a-trust-problem-and-google-wants-to-fix-it/ appeared first on The Decoder https://the-decoder.com .