Google DeepMind Pilots Double-Blind Testing to Stop AI Models From Gaming Exams
The pilot tests a Gemini Flash Lite model against outside benchmarks using encryption so Google and its evaluators never see each other's data.
A New Way to Test AI Models
Google DeepMind said it has launched what it called the industry's first double-blind evaluation of a proprietary, frontier AI model, a pilot meant to prevent AI systems from gaming the exams used to judge their abilities. The lab is testing a version of its Gemini Flash Lite model against confidential outside benchmarks in a setup where neither Google nor the evaluators can see the other's data, according to Google DeepMind's announcement.
The project pairs Google DeepMind with the Singapore AI Safety Institute, the privacy-focused nonprofit OpenMined, the AI evaluation group AVERI, and the nonprofit MLCommons, which builds standardized benchmarks for AI systems, the company said.
The pilot targets a problem researchers call benchmark contamination: when an AI model has effectively seen test questions during training, its high scores on those tests become meaningless, much like a student who peeked at exam answers beforehand. Google DeepMind said that as AI models grow more capable, policymakers, researchers, and companies need confidence that published benchmark results reflect real ability rather than inflated numbers from a model that trained on the answers.
Encryption Instead of Trust
Until now, according to Google DeepMind, outside evaluators testing a company's AI model faced a tradeoff. They could hand their test questions to the company, risking that the questions leaked into future training data, or the company could hand over its model's underlying code and data, known as weights, risking exposure of a valuable trade secret.
Google DeepMind said its new system avoids that choice by running the test inside Confidential Space, a feature of Google Cloud's Confidential Computing service that uses cryptography to keep both sides sealed off from each other. The company said the setup lets it verify, through cryptographic proof, that the evaluator never sees Gemini's model weights and Google never sees the evaluator's test prompts.
Google DeepMind said the approach could matter most for sensitive evaluations, such as tests of a model's cybersecurity risks or reviews conducted by government bodies. The company said it hopes the pilot leads other AI developers to adopt similar cryptographic safeguards, though it did not say when or whether the method would expand beyond this initial test with Gemini Flash Lite, and it did not disclose specific benchmark results from the pilot in its announcement.