Piloting the world's first double-blind AI evaluations
Google is introducing the world’s first double-blind evaluation of a proprietary frontier-class artificial intelligence model to prevent benchmark contamination. Historically, external evaluations forced a compromise where evaluators risked exposing their test questions to model providers, or providers risked exposing their intellectual property by sharing model weights.
To solve this, the pilot uses Confidential Space within Google Cloud to create a cryptographically secure environment. This setup allows independent organizations to rigorously test advanced models while ensuring that neither the evaluator can view the model weights nor the provider can see the test prompts.
For the pilot, Google is partnering with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons to test a Gemini Flash Lite model against confidential benchmarks. This approach maintains the privacy of both parties and prevents models from peeking at test questions in advance, which can artificially inflate scores.
This technical safeguard is critical as AI models become more capable and are evaluated for sensitive areas like cybersecurity and government use. Policymakers, researchers, and enterprises require trustworthy benchmarks to accurately reflect true model capabilities and safety.