Piloting the world's first double-blind AI evaluations
Google and partners have launched the first double‑blind AI evaluation, using cryptographic confidential computing to keep both test data and model weights hidden from each other, preventing benchmark contamination. The pilot aims to restore trust in AI benchmarks by ensuring models cannot ‘peek’ at evaluation prompts, a step seen as crucial for safe, reliable AI deployment.