HomeAIDeepMind Pilots Double-Blind AI Evalua
AI

DeepMind Pilots Double-Blind AI Evaluations

Google DeepMind is partnering with safety institutes and research groups to test Gemini Flash Lite inside secure cloud environments.

WHAT YOU NEED TO KNOW
  • Google DeepMind is testing its Gemini Flash Lite model using double-blind evaluations.
  • The pilot is conducted in partnership with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons.
  • Google Cloud Confidential Space is used to ensure evaluators cannot see model weights and DeepMind cannot see evaluation prompts.

Google DeepMind launched a pilot to conduct double-blind evaluations of its proprietary artificial intelligence models, the company announced on August 27, 2026.

The initiative evaluates a Gemini Flash Lite model against confidential benchmarks inside a privacy-preserving cryptographic environment. DeepMind partnered with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons to carry out the tests, targeting the issue of benchmark contamination where models encounter evaluation questions prior to testing.

Cryptographic safeguards

External evaluations historically required a compromise between testing integrity and intellectual property protection. Evaluators had to provide their test prompts directly to the AI developer, which risked exposing the questions in advance. Alternatively, the model developer had to transfer model weights to the evaluator, creating intellectual property risks.

Google DeepMind said the new process relies on Confidential Space within Google Cloud’s Confidential Computing portfolio. Under this architecture, Google cannot view the evaluator’s confidential testing prompts, and external evaluators cannot inspect the Gemini model weights. While companies have previously relied on zero-logging protocols and contractual agreements to protect testing data, the pilot introduces technical cryptographic verification.

DeepMind researchers William Isaac, Sol Messing, and Kristian Lum noted that the setup allows independent groups, civil society organizations, and national AI Safety and Security Institutes to test frontier systems. The company pointed to sensitive domains such as cybersecurity and government evaluations as key use cases, and published a technical report detailing the pilot's methodology and findings.

Xentir Media
Xentir Media NewsroomSource-backed AI and technology coverage, drafted by Xentir's automated editorial system under fixed human-set rules. See our editorial policy and AI usage policy.
J
Jomon · Founder & EditorFounder and editor of Xentir Media. Sets the editorial rules the newsroom system runs under, and is accountable for its corrections. About Jomon · [email protected]
The Xentir Brief
The developments worth knowing — one useful email.
Get the Brief →