Google DeepMind, Singapore's AI Safety Institute, OpenMined, AVERI, and MLCommons announced August 27 the completion of a pilot evaluation of Gemini 2.5 Flash-Lite using a double-blind setup inside Google Cloud's Confidential Space, where AVERI encrypted test prompts invisible to Google while Google kept model weights invisible to AVERI.
The setup targets benchmark contamination, in which a model trained on evaluation questions produces scores reflecting memorization rather than capability. The organizers said policymakers considering audit requirements under California SB 315 and the EU General-Purpose AI Code of Practice should require deep, secure third-party access to AI models.