Loading prices...
All news
Two glowing locked vaults facing each other across a tunnel of data particles, symbolizing a cryptographic double-blind evaluation where neither side sees the other's private data

Google pilots cryptographic double-blind testing for AI models

14:00 · 29.08.2026
Source: The Decoder
0

Google DeepMind is piloting the first double-blind evaluation of a proprietary frontier AI model, The Decoder reported. The project, run with the Singapore AI Safety Institute and other partners, uses cryptography to keep test questions and model weights hidden from each other.

The setup targets benchmark contamination: if a model has already seen a test's questions during training, its score tells you less than it looks like it does. Google DeepMind compares it to a test-taker who already knows the answers. External benchmark questions now stay locked inside a cryptographic "box," so a model can't later train on them to game the result.

Sensitive external evaluations used to force a choice. Evaluators could hand over their test prompts and let the model provider see the questions in advance, or the provider could hand over its model weights and risk its intellectual property. Google points to a recent case: the delayed ARC-AGI benchmark evaluation of Anthropic's Fable 5, complicated by Anthropic's 30-day data retention policy for its strongest models.

The double-blind setup is built to remove that choice entirely. Google runs it on Confidential Space, part of Google Cloud's confidential computing portfolio, which cryptographically verifies that both sides' data stays private. The evaluator never sees Gemini's weights. Google never sees the test prompts. Zero-logging protocols and contracts used to be what kept prompts confidential; this adds a cryptographic guarantee on top, closer to the kind of verification the US government's own secret AI benchmarking program is trying to build through policy instead of code.

  • Problem: benchmark contamination, when a model has seen test questions during training
  • Fix: cryptographic double-blind evaluation, no shared access to test prompts or model weights
  • Infrastructure: Confidential Space, part of Google Cloud's confidential computing portfolio
  • Pilot scope: a Gemini Flash Lite model tested against confidential benchmarks
  • Priority use case: cybersecurity and government-agency evaluations, where data sovereignty matters most

For the pilot, Google is testing a model from its Gemini Flash Lite line against confidential benchmarks under this setup.

Google says the method matters most for highly sensitive evaluations, cybersecurity testing or assessments run by government agencies, where independent organizations need to test a model rigorously without giving up data sovereignty or security. The push fits a pattern at DeepMind, whose leadership has separately called for an urgent AI oversight body ahead of more capable systems arriving.

Google published a technical report on the methodology and results, and says it hopes the approach becomes a standard other labs adopt, not a one-off pilot.

This piece is informational, not a recommendation to buy, sell, or hold any asset.

Published: 14:00 · 29.08.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!