메뉴
BL
The Decoder • 28일 전

AI 벤치마크 오염 문제, 구글이 암호화로 해결 나서

IMP
7/10
핵심 요약

구글 딥마인드는 기밀 컴퓨팅(Confidential Space)을 활용한 이중 블라인드 평가를 통해 AI 모델이 평가용 테스트 문항을 사전에 학습하는 '벤치마크 오염(benchmark contamination)' 문제를 방지하는 파일럿 프로젝트를 진행한다. 평가자는 모델 가중치를, 구글은 테스트 문항을 볼 수 없도록 암호학적으로 보장하는 방식으로, 사이버보안·정부 평가 등 민감한 분야의 독립적 모델 검증에 새 표준이 될 것으로 기대된다.

번역된 본문

AI 벤치마크에는 신뢰 문제가 있고, 구글이 이를 해결하려 한다

구글 딥마인드는 암호학적 기법을 활용해 AI 모델이 테스트 문제를 사전에 미리 보지 못하도록 하려 한다. 싱가포르 AI 안전 연구소(AI Safety Institute) 및 다른 파트너들과 함께 진행하는 파일럿 프로젝트에서 제미나이(Gemini) 모델에 대해 이중 블라인드(double-blind) 테스트를 실시한다.

시험 응시자가 문제를 미리 알고 있다면 만점을 받아도 아무 가치가 없다. 구글 딥마인드는 이 비유로 AI 모델 평가의 핵심 문제인 '벤치마크 오염(benchmark contamination)'을 설명한다. 모델이 학습 과정에서 이미 테스트 문제를 접했다면, 그 결과를 그만큼만 신뢰할 수 있을 뿐이다.

이를 방지하기 위해 구글 딥마인드는 독점 프론티어 AI 모델에 대한 최초의 이중 블라인드 평가를 시작한다고 밝혔다. 외부 테스트는 암호화된 '상자'에 잠겨 보관되어, 모델이 이후 그 문제들을 활용해 테스트에 맞춰 최적화하는 것을 막는다. 이번 파일럿에서 구글은 제미나이 플래시 라이트(Gemini Flash Lite) 계열 모델을 기밀 벤치마크와 대상으로 테스트한다.

이 방법이 해결하려는 트레이드오프

딥마인드에 따르면, 지금까지 고도로 민감한 외부 평가는 타협이 필요했다. 평가자가 테스트 프롬프트를 넘기면 모델 제공자가 문제를 미리 보게 되고, 반대로 제공자가 모델 가중치를 넘기면 지식재산권을 위험에 빠뜨리는 상황이었다. 최근 이런 딜레마의 사례로는 앤스로픽(Anthropic)의 Fable 5에 대한 ARC-AGI 벤치마크 평가 지연을 들 수 있다. 해당 AI 기업이 최강 모델에 대해 30일 데이터 보존 정책을 고수했기 때문이다.

이중 블라인드 평가는 이런 트레이드오프를 제거하는 것을 목표로 한다. 구글은 구글 클라우드의 기밀 컴퓨팅 포트폴리오 중 '컨피덴셜 스페이스(Confidential Space)'를 활용한다. 이 설정은 외부 테스트 데이터와 모델 모두 각 소유자에게만 비공개로 유지됨을 암호학적으로 검증한다. 평가자는 제미나이 가중치를 절대 볼 수 없고, 구글은 테스트 프롬프트를 절대 볼 수 없다.

지금까지는 제로 로깅(zero-logging) 프로토콜과 계약상의 안전장치로 외부 프롬프트의 기밀을 유지해 왔다. 여기에 기술적·암호학적 보호를 더하는 것은 안전한 모델 평가에 있어 큰 진전이라고 회사는 말한다.

민감한 분야에서 특히 중요한 이유

이 암호학적 증명은 오염을 방지하고 민감 데이터를 보호하는 것을 목표로 한다. 딥마인드는 이것이 사이버보안이나 정부 기관이 실시하는 테스트처럼 고도로 민감한 평가에서 가장 중요하다고 밝혔다. 독립 기관들이 데이터 주권이나 보안을 포기하지 않고도 첨단 모델을 엄격하게 테스트할 수 있게 된다.

구글은 이 노력이 모델 감독의 새로운 표준을 정하고, 업계가 더 신뢰할 수 있고 널리 신뢰받는 AI 시스템을 구축하는 데 도움이 되기를 희망한다. 구글은 방법론과 결과에 대한 세부 내용을 기술 보고서에서 공개했다.

원문 보기
원문 보기 (영어)
AI benchmarks have a trust problem and Google wants to fix it Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 28, 2026 GPT-Image-2 prompted by THE DECODER Google Deepmind wants to use a cryptographic method to stop AI models from seeing test questions in advance. A pilot project with the Singapore AI Safety Institute and other partners runs a double-blind test on a Gemini model. If a test-taker knows the questions ahead of time, even a perfect score is worthless. Google Deepmind uses this image to describe a core problem in evaluating AI models: benchmark contamination. If a model has already seen the test questions during training, you can only trust the results so far. To prevent that, Google Deepmind says it's launching the first double-blind evaluation of a proprietary frontier AI model. External tests stay locked in a cryptographic "box," so a model can't later use those questions to optimize itself specifically for the test. For the pilot, Google is testing a model from the Gemini Flash Lite line against confidential benchmarks. The tradeoff the method aims to fix Highly sensitive external evaluations used to require a compromise, Deepmind says. Either the evaluators handed over their test prompts, which let the model provider see the questions in advance. Or the provider handed over its model weights and risked its intellectual property. A recent example of this dilemma was the delayed evaluation for the ARC-AGI benchmark of Anthropic's Fable 5 , since the AI company enforces a 30-day data retention policy for its strongest models. The double-blind evaluation is meant to eliminate that tradeoff. Google uses Confidential Space from Google Cloud's confidential computing portfolio to do it. The setup cryptographically verifies that both the external test data and the model stay private to their respective owners. The evaluator never sees the Gemini weights, and Google never sees the test prompts. Until now, zero-logging protocols and contractual safeguards kept external prompts confidential. Adding technical and cryptographic protection is a big step forward for secure model evaluation, the company says. Why this matters most for sensitive areas The cryptographic proof aims to prevent contamination and protect sensitive data. Deepmind says this matters most for highly sensitive evaluations, such as cybersecurity or tests run by government agencies. Independent organizations could rigorously test advanced models without giving up data sovereignty or security. Google hopes the effort sets a new standard for model oversight and helps the industry build more reliable and widely trusted AI systems. Google lays out the details on methodology and results in a technical report . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->