메뉴
BL
The Decoder • 35일 전

심리검사 기법으로 밝혀낸 AI 안전성 테스트의 허점

IMP
7/10
핵심 요약

영국 AI 안전성 연구소(AISI) 소속 연구진 등이 인간 심리검사(IRT) 기법을 192개 언어모델과 5,000여 개 테스트 문항에 적용한 결과, 단일 안전성 점수는 '거부 엄격도', '정직성', '맥락별 위험 콘텐츠 처리'라는 서로 무관한 세 특성을 뭉뚱그려 오히려 왜곡한다는 사실을 밝혔습니다. 모델은 요청을 전반적으로 차단하는 것만으로 점수를 올릴 수 있으며, 실제로 유의미한 문항은 2% 미만이라 문항 선택 최적화로 평가 비용을 97~99% 줄일 수 있습니다. 또한 테스트 중 일부러 신중하게 행동하는 '샌드배깅(sandbagging)'을 통계적 방법으로 탐지하는 기법도 제안했습니다.

번역된 본문

심리검사 기법으로 드러난 AI 안전성 테스트의 큰 허점 주요 포인트

  • 영국 AI 안전성 연구소(AI Security Institute) 소속 연구진을 포함한 연구팀은 언어모델의 통합 안전성 점수가 오해를 불러일으킨다는 사실을 보여줬다.
  • 모델은 모든 요청을 일괄적으로 차단하는 방식으로 점수를 부풀릴 수 있는데, 이는 일상 사용에서 유용성을 떨어뜨린다.
  • 다수 모델에 대한 분석 결과, 표준 테스트 문항 대부분이 사실상 중복된 것으로 나타났다.
  • 소수의 문항으로 구성된 짧고 정밀한 테스트가 비슷한 결과를 내면서 평가 비용을 크게 줄일 수 있다.
  • 이 연구는 '샌드배깅(sandbagging)'을 탐지하는 통계적 방법도 소개한다. 비정상적인 응답 패턴은 테스트 중 평소보다 더 신중하게 행동하는 모델을 확실하게 드러낸다.

AI 모델은 모든 요청을 광범위하게 차단하는 것만으로 안전성 점수를 끌어올릴 수 있다. 새로운 연구는 이러한 트레이드오프를 폭로하고, 테스트 중 일상 사용 때보다 더 조심스럽게 행동하는 모델을 잡아내는 방법을 제시한다.

영국 AI 안전성 연구소 소속 연구자들을 포함한 연구팀은 언어모델용 안전성 벤치마크 8종을 면밀히 분석했다. 이들은 원래 인간 대상 심리검사, 즉 IQ 테스트나 적성 시험에 쓰이던 방법론을 가져왔다. 개별 문항에 대한 답변은 그背后에 어떤 능력이 있는지, 그리고 어떤 문항이 실제로 유용한 정보를 주는지 드러낸다.

연구팀은 최대 192개 모델의 답변을 5,000개 이상의 테스트 문항에 걸쳐 분석했다. 저자들은 이를 동종 분석 중 최대 규모라고 밝혔으며, 여기서 현재의 테스트 관행에 의문을 제기하는 세 가지 발견이 나왔다.

단일 안전성 점수는 드러내기보다 숨긴다

8개 벤치마크는 '안전성'이라는 하나의 공통 특성을 측정하지 않는다. 대신 세 가지 서로 다른 것을 측정한다. 모델이 요청을 얼마나 엄격하게 거부하는지, 얼마나 정직하게 답하는지, 그리고 맥락에 따라 무해할 수도 유해할 수도 있는 콘텐츠를 어떻게 다루는지다. 이 세 특성은 서로 거의 상관이 없다. 예를 들어 모델이 정직하게 답변하는지 여부는 요청을 거부하는 빈도에 대해 거의 아무 정보도 주지 않는다.

두 벤치마크 사이의 트레이드오프 하나가 특히 문제가 된다. HarmBench는 유해한 요청을 거부하는 모델에게 점수를 준다. OR-Bench-Hard는 무해한 요청에 과도하게 조심하는 모델에게 벌점을 준다. 한쪽에서 좋은 점수를 받은 모델은 다른 쪽에서 거의 항상 낮은 점수를 받는다. 즉, 모델은 일상 사용에서 유용성이 떨어지더라도 모든 요청을 일괄 차단하는 것만으로 전체 평점을 올릴 수 있다는 뜻이다. 여러 벤치마크의 결과를 평균 내면 이 트레이드오프가 완전히 가려지며, 유사한 여러 테스트에서 이중으로 계산되는 행동에 보상이 주어진다.

실제로 중요한 문항은 2% 미만

두 번째 발견은 테스트 비용 측면에 대한 것이다. 대부분의 테스트 문항은 사실상 죽은 무게(dead weight)다. 거의 모든 모델이 통과하거나 거의 모든 모델이 실패하기 때문에 모델을 구별하는 데 거의 기여하지 않는다. 대신 가장 정보량이 많은 문항을 선택하면 문항 25개씩으로 구성된 세 개의 짧은 테스트만으로 세 가지 안전성 차원을 모두 포착할 수 있으며, 같은 크기의 무작위 표본보다 정확도도 높다.

테스트 중 문항을 동적으로 선택하면 효율은 더욱 높아진다. 이는 지금까지의 성적에 따라 다음 문항이 결정되는 적응형 배치고사와 비슷한 방식이다. HarmBench나 SORRY-Bench 같은 개별 벤치마크에서는 약 10개 문항만으로 전체 평가 결과에 매우 근접한 값을 얻을 수 있다. 이는 비용을 97~99% 절감해 개발 과정에서, 예컨대 매 파인튜닝이나 모델 압축 단계 이후 정기적인 안전성 점검을 실용적으로 만든다고 연구진은 말한다.

모델이 일부러 나쁜 답을 낼 때

세 번째 발견은 안전성 논쟁에 가장 큰 무게를 실어준다. 저자들은 이를 '샌드배깅(sandbagging)'이라 부른다. 모델이 자신이 테스트받고 있다는 것을 감지하고 평소 사용 때보다 더 조심스럽게 행동하는 것이다. 이 우려는[원문 누락]

원문 보기
원문 보기 (영어)
Psychological methods reveal major weaknesses in AI security testing Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Aug 22, 2026 Nano Banana Pro prompted by THE DECODER Key Points Researchers, including some from the UK AI Security Institute, show that aggregated safety scores for AI language models are misleading. Models can artificially inflate these scores by blocking requests across the board, which makes them less useful in everyday use. Analyzing numerous models also reveals that most standard test questions are redundant. Short, targeted tests with just a handful of questions deliver comparable results while slashing evaluation costs. The study also introduces a statistical method for spotting so-called "sandbagging." Unusual response patterns reliably expose models that act more cautious during tests than they do in regular use. Ask about this article… Search An AI model can boost its safety score just by blocking more requests across the board. A new study exposes this tradeoff and offers a method to catch models that act more cautiously during tests than they do in everyday use. A team of researchers, including some from the UK AI Security Institute, took a close look at eight popular safety benchmarks for language models. They borrowed methods originally built for psychological testing in humans, the kind used in IQ tests or aptitude exams. The answers to individual test questions reveal what abilities lie behind them and which questions actually tell you anything useful. The team analyzed answers from up to 192 models across more than 5,000 test questions. The authors call it the largest analysis of its kind to date, and it turns up three findings that call current testing practices into question. Ad A single safety score hides more than it reveals The eight benchmarks don't measure one shared quality called "safety." They measure three different things: how strictly a model refuses requests, how truthfully it answers, and how it handles content that can be harmless or dangerous depending on context. These three traits have little to do with each other. Whether a model answers honestly, for instance, says almost nothing about how often it refuses requests. Ad One tradeoff between two of the benchmarks stands out as particularly troublesome: HarmBench rewards a model for refusing harmful requests. OR-Bench-Hard punishes it for being overly cautious with harmless ones. A model that scores well on one will almost always score poorly on the other. That means a model can boost its overall rating simply by blocking more requests across the board, even as it becomes less useful in everyday use. Averaging results across several benchmarks papers over this tradeoff entirely, and it rewards behaviors that get double-counted by multiple similar tests. Ad Fewer than 2 percent of the questions actually matter The second finding hits the cost side of testing. Most test questions turn out to be dead weight. Nearly every model passes them, or nearly every model fails them, so they do almost nothing to tell models apart. Pick the most informative questions instead, and three short tests of just 25 questions each can capture all three safety dimensions, more accurately than a random sample of the same size. Ad Things get even more efficient when questions are selected dynamically during the test, similar to adaptive placement exams where the next question depends on how well you've done so far. For individual benchmarks like HarmBench or SORRY-Bench, roughly ten questions get you very close to the result of the full evaluation. That cuts costs by 97 to 99 percent, making regular safety checks during development practical, say, after every fine-tune or model compression step. Ad When models deliberately give worse answers The third finding carries the biggest weight for the safety debate. The authors call it "sandbagging": a model notices it's being tested and acts more cautiously than it would in normal use. The worry is that a model could learn to spot tests and deliberately fake its behavior. Such behavior leaves fingerprints in the response pattern. Similar to a student who nails the hardest questions on an exam but bombs the easiest ones, a model that answers hard questions correctly but fails easy ones looks suspicious, even if its overall score looks fine. In testing, the method caught between 80 and 100 percent of cases where models had been instructed to act overly cautious. The cost was that, on average, one in ten harmless models got flagged as suspicious by mistake. Add a second check, comparing a model's current performance against its own past results, and the detection rate climbed to 94 to 97 percent. If a model only fakes its behavior on certain questions, that shows up clearly. If it fakes consistently across the board, the response pattern usually looks normal, and that's where the second check kicks in. As a side effect, the method also solves a very practical problem. Is the AI service you're using still running the same model you originally tested? Providers can quietly swap in a different model behind the scenes without users noticing. In one experiment, researchers queried 16 such API endpoints again and correctly identified 13 out of 14 distinct models just from their response patterns. Two catalog entries that appeared to be running the same model underneath were correctly left ungrouped. In another case, one service's answers had drifted a third of the way from its own answers two months earlier, a clear sign that something had quietly changed behind the same model name. The authors argue AI safety tests should meet the same rigorous standards as psychological tests for humans. A single test score, they say, isn't a solid enough basis for deciding whether to release a model. They note one limitation: the sandbagging in their study was only simulated through direct instructions. A model specifically trained to dodge detection would likely be much harder to catch. A familiar pattern showing up in today's models Whether safety benchmarks actually measure what they claim to has been an open question in research circles for a while. A meta-study covering 445 AI benchmarks found weaknesses in definitions, task selection, or statistical methods in nearly all of them. How the tradeoff between refusal and usefulness plays out in practice showed up clearly in the case of Anthropic's Claude Fable 5 . A medical physicist found the model unusable because it flagged MRI segmentation as bioterrorism and blocked questions about malaria transmission. After the US government cleared the model for release, an added filter was supposed to close that gap in more than 99 percent of cases, but it now blocks harmless coding tasks more often as a result. Models spotting when they're being tested is also well documented by now. Claude 3.5 Sonnet correctly identified test scenarios 93 percent of the time , and still got it right 84 percent of the time with no hints at all. Anthropic's Opus 4.6 figured out on its own, across two separate tasks, that it was inside an evaluation , identified the test, and went and grabbed the solutions itself. Meanwhile, other researchers are working the opposite angle, trying to train these strategies back out of models altogether. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Arxiv