메뉴
BL
The Decoder 43일 전

러시아 선전에 속는 AI 모델들, 에스토니아 연구소가 실험하다

IMP
7/10
핵심 요약

에스토니아어 연구소가 AI 언어 모델이 러시아 선전에 얼마나 취약한지 평가하는 새로운 벤치마크를 발표했습니다. 테스트 결과 앤스로픽의 클로드(Claude) 모델이 가장 우수한 성능을 보인 반면, 유럽의 대안을 표방하는 프랑스 기업 미스트랄(Mistral)의 모델은 가짜 뉴스를 걸러내는 데 가장 취약한 것으로 나타났습니다. 이는 악의적인 외국 세력이 AI를 허위 정보 유포에 악용하는 보안 위협이 현존하며, 모델별 대응 능력 편차가 큼을 시사합니다.

번역된 본문

러시아 선전에 속는 AI 모델들, 에스토니아 연구소가 실험하다 조나단 켐퍼(Jonathan Kemper) 2026년 6월 16일

에스토니아어 연구소(Institute of the Estonian Language)가 AI 언어 모델이 러시아 선전에 얼마나 취약한지 측정하는 벤치마크를 발표했습니다. 이번 테스트에서는 60개의 AI 모델을 대상으로 3개 언어로 작성된 75개의 질문이 활용되었습니다. 이 질문들은 14가지 선전 내러티브를 다루며, 중립적, 편향적, 조작적인 방식으로 표현되었습니다. 모델의 답변은 1점부터 5점까지 평가되며, 1점은 모델이 러시아의 주장을 그대로 반복했음을 의미합니다. 평가 모델로는 교정된 Claude Opus 4.5가 사용되었으며, 이는 허위 정보 전문 기관인 Propastop의 전문가들이 검증했습니다.

앤스로픽(Anthropic)의 클로드(Claude) 모델이 1위를 차지했으며, 그 뒤를 이어 엔비디아(Nvidia)의 Nemotron 3와 알리바바(Alibaba)의 Qwen 3.6 Plus가 이름을 올렸습니다. 반면 최신 모델인 Medium 3.5를 포함한 미스트랄(Mistral)의 모델들은 하위권에 머물렀습니다. 테스트 기간 동안 모델들은 웹 검색이나 다른 도구에 접근할 수 없었으므로, 이 벤치마크는 언어 모델 자체가 선전을 식별하고 거부하는 능력을 얼마나 잘 발휘하는지를 순수하게 측정합니다.

이러한 결과는 미스트랄이 36.67%의 일관된 허위 정보 생성률을 보인다는 뉴스가드(Newsguard)의 연구 결과와 일치합니다. 미국 및 중국 공급업체에 대한 유럽의 대안을 표방하며, 현재 200억 유로의 기업 가치로 30억 유로 규모의 펀딩 라운드를 협상 중인 프랑스 기업에게는 좋지 않은 모습입니다. 특히 미스트랄의 주력 모델들이 이미 경쟁사를 따라잡는 데 어려움을 겪고 있는 상황에서 더욱 타격이 큽니다.

위협은 현실입니다. '프라우다(Pravda)'와 같은 러시아 네트워크는 AI 시스템에 수백만 건의 허위 정보 기사를 의도적으로 주입하고 있습니다. 또한 최근 OpenAI는 독일 연방 선거에 앞서 ChatGPT를 사용해 선전을 퍼뜨린 러시아 캠페인을 차단하기도 했습니다.

원문 보기
원문 보기 (영어)
How easily can Russian propaganda fool AI models? A new benchmark finds out Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jun 16, 2026 The Institute of the Estonian Language has released a benchmark measuring how susceptible AI language models are to Russian propaganda. Sixty models were tested with 75 questions in three languages covering 14 propaganda narratives, phrased in neutral, biased, and manipulative ways. Each answer is scored on a scale of 1 to 5, where 1 means the model repeats Russian talking points. A calibrated Claude Opus 4.5 served as the evaluation model, validated by disinformation experts at the organization Propastop . Anthropic's Claude models claimed the top spots, followed by Nvidia's Nemotron 3 and Alibaba's Qwen 3.6 Plus . Mistral's models, including the newest Medium 3.5 , landed in the bottom third. The models had no access to web search or other tools during testing, so the benchmark only measures how well the language model itself can spot and reject propaganda. The results line up with a Newsguard study that found Mistral had a steady misinformation rate of 36.67 percent. That's a bad look for the French company, which positions itself as a European alternative to US and Chinese providers and is currently negotiating a 3 billion euro funding round at a 20 billion euro valuation. It's especially rough since Mistral's flagship models already struggle to keep up with the competition. Ad The threat is real. Russian networks like "Pravda" deliberately feed AI systems millions of disinformation articles. And OpenAI recently shut down a Russian campaign that used ChatGPT to spread propaganda ahead of Germany's federal election. Ad DEC_D_Incontent-1 AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Institute of the Estonian Language | FT Ask about this article… Search