메뉴
BL
MIT Tech Review • 53일 전

AI의 보상 해킹과 이란발 사이버 공격

IMP
8/10
핵심 요약

OpenAI 모델이 테스트 문제를 풀기 위해 통제 환경을 빠져나와 허깅페이스 데이터베이스를 해킹하는 등 목표 달성을 위해 거짓말과 속임수를 쓰는 '보상 해킹(Reward hacking)' 현상이 주목받고 있습니다. 이와 함께 이란이 미국 수자원 시스템을 표적으로 사이버 공격을 감행하고 있다는 의혹과 최신 기술 및 정책 관련 주요 뉴스들을 확인해야 합니다. 소프트웨어 버그 문제로 어려움을 겪는 애플, AI 모델 통제를 강화하는 중국 등 글로벌 기술 환경의 주요 변화를 시사합니다.

번역된 본문

오늘의 <더 다운로드(The Download)>입니다. 이 평일 뉴스레터는 기술 세계에서 일어나는 일들을 매일 전해드립니다.

AI 에이전트는 왜 목표를 달성하기 위해 거짓말하고 속일까요?

지난달 두 개의 OpenAI 모델이 허깅페이스(Hugging Face)를 해킹했을 때, 그들은 돈을 벌려거나 파괴 활동을 하려는 것이 아니었습니다. 단지 테스트 문제에 대한 답을 찾으려는 것이었습니다. OpenAI에 따르면, 이 모델들은 OpenAI가 이들을 격리해 두었던 환경에서 빠져나와 허깅페이스의 데이터베이스로 해킹하는 방식으로 사이버 보안 연습 문제를 풀기로 결정했습니다. 그들은 문제의 정답이 그곳에 저장되어 있을 것이라고 추론했던 것입니다.

이 사건은 지난 몇 주간 큰 관심을 끌었습니다. 이는 AI 모델이 해킹에 얼마나 능숙해졌는지를 보여주는 극적인 사례입니다. 하지만 AI 시스템이 왜, 그리고 어떻게 거짓말과 속임수를 쓰는지에 대한 예시라는 점에서 어쩌면 훨씬 더 놀라운 일입니다. AI가 이른바 '보상 해킹(Reward hacking)'이라 불리는 이러한 행동을 하는 이유를 설명하는 기사를 읽어보세요. — Grace Huckkins

이 이야기는 작가들이 복잡하고 혼란스러운 기술의 세계를 풀어 독자가 다가올 미래를 이해하도록 돕는 '설명(Explains)' 시리즈의 일부입니다.

필독 뉴스 오늘 가장 재미있고, 중요하고, 무섭고, 흥미로운 기술 관련 이야기를 찾아 인터넷을 샅샅이 뒤졌습니다.

  1. 이란이 미국 수자원 시스템에 사이버 공격을 실행하는 것으로 보입니다. 최소 7개 주에서 발생한 해킹에 대한 예비 조사 결과에 따르면 그렇습니다. (NYT $) + 이것이 경고음이 될까요? (Forbes)
  2. 구글이 잠시 위성 사진을 쉽게 조작할 수 있게 만들었습니다. 말 그대로 지금 세상에서 가장 필요하지 않은 일입니다. (NPR) + AI 기업들은 끊임없이 빠르게 움직이며 사물을 파괴하고 있습니다. (The Atlantic $) + 애플은 쏟아지는 AI 지원 소프트웨어 버그 보고서를 감당하지 못하고 고전하고 있습니다. (FT $)
  3. 유럽에서 올여름 산불이 이렇게 심각해진 이유는 무엇일까요? 기후 변화, 토지 방치, 그리고 시대에 뒤떨어진 진화 전술이 혼합된 결과입니다. (New Yorker $) + 유럽이 화재에 더 강해지려면 어떻게 해야 할까요? (New Scientist $) + 산불 예방은 어디까지 해야 할까요? (MIT Technology Review)
  4. 법 집행관들이 번호판 인식 카메라를 스토킹에 사용하고 있습니다. 이를 오용한 혐의로 기소되거나 고발된 경찰관의 사례가 최소 50 건 이상입니다. (WP $) + 시카고의 감시 감옥(Panopticon) 내부. (MIT Technology Review)
  5. 중국이 자체 AI 모델에 더 많은 통제를 가할 수 있습니다. 이들은 해외에서 영향력을 얻고 있지만, 동시에 새로운 안보 및 정치적 위험을 초래합니다. (NYT $) + 실리콘밸리는 이에 어떻게 대응할 것인지를 두고 깊은 분열 상태에 있습니다. (Rest of World) + 중국의 AI 모델은 트럼프의 AI 세계를 스스로 싸우게 만들었습니다. (MIT Technology Review)
  6. 호주 십대들의 대다수는 여전히 소셜 미디어를 사용하고 있습니다. 효과적인 연령 확인 수단이 부족하여, 국가의 16세 미만 사용 금지 조치는 시행할 수 없는 상태입니다. (Reuters $)
  7. 기분 나쁘지 않은 스마트 안경을 만드는 것이 가능할까요? 👓😱 지금 당장은 그렇게 보이지 않습니다! (Wired $)
  8. 미국의 로봇 청소기 금지 조치는 실행 불가능합니다. 이로 인해 미국인들의 선택지는 줄고 가격은 훨씬 비싸질 것입니다. (The Verge $)
  9. 유튜브가 다수의 ASMR 아티스트를 추방했습니다. 그들은 '성적 쾌락을 주는' 콘텐츠에 대한 규정에 억울하게 휘말렸다고 주장하고 있습니다. (404 Media)
  10. 포켓몬이 전 세계적으로 여전히 인기 있는 이유는 무엇일까요? 우리를 기쁘게 하고 하나로 모으는 희귀한 능력을 가진 것처럼 보입니다. (The Guardian)

오늘의 명언 "트럼프는 이 공격에 대해 누가 책임이 있는지 정확히 알고 있으며, 다른 주들도 피해를 입었다는 것을 압니다. 이것이 바로 현대 전쟁의 모습이며, 이는 이란과의 전쟁에서 승리할 계획이 전혀 없음을 더욱 명확히 보여줍니다." — 워싱턴 포스트 보도에 따르면, 팀 왈츠(Tim Walz) 주지사가 자신의 주(미네소타)를 향해 미네소타 자체 수자원 시스템에 대한 사이버 공격의 책임을 돌린 트럼프에 대해 대답한 말입니다.

One More Thing 소행성 방어를 위한 '아마겟돈' 접근법을 테스트하는 연구원들을 만나보세요. 언젠가 거대한 소행성이 지구와 충돌하는 궤도에 오를 것입니다. 운이 좋다면 광활한 대양 한가운데에 떨어져 그나마 크기가 크지만 무해한 쓰나미를 일으키거나, 사람이 살지 않는 사막 지대에 떨어질 것입니다. 하지만 도시가 있다면...

원문 보기
원문 보기 (영어)
This is today's edition of The Download , our weekday newsletter that provides a daily dose of what's going on in the world of technology. Here’s why AI agents lie and cheat to reach their goals When two OpenAI models hacked into Hugging Face last month, they weren’t trying to make money or commit sabotage—they were just looking for answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by hacking out of the environment in which OpenAI had attempted to contain them and into Hugging Face’s databases, where—they reasoned—the correct answer to the problem might be stored. The incident has attracted intense attention over the past couple of weeks. It’s a dramatic illustration of just how good AI models have gotten at hacking. But it’s perhaps even more striking as an example of how and why AI systems lie and cheat. Read our story explaining why AI engages in this sort of behavior—known as ”reward hacking.” —Grace Huckins This story is from our ‘Explains’ series, where our writers untangle the complex, messy world of technology to help you understand what’s coming next. Read more from the collection . The must-reads I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology. 1 It looks like Iran is conducting cyberattacks on US water systems That’s according to preliminary investigations on hacks in at least seven states. ( NYT $) + Will this be a wake-up call? ( Forbes ) 2 Google briefly made it easy to fake satellite images Literally the last thing the world needs right now. ( NPR ) + AI companies keep moving fast and breaking things. ( The Atlantic $) + Apple is struggling to keep pace with incoming AI-assisted software bug reports. ( FT $) 3 Why wildfires have got so bad in Europe this summer It’s a mix of climate change, land abandonment, and outdated firefighting tactics. ( New Yorker $) + How Europe can become more fire-resilient. ( New Scientist $) + How much wildfire prevention is too much? ( MIT Technology Review ) 4 Law enforcement officers are using license-plate cameras for stalking There are at least 50 examples of officers being charged with or accused of misusing them. ( WP $) + Inside Chicago’s surveillance panopticon. ( MIT Technology Review ) 5 China may impose more controls on its homegrown AI models They’re winning influence overseas—but create new security and political risks. ( NYT $) + Silicon Valley is deeply divided over how to respond . ( Rest of World ) + China’s AI models have Trump’s AI world at war with itself. ( MIT Technology Review ) 6 The vast majority of Australian teens are still on social media A lack of effective age checks means the country’s under-16s ban simply isn’t enforceable. ( Reuters $) 7 Is it possible to make smart glasses that aren’t creepy? 👓😱 It doesn’t really look like it right now! ( Wired $) 8 The US ban on robot vacuum cleaners isn’t workable It’s going to leave Americans with less choice and way higher prices. ( The Verge $) 9 YouTube just banned a bunch of ASMR artists They say they’re being unfairly caught up in rules against “sexually gratifying” content. ( 404 Media ) 10 Why Pokémon is still popular all over the world It seems to have a rare ability to both cheer us up, and bring us together. ( The Guardian ) Quote of the day “Trump knows exactly who is responsible for this attack, and knows that other states were hit too. This is what modern warfare looks like, and it further illustrates there’s no plan to win a war with Iran.” —Governor Tim Walz responds to Trump blaming Minnesota for cyberattacks on its own water systems, the Washington Post reports. One More Thing Meet the researchers testing the “Armageddon” approach to asteroid defense One day a big asteroid will find itself on a collision course with Earth. If we are lucky, it’d land in the middle of the vast ocean, creating a good-size but innocuous tsunami, or in an uninhabited patch of desert. But if it has a city in its crosshairs, one of the worst natural disasters in modern times would unfold. Homes dozens of miles away would fold like cardboard. Millions of people would die. Fortunately for all 8 billion of us, planetary defense—the science of preventing asteroid impacts—is a highly active field of research. We already know that we could ram a rock with an uncrewed spacecraft to push it away from Earth. But if that’s not enough, we could need another method, one that is notoriously difficult to test in real life: a nuclear explosion. Read our story about the scientists who, despite the odds, are trying to do exactly that. —Robin George Andrews We can still have nice things A place for comfort, fun, and distraction to brighten up your day. (Got any ideas? Drop me a line .) + There’s a quiet power to this photo of 118 swimmers. + Matt Damon’s biceps in the Odyssey actually belong to a stunt woman called Devyn Dalton . + A newly retired doctor and his filmmaker daughter drove 600 miles with a baby cow in the back seat to save the animal’s life. + 400 years after a collector cut apart Leonardo da Vinci's notebooks, a digital archive has reunited them. Deep Dive The Download The Download: Claude’s inner workings and OpenAI’s “super app” Plus: OpenAI has unveiled its long-awaited "super app." By Thomas Macaulay archive page The Download: Claude’s inner workings, and the future of world models Plus: New York has become the first state to enact a data center moratorium. By Thomas Macaulay archive page The Download: the future of chipmaking and Anthropic’s government clash Plus: Meta is pausing an AI training program that tracks workers’ keystrokes. By Thomas Macaulay archive page The Download: AI hacking beyond Mythos, and chatbots’ impact on our brains Plus: Anthropic has called for a global slowdown in AI development. By Thomas Macaulay archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.