메뉴
BL
MIT Tech Review 1일 전

OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.

IMP
3/10
핵심 요약

[요약 오류] OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.

원문 보기
원문 보기 (영어)
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here . Reading OpenAI’s account last week of how some of its models broke their containment and hacked into the computer systems of Hugging Face , another AI company, was the first time I got genuine chills about what large language models are now able to do. But this is a case of human hubris, not rogue AI. I am not an alarmist. In fact, I have been pushing back against AI scare stories for years. Even so, this incident crossed a line. I think it’s the clearest illustration yet of how the people building and testing this technology do not fully understand what they’re doing. OpenAI could—and should—have seen this coming. Here’s what happened, at least according to the two companies involved. A couple of weeks ago, OpenAI started testing the hacking abilities of some of its new models, including GPT‑5.6 Sol (released in June) and what OpenAI describes as “an even more capable pre-release model.” OpenAI pitted its models against a benchmark called ExploitGym , released in May, which challenges LLMs to find ways to exploit real-world vulnerabilities found in commonly used software. To see what they could do, the researchers removed most of their cybersecurity guardrails. Then they ran the models inside a sandbox that was cut off from the internet except for one link to a third-party piece of software that acted as a proxy to the outside world, and let them install code that they needed to beat ExploitGym. On July 9, according to reporting by Reuters , OpenAI’s models started trying to break through the proxy. They found an unknown bug in the proxy’s software and used it to access the internet. From there, they broke into Hugging Face’s computer systems on July 11, apparently looking for data sets and solutions that would help them complete their task. Hugging Face announced the hack on July 16. OpenAI did not realize (or at least did not reveal) that its models were involved until July 21, around 10 days after they broke containment and a week after Hugging Face had shut down the attack and alerted the FBI. In a statement given to MIT Technology Review , OpenAI says: “We are conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we will publish a technical report of our learnings for everyone.” The firm also confirmed that its researchers were properly using existing safety guidelines and procedures at the time. Wake-up call OpenAI has said the event was unprecedented—and in many ways it was. This was the first time outside of a simulation that LLMs escaped what was thought to be a secure sandbox, accessed the open internet, and attacked an unrelated organization. It’s a wake-up call that shows just how good the latest LLMs are at finding and exploiting vulnerabilities in real-world software with little or no human guidance. And yet at the same time, what OpenAI’s models did is something this technology has done for years. Give a model a goal and it will very often achieve that goal in unexpected ways, finding loopholes that look like cheats. OpenAI itself has studied this behavior. A decade ago, it shared results of an experiment in which a model was tasked with beating a video game called CoastRunners . Human players take it for granted that the way to do this is by racing a boat through a series of flags to the finish line, racking up points for each flag you hit. OpenAI’s model figured out that you could get a high score by spinning in a circle and hitting the same three flags over and over again. There have been dozens of similar examples from researchers since. AI will always find a way. “Despite repeatedly catching on fire, crashing into other boats, and going the wrong way on the track, our agent manages to achieve a higher score using this strategy than is possible by completing the course in the normal way,” OpenAI wrote in a blog post about the CoastRunners experiment in 2016. “While harmless and amusing in the context of a video game, this kind of behavior points to a more general issue … it is often difficult or infeasible to capture exactly what we want an agent to do.” I couldn’t help thinking about CoastRunners when I read OpenAI’s blog post about the Hugging Face attack: “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal … After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.” Last week’s news was not about rogue AI, despite the headlines. It was about models achieving the goal they had been given: Find ways to exploit vulnerabilities in software. The fact that those models then behaved in a way OpenAI had not anticipated isn’t surprising. But it is worrying. Back in 2016, OpenAI had this to say about its CoastRunners bot: “More broadly it contravenes the basic engineering principle that systems should be reliable and predictable.” A decade on, those basic engineering principles are still AWOL. Deep Dive Artificial intelligence A startup claims it broke through a bottleneck that’s holding back LLMs Subquadratic has now shared more details about its new model. But some are still skeptical. By Will Douglas Heaven archive page Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. By Will Douglas Heaven archive page Claude Science is Anthropic’s newest flagship product The company is doubling down on AI for science. By Grace Huckins archive page Google DeepMind is worried about what happens when millions of agents start to interact The firm is calling for more scientists to study the risks of multi-agent systems. By Will Douglas Heaven archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.
관련 소식
TD
The Decoder 1일 전
IMP 8

OpenAI: 전 직종의 절반, ChatGPT로 타인 업무 대체

OpenAI의 분석에 따르면 직장인들이 자신의 전문 분야가 아닌 다른 직군의 업무를 처리하는 '태스크 크로스오버(Task Crossover)' 현상이 43.5%에 달하는 것으로 나타났습니다. 특히 전문 인력이 부족한 중소기업일수록 마케팅, 엔지니어링, 데이터 분석 등의 타 직무 영역을 AI로 대체하는 경향이 뚜렷했습니다. 이는 기존 직무의 경계가 허물어지고 실제 업무 형태가 빠르게 변화하고 있음을 보여주는 중요한 지표입니다.

오픈에이아이 챗지피티 업무 생산성
TD
The Decoder 1일 전
IMP 8

마이크로소프트, 자체 보안 AI 모델 발표... 여전히 복잡한 작업은 OpenAI에 의존

마이크로소프트가 비용 절감과 성능 향상을 위해 자체적인 컴팩트 보안 AI 모델인 MAI-Cyber-1-Flash를 출시했습니다. 이 모델은 대부분의 보안 작업을 처리하지만, 여전히 가장 복잡한 추론 작업에 대해서는 OpenAI 모델에 의존하는 하이브리드 방식을 취하고 있습니다. 또한, 실시간 위협 모니터링 시스템인 Perception을 도입하며 방대한 보안 데이터를 기반으로 한 AI 오케스트레이터로서의 입지를 강화하고 있습니다.

마이크로소프트 사이버보안 OpenAI
TD
The Decoder 1일 전
IMP 8

인도 법원, OpenAI 저작권 가처분 신청 기각… AI 학습 공정 이용 인정

인도 델리 고등법원이 대형 통신사 ANI가 OpenAI를 상대로 제기한 저작권 침해 가처분 신청을 기각했습니다. 법원은 챗GPT가 기사를 그대로 복제했다는 ANI의 주장에 증거가 부족하며, AI 모델 학습 과정이 인도 저작권법상 '연구'를 포함한 사적 이용 예외에 해당될 수 있다고 판단했습니다. 이는 전 세계 법원이 AI 학습 데이터 활용의 적법성을 공정 이용으로 인정한 첫 사례로, 향후 AI 저작권 소송에 큰 선례를 남길 중요한 판결입니다.

저작권 ai-학습 공정-이용