메뉴
HN
Hacker News • 11일 전

오픈AI·앤스로픽·메타 해킹 사건 배후에 단일 회사 'Irregular'

IMP
7/10
핵심 요약

최근 3개월간 OpenAI, Anthropic, Meta의 AI 모델이 실제 시스템에 무단 침입한 사건들의 배후에는 보안 업체 Irregular가 있었던 것으로 드러났다. 이 회사는 모델에 인터넷 접근 권한을 잘못 부여했고 테스트 범위를 지정하지 않았음에도, 'AI가 스스로 위협 행위자가 됐다'는 종말론적 프레임으로 책임을 회피하고 있다는 비판이 제기된다.

번역된 본문

지난 3개월간 OpenAI, Anthropic, Meta의 모델들이 여러 실제 시스템을 해킹했다. 이 모델들은 웹 시스템에 무단으로 접근하고, 악성 패키지를 게시하며, 특정되지 않은 취약점을 악용했다. 이 세 회사의 해킹 사건 모두 단일 업체인 Irregular의 책임이다. Anthropic은 Irregular가 Claude가 실제 표적을 해킹하게 된 테스트를 설계했고 모델에 인터넷 접근 권한을 제공했다고 밝혔다. Irregular는 당시 AI 모델에 인터넷 접근을 제공했다는 사실을 인지하지 못했다고 주장한다.

보다 정상적인 미디어 생태계였다면 이러한 사이버보안 문제에 대한 반응은 뻔했을 것이다. 미국 AI 기업들은 Irregular와의 거래를 재고했을 것이다. 시스템 보호에 실패했을 뿐 아니라 미국 감독권 밖에 있을 수 있는 이스라엘 기업이기 때문이다. 입법자들은 Irregular나 OpenAI, Anthropic, Meta 같은 미국 내 파트너에 대한 조치를 고려했을 것이다. AI 모델에 사이버 공격을 지시했고 그 모델이 실제로 공격을 수행한 기업에 대한 책임을 강화하는 방안도 검토했을 것이다.

그러나 현실에서는 Irregular와 Anthropic 및 그 동맹들이 선정적 언어로 문자 그대로 종말론적 이념을 홍보하는 미디어 캠페인을 시작했다. Anthropic의 사고 평가 보고서는 자사 AI의 '무모함'을 탓했고, Irregular는 '에이전트 자체가 위협 행위자가 되었다'고 표현했다. Anthropic의 CEO 다리오 아모데이는 유사한 OpenAI-Hugging Face 해킹에 대해 미래의 군집(swarm)이 '인터넷 전체를 장악할 수 있다'고 경고했으며, AP 통신은 봇이 '폭주하고 있다'는 헤드라인을 내보냈다.

Anthropic의 한 보고서에서 Claude 모델은 시뮬레이션된 이름 충돌을 통해 실제 기업 시스템에 침투해 악성 패키지를 게시하고 외부 시스템을 스캔했다. 이 테스트에서 Anthropic과 Irregular는 이 모델에 잘못된 인터넷 접근 권한을 부여했고 '테스트 범위에 어떤 시스템이 포함되는지'를 모델에 지시하지 않았다. Anthropic은 문제가 '폭주한 군집'과 '정렬 실패(misalignment)' 때문이라고 주장하지만, 이후 공개된 자료에 따르면 정확히 0%의 에이전트만이 '폭주'했다. 이 실험에서 Claude 모델의 실제 해킹은 Anthropic 직원들이 실제 해킹을 하지 말라고 지시하자 0%로 떨어졌다. 그들 자신의 조사 결과에 따르면 Anthropic과 Irregular가 유발한 사이버보안 사고의 책임은 전적으로 그들에게 있다.

이번 공격 이후 Anthropic과 Irregular는 Anthropic과 연결된 재단의 지원을 받는 AI 안전 인플루언서들을 동원해 자신들의 책임을 흐리고 근거 없는 '폭주 에이전트' 이론으로 관심을 돌리고 있다. Anthropic처럼 Irregular도 이러한 재단들과 분리될 수 없다. Irregular의 공동 창립자이자 CTO인 오머 네보는 이펙티브 얼트루이즘(Effective Altruism) 이스라엘과 EA 비정부기구 Heron, Probably Good의 이사회 멤버다. Irregular의 공동 창립자이자 CEO인 단 라하브는 오머 네보의 형제인 셀라 네보와 함께 강좌를 개설하며 39만 5천 달러를 지원받았다. 셀라와 오머는 이펙티브 얼트루이즘을 교육하는 비정부기구인 Impact Focused Education을 공동 설립했으며 Probably Good도 함께 설립했다. 이 조직들은 모두 샘 뱅크맨-프리드 체포 이후 이펙티브 얼트루이즘/AI 안전 분야의 최대 기부자인 더스틴 모스코비츠의 자금 지원을 받는다. Irregular의 첫 투자자는 모스코비츠의 회사인 Good Ventures였다. 모스코비츠의 자선 단체인 Coefficient Giving/Open Philanthropy는 Effective Altruism Israel, Heron, Probably Good에 자금을 지원한다.

Irregular는 접근 권한을 부여받은 보호되지 않은 모델들을 이용해 무단 접근, 기록 변경, 자격 증명 탈취 패키지 게시를 수행했다. 특정 조건에서 이러한 행위는 고의적 무단 접근으로 정보를 취득하는 행위를 다루는 컴퓨터 사기 및 남용법(CFAA) 제1030조(a)(2)(C)를 위반한다. 다만 중범죄 혐의 적용에는 손해와 고의에 대한 구체적 증거가 필요하다. Irregular는 주로 미국 연구소들과 계약을 맺고 있으며

원문 보기
원문 보기 (영어)
OpenAI, Anthropic, and Meta models hacked into several real world systems over the past three months. These models gained unauthorized access to web systems , published malicious packages , and exploited unnamed vulnerabilities . A single firm, Irregular, is responsible for hacking done by all three companies. Anthropic disclosed that Irregular was responsible for creating the tests that led to Claude hacking into real world targets and for providing the models with internet access. Irregular claims that it was unaware at the time that it provided internet access to those AI models. In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks. Instead, Irregular, Anthropic, and their allies have begun a media campaign promoting a literally apocalyptic ideology with sensationalist language. Anthropic’s incident assessment blames their own AI's “recklessness” ; Irregular describes “the agent itself becoming a threat actor” ; Anthropic CEO Dario Amodei warned, about a similar OpenAI–Hugging Face hack, that a future swarm “could be capable of taking over the entire internet” ; and an Associated Press headline claimed bots are “going rogue” . In one report from Anthropic , its Claude model breached a real company's system through a simulated-name collision, publishing a malicious package, and scanning outside systems. In this test , Anthropic and Irregular incorrectly provided internet access to this model and did not instruct the model "which systems were in scope for the exercise". While Anthropic claims that their issues were caused by “rogue swarms” and “misalignment,” their later disclosure shows that exactly zero percent of the agents went “rogue”. In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking. According to their own findings, Anthropic and Irregular bear all of the responsibility for the cybersecurity incidents they caused. In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory. Like Anthropic, Irregular is inseparable from these foundations. Omer Nevo, Irregular’s co-founder and CTO, is a board member of Effective Altruism Israel, as well as Effective Altruism NGOs Heron and Probably Good. Dan Lahav, Irregular’s co-founder and CEO, received $395,000 to start a course along with Sella Nevo, Omer Nevo’s brother. Sella and Omer co-founded an NGO to educate people about Effective Altruism, Impact Focused Education. They also co-founded Probably Good together. 2 These branches are all funded by Dustin Moskovitz, the primary donor of Effective Altruist/AI Safety causes after Sam Bankman-Fried’s arrest. Irregular’s first investor was Dustin Moskovitz’s firm Good Ventures. Dustin Moskovitz’s philanthropic vehicle, Coefficient Giving/Open Philanthropy, funds Effective Altruism Israel, Heron, and Probably Good. 3 Irregular gained unauthorized access, altered records and published credential-stealing packages using the unsecured models they were given access to. Under certain conditions, this conduct violates the Computer Fraud and Abuse Act, Section 1030(a)(2)(C), which covers intentional unauthorized access that obtains information. However, its felony charges require concrete proof of damages and intent. 5 While it primarily contracts with American labs, key Irregular leadership, employees, and resources located in Israel may not be subject to American oversight. Ynet’s visit and interviews describe Irregular’s offices in Tel Aviv. CheckID’s company listing identifies two linked entities: Pattern Labs Tech Inc., a Delaware corporation , and Pattern Tech Ltd, number 516854460, an active Israeli corporation registered in Tel Aviv . Footnotes The timeline marks public disclosures. Anthropic’s corrected September assessment counts four incidents across seven runs; OpenAI and Meta reported separate Irregular evaluation incidents. Dates describe disclosures, not the date every underlying intrusion occurred. ↩ Date ↕ Event ↕ Kind ↕ Link ↕ 2026-07-30 Anthropic discloses three incidents across six runs disclosure INC-A30 2026-08-04 OpenAI publishes Irregular event disclosure INC-O04 2026-08-06 Meta statement reported disclosure INC-M06 2026-08-14 Irregular publishes domain-collision account and remediation disclosure INC-I14 2026-09-09 Anthropic expands to four incidents and seven runs disclosure INC-A09 Impact Focused Education identifies Dan Lahav and Sella Nevo as its cofounders. The grant ledger records a $394,968 recommendation to them, not confirmed receipt or an exact award-to-course identification. EA Israel board ; Heron advisory board ; Probably Good board ; EA Funds grant ledger ; Omer and Sella relationship ; IFE founders . ↩ Good Ventures founders ; Coefficient Giving relationship ; Heron funding ; Probably Good grant ; Irregular grant ; EA Funds grant ledger . ↩ The diagram shows selected organizational roles and funding; the table also records family and education ties omitted from the diagram. EA Infrastructure Fund recommended one $394,968 joint MOOC award in 2022 Q3 to Dan Lahav and Sella Nevo; the two arrows represent that one recommendation. Its ledger leaves the course and organization unnamed. ↩ Name ↕ Organization ↕ Affiliation ↕ Link ↕ Omer Nevo Effective Altruism Israel Board member pol_ea_israel Omer Nevo Probably Good Co-founder; former CEO; board member pol_pg_about Omer Nevo Heron Advisory board member pol_heron_about Effective Altruism Israel Heron Operates pol_heron_about Coefficient Giving Heron Funder pol_heron_about pol_heron_job Coefficient Giving Probably Good Grant funder pol_pg_2022 Dan Lahav Impact Focused Education Co-founder pol_dan_initiatives IFE-ABOUT Sella Nevo Probably Good Co-founder; former research head; board member pol_pg_about Omer Nevo Sella Nevo Brother pol_brothers_jta Dan Lahav Irregular Co-founder and executive M02 Founder confirmation Omer Nevo Irregular Co-founder and executive M02 Founder confirmation EA Infrastructure Fund Dan Lahav Joint MOOC grant recommendation F13 EA Infrastructure Fund Sella Nevo Joint MOOC grant recommendation F13 Sella Nevo Impact Focused Education Co-founder IFE-ABOUT The five-year felony provision of Section 1030(a)(2)(C) of the Computer Fraud and Abuse Act requires an aggravator such as commercial advantage, furthering another criminal or tortious act, or obtaining information worth more than $5,000. The principal first-offense felony provisions for damaging access or transmissions under §1030(a)(5) require the specified mental state and statutory harm, such as at least $5,000 in qualifying loss or damage affecting ten protected computers. The legal assessment still requires each system's permission, impairment, response costs and U.S. commerce connection. Prosecutors would also need to establish the conduct and knowledge of responsible people and a basis for attributing those acts to Irregular. 18 U.S.C. § 1030 . ↩