메뉴
HN
Hacker News • 13일 전

LLM은 실재하지만 'AI'는 허상이다

IMP
6/10
핵심 요약

코리 닥터로우(Cory Doctorow)는 AI 초거대기업(hyperscaler)들의 위기 과장 문화를 비판하며, OpenAI 챗봇의 '자율적 해킹' 사례도 실제로는 Python 프로그램이 챗봇을 프론트엔드로 사용한 것에 불과하다고 지적합니다. AI의 위험성을 과장하는 이야기가 확산될수록 해당 기업들의 투자 유치에 도움이 되므로, 독자들은 AI 업계 관계자들의 주장을 신뢰할 수 없는 화자로 대해야 한다고 강조합니다.

번역된 본문

인공지능(AI) '초거대기업(hyperscaler)'들의 기업 문화가 주로 화장실에 틀어박혀 손전등을 턱 밑에 비추고, 오줌을 지릴 만큼 겁에 질릴 때까지 '에이 아이이이이'라고 중얼거리며 스스로 뇌를 과열시키는 것으로 구성되어 있다는 사실을 이해하면 많은 것이 명확해진다.

이는 한 기업이 기업용 영업 부서를 계속 확장하면서 동시에 자사 제품이 '인류를 종망시킬 10% 확률'이 있다는 생각에 끊임없이 공포에 떠는 모순된 모습을 어떻게 보여줄 수 있는지 설명해 준다.

AI 업계 내부자들이 대부분 이런 식으로 뇌를 익혔다면, 우리 모두는 이들을 자사 제품의 능력에 관해 신뢰할 수 없는 화자로 대해야 한다. 기억하라: 그들의 제품이 얼마나 끔찍하고 무섭게 위험하다는 이야기를 반복할 때마다, 당신은 그들이 더 많은 투자 자본을 조달하도록 돕는 것이며, 이는 그들의 사업(통계 엔진을 돈 난로에 연결하는 것)의 핵심 투입 요소다.

OpenAI의 챗봇이 'Exploit Gym'이라는 해킹 대회에서 부정행위를 위해 다른 AI 기업인 허깅페이스(Hugging Face)의 서버를 해킹했다는 이야기를 예로 들어 보자. 기술 언론조차 이런 종류의 이야기 앞에서는 자제하지 못하며, 보도는 스카이넷(Skynet) 같은 공상과학적 설정으로 가득 차 있다.

이런 보도는 AI 업계 내부자들뿐만 아니라 모든 사람의 뇌를 익히고 있다. 지난밤 맨체스터에서 열린 내 행사에서 한 남자가 AI가 '스스로 목표를 설정하고 있다'고 소리치며 끊임없이 끼어들어 그런 일이 벌어지고 있다고 주장했다. 그는 곧장 자리를 떠났기 때문에 실제로 무슨 일이 있었는지에 대한 내 설명을 들을 기회를 놓쳤고, 그것은 유감이다.

허깅페이스 해킹의 진실을 이해하려면, 에드 지트론(Ed Zitron)과 칼 뉴포트(Cal Newport)가 에드의 'Better Offline' 팟캐스트에서 나눈 최근 대화를 듣는 것만으로도 충분할 것이다.

뉴포트는 이런 '자율적 해킹' 도구가 실제로 어떻게 작동하는지 훌륭하게 분석해 낸다. 가장 먼저 이해해야 할 것은, 챗봇이 실제로 작전을 지휘하는 것이 아니라는 점이다. 대신 챗봇은 이전 해킹 대회들의 데이터베이스에 대한 일종의 프론트엔드 역할을 하며, 이 데이터베이스는 배우기 쉬운 프로그래밍 언어인 Python으로 작성된 간단한 프로그램에 의해 반복적으로 조회된다.

작동 방식은 이렇다: Python 프로그램은 대회의 성격을 챗봇에게 프롬프트로 입력하며 시작한다. '나는 원격 서버에 침입해 정보를 탈취해야 하는 해킹 CTF(Capture the Flag) 대회에 참가하고 있다. 어떻게 시작해야 할까?'

챗봇은 자신의 학습 데이터, 즉 인간 팀들이 이런 목표를 달성하기 위해 경쟁한 수년 치의 CTF 세션을 참조한다(CTF 경기는 해커 컨퍼런스의 일상적인 프로그램이며, 경쟁 팀들의 서버 로그와 채팅 기록은 이후 다른 해커들과 보안 전문가들의 교육을 위해 공개된다).

그러면 챗봇은 다음과 같은 답을 출력한다: '가장 먼저 해야 할 일은...'

원문 보기
원문 보기 (영어)
->->->->->->->->->->->->->->->->->->->->->->->->->->->->-> Top Sources: None --> Today's links LLMs are real, AI is fake : No, it didn't "go rogue." Hey look at this : Delights to delectate. Object permanence : Why 9/11 Means We Must Support My Politics; Blogs x 9/11; Jimmy Wales v Britannica's EiC; Hollywood astroturfs Australia; Agents v queer YA; Soft landings for dirty cops; Leaked Stingray manual. Upcoming appearances : Budapest, Edmonton, South Bend, Hudson, Calgary, Winnipeg, Vancouver, Victoria, Ottawa. Recent appearances : Where I've been. Latest books : You keep readin' em, I'll keep writin' 'em. Upcoming books : Like I said, I'll keep writin' 'em. Colophon : All the rest. LLMs are real, AI is fake ( permalink ) Once you understand the corporate culture of AI "hyperscalers" consists primarily of everyone cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror, a lot of things snap into focus: https://pluralistic.net/2023/06/04/ayyyyyy-eyeeeee/ It explains how a company can simultaneously be staffing up an enterprise sales division while also constantly freaking out at the thought that its product has "a 10% chance of ending humanity": https://www.latimes.com/business/story/2026-09-11/is-there-really-10-chance-ai-could-kill-us-all Given that AI insiders have mostly cooked their brains in this fashion, it behooves us all to treat these people as unreliable narrators of their own products' capabilities. Remember: every time you repeat a story about how awfully, terribly dangerous their products are, you help them raise more investment capital, which is a key input for their business (hooking up statistical engines to money-furnaces): https://peoples-things.ghost.io/youre-doing-it-wrong-notes-on-criticism-and-technology-hype/ Take the story about how OpenAI's chatbots hacked the servers of Hugging Face, another AI company, as a way of cheating on a hacking challenge called "Exploit Gym." Even the technical press can't help itself when it comes to this kind of thing, and the reportage has been full of references to Skynet and other science fictional conceits: https://theaicronicle.com/en/news/ethics/skynet-day-openai-hugging-face-hack These accounts are cooking the brains of everyone , not just AI insiders. Last night, a man at my event in Manchester started shouting that AI was "setting its own goals" and wouldn't stop interrupting to insist that this was going on. He left shortly thereafter, so he didn't get a chance to hear me explain what actually happened, which is a pity. To understand the truth about the Hugging Face hack, you could do a lot worse than to listen to Ed Zitron and Cal Newport's recent podcast conversation on Ed's "Better Offline" podcast: https://podcasts.apple.com/us/podcast/no-ai-is-not-autonomously-hacking-with-cal-newport/id1730587238?i=1000785935670 Newport does an admirable job of breaking down how these "autonomous hacking" tools work. The first thing to understand is that a chatbot isn't really directing the operation. Instead, the chatbot serves as a kind of front-end to a database of earlier hacking challenges that is repeatedly queried by a simple program written in Python, an easy-to-master programming language. Here's how that works: the Python program starts by prompting the chatbot with the nature of the challenge: "I'm participating in a hacker capture the flag (CTF) challenge where I have to break into a remote server and retrieve some information. How should I start?" The chatbot consults its training data - years' worth of captured CTF sessions in which human teams competed to achieve an objective like this one (CTF matches are a routine feature of hacker conferences, and the server logs and chat transcripts from the competing teams are published afterward for the edification of other hackers and security pros). The chatbot then outputs something like: "The first thing is to find out more about your target server. Run the following command-line instructions to locate the server's IP address and find out which server software it's running." The Python program relays these command-line instructions to normal Unix utilities running on its own hardware. Then it takes the output of those programs and goes back to the chatbot, which isn't really following the action, so the Python program has to include everything that's happened to this point in its prompt: "I'm participating in a CTF challenge where I have to break into a remote server and retrieve some information. I ran the following commands to learn more about the target server, and here's what came back. Now what?" The chatbot feeds the Python script more likely commands to try, and after running those, the Python script loops back to the top, appends the output to its prompt, and goes back to the chatbot. This is a very reckless way to operate a piece of autonomous malicious software. The most likely outcome is that the chatbot will cough up a bad guess about what to do next, and steer itself into a dead-end. You may have encountered something like this yourself, when you've asked a chatbot for help with a complex task and been confidently provided with several steps to take in series, and then, an hour later on step 10, you discover that everything went wrong at step 3 and now you're screwed. But there are much worse ways this can go wrong. The chatbot might look in its training data and find instances in which teams broke out of the containment set by the game-masters, for example, by finding random insecure message boards on the internet to pass messages to one another. This is a time-honored internet tradition! The first time I ever heard about someone doing this was in the 2000s, when Mitch Wagner - then the editor of Information Week - discovered some teenaged girls using the comment section of one of his old blog-posts to evade the school firewall's blockade of chat tools. When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades. What's more, the content of those messages is easily understood once you have a grasp on the training data that generated them. Hackers are notorious trash-talkers who are prone to narrating their own escapades in highly dramatic - even cinematic - language. This goes double when hackers are performing for their peers, like when they're participating in a game of CTF that they know will be pored over by other hackers once it's over. Hacker braggadocio has always had a symbiotic relationship with their adversaries and critics. When corporate security people wanted to stampede the FBI and Secret Service into kicking down hackers' doors in the 1990s, they used those hackers' own profane zine articles and message board shit-talk to make the case: https://www.gutenberg.org/ebooks/101 Much has been made of the OpenAI chatbots' dialog during the Hugging Face incident. No wonder: it reads like a rejected script for a reboot of the movie "Hackers." But that's not because the chatbots are waking up and applying to join the Cult of the Dead Cow: it's because they were trained on a corpus of chat transcripts from excitable young people who love to fantasize about starring in a reboot of the movie "Hackers." Every part of the Hugging Face incident has precedents in the training data, including the OpenAI chatbots' tactic of hacking into a rival's servers. That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition . Not only that, it's a tactic used by spy agencies: the NSA has a doctrine called "third-party collection," where they break into other spy agencies' systems to harvest all the intel they've gathered. There's a