메뉴
BL
MIT Tech Review • 15일 전

AI 전력 문제는 결국 아키텍처 문제다

IMP
7/10
핵심 요약

버지니아 애쉬번 데이터센터 클러스터의 대규모 정전 사례를 통해, AI 데이터센터의 급격한 부하 변동이 발전량 부족이 아닌 전력 아키텍처의 구조적 결함 때문임을 지적한다. 기존 UPS 중심의 저전압 전력 스택을 중전압·변전소 인근 모듈형·상시 경로 방식으로 전환해야 그리드 신뢰성을 지킬 수 있다고 제안한다.

번역된 본문

후원 제공: ON.energy

2026년 7월 22일, 세계 최대 데이터센터 클러스터의 중심지인 버지니아주 애쉬번에서 송전선로 고장이 발생해 몇 초 만에 3기가와트 이상의 부하가 그리드에서 이탈했다. 그리고 이는 처음 있는 일이 아니었다. 2년 전에는 서지 어레스터(피뢰기) 하나의 고장으로 약 60개의 버지니아 시설과 1,500메가와트가 한꺼번에 연결이 끊겼다. 이렇게 동일한 부하가 그리드 고장에 같은 방식으로, 같은 시점에 반응하리라고는 아무도 예상하지 못했다.

AI 전력 논쟁은 대부분 발전에 관한 것이다. 더 많은 터빈, 더 많은 태양광, 더 많은 송전선. 그리드에는 더 많은 전자(electron)가 필요하다는 것이다. 하지만 버지니아의 정전은 공급 실패가 아니라 아키텍처 실패였다. 그리고 거대한 규모의 계통연동(인터커넥션) 물결이 바로 그 아키텍처 위로 밀려오고 있어, 그리드 신뢰성이 위험에 처해 있다. 아무도 책임지려 하지 않는 문제다.

그리드에 더 많은 것을 요구하는 시대

그리드는 예측 가능한 부하를 중심으로 설계되었다. 제철소, 정유시설, 저녁 식사 시간의 가정들. 부하 크기는 달라도 과정은 같았다. 전력을 부드럽게 끌어다 쓰고, 가끔 문제를 일으키며, 우아하게 회복하는 것이다. 하지만 AI 데이터센터는 그렇게 움직이지 않는다.

AI 캠퍼스 하나는 학습 실행 중 수 밀리초 만에 부하의 70%를 오갈 수 있고, 상류에 문제 징후가 보이면 수십억 달러의 컴퓨팅 자산을 보호하기 위해 그만큼 빠르게 오프라인으로 이탈한다. 개별적으로는 합리적인 행동이다. 하지만 기가와트 규모에서 동시에 일어나면 그리드가 한 번도 해결한 적 없는 문제가 된다. 그리고 다음 물결의 데이터센터 캠퍼스는 정확히 그 규모로 계획되어 있다.

기존 스택이 무너지는 지점

표준적인 데이터센터 전력 스택은 수십 년간 변하지 않았다. 중전압 전력이 들어오고, 변압기가 전압을 낮추며, 저전압 무정전전원장치(UPS)가 전력을 정제해 랙에 전달한다. 이 설계를 AI 규모로 밀어붙이면 세 곳에서 균열이 생긴다.

첫째, UPS는 건물 깊숙이, 랙 가까이에 있다. 하지만 그 배터리는 과소한 스페어 타이어다. 몇 분간의 정전을 버티도록 설계되었을 뿐, 이토록 빠르고 변동성 큰 부하 변화를 연중무휴로 흡수하도록 만들어지지 않았다.

둘째, UPS는 평생의 대부분을 바이패스 모드로 보낸다. 레거시 컨버터의 손실이 커서 운영자들은 에코 모드로 운영한다. 정적 스위치가 그리드에서 랙으로 직접 전력을 공급하며 어느 방향으로도 필터링이 없다. 컴퓨팅의 부하 변동이 가공되지 않은 채 밖으로 나가고, 그리드의 과도 현상—장비를 손상시키거나 다운시킬 수 있는 1밀리초 미만의 사건들—은 스위치가 잡아낼 수 없을 만큼 빠르게 안으로 들어온다.

셋째, 보호 로직은 '대형 부하'가 50메가와트를 의미하던 시절에 작성되었다. 이 보호 로직은 자신이 이제 속한 그리드를 볼 수 없어서, 상류에 문제가 생기면 정확히 잘못된 행동, 즉 이탈을 한다. 2024년 버지니아 사고에서 유실된 부하 대부분은 전압 강하를 세다가 세 번째에 연결을 끊는 보호 방식—설계된 대로, 최악의 순간에—때문이었다. 이는 엉성한 엔지니어링이 아니다. 부하가 성장을 넘어선, 정교한 엔지니어링이다.

경로 안으로 이동하기

해법은 세 가지 이동을 동시에 하는 것이다.

위로 올린다 — 480볼트에서 중전압(13.8킬로볼트 이상), 즉 대형 사이트가 그리드에서 끌어오는 전압으로.

밖으로 이동시킨다 — 데이터 홀에서 변전소 인근의 모듈형 인클로저로 옮겨, 건물에는 컴퓨팅과 이를 지탱하는 냉각만 남긴다.

경로 안으로 넣는다 — 지켜보다 반응하는 배터리 대신, 모든 전자가 항상 통과하는 시스템으로. 감지할 것도, 전환할 것도 없다. 애초에 우회 경로가 없기 때문이다.

서류상으로는 세 가지 간단한 업그레이드다. 실제로는 하류의 모든 항목을 다시 쓰는 일이다.

변화 만들기

수천 개의 GPU가 동시에 구동될 때, 이 시스템은 부하 변동을 흡수하고 그리드에는 평평한 부하 프로파일을 넘긴다. 장애가 닥치면, 그 뒤의 장비는 아무것도 눈치채지 못한다. 성가운 이웃이 예측 가능한 이웃이 된다. 그리고 전력회사가 도움이 필요할 때는 유용한 이웃이 된다.

계통연동도 달라진다. 전력회사는 뒤따르는 모든 변압기, UPS, 냉동기, 펌프, 개폐기를 풀어헤치는 대신 중전압 박스 하나만 인증하면 된다. 엔지니어들은 새로운 계통연동 연구 없이도 칩 세대를 교체할 수 있다. 수개월이 (줄어든다)

원문 보기
원문 보기 (영어)
Sponsored Provided by ON.energy On July 22, 2026, a transmission line fault in Ashburn, Virginia—the heart of the world's largest data center cluster—knocked more than 3 gigawatts of load off the grid in seconds. And it wasn't the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time. The AI power debate is mostly about generation: more turbines, more solar, more transmission. The grid needs more electrons. But the outages in Virginia weren't supply failures; they were architecture failures. And a giant wave of interconnections is arriving on that same architecture, putting grid reliability at risk. It's a problem nobody wants to own. Asking more from the grid The grid was built around predictable loads: steel mills, refineries, and houses at dinnertime. Different load sizes, same process—drawing power smoothly, misbehaving occasionally, and recovering gracefully. But AI data centers don't behave that way. An AI campus can swing 70% of its load in milliseconds during a training run, then trip offline just as fast at the first sign of trouble upstream to protect billions in compute. Each is rational alone. Together, at gigawatt scale, they're a problem the grid has never solved—and the next wave of data center campuses is planned at exactly that scale. Where the old stack breaks The standard data center power stack hasn't changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage uninterruptible power supply (UPS) units condition it, and it reaches the racks. Push that design to AI scale, and it cracks in three places. First, the UPS sits deep inside the building, close to the racks. But its batteries are an undersized spare tire, designed to handle an outage for a few minutes, not to absorb load swings this fast and volatile around the clock. Second, the UPS spends most of its life in bypass. Legacy converters waste enough power that operators run in eco-mode: A static switch feeds the racks directly from the grid and nothing filters in either direction. The compute's swings go out raw, and grid transients—sub-millisecond events that can damage or take down equipment—come in too fast for any switch to catch. Third, the protection logic was written when "large load" meant 50 megawatts. This protection logic can't see the grid it is now a part of, so when trouble hits upstream, it does exactly the wrong thing: it drops out. In the 2024 Virginia event, most of the lost load traced to protection schemes that count voltage dips and disconnect on the third one —as designed, at the worst moment. This isn't sloppy engineering. It's careful engineering the load has outgrown. Moving into the path The fix is three moves, made together. Move it up —from 480 volts to medium voltage (13.8 kilovolts and higher), the voltage large sites draw from the grid. Move it out —from the data hall to modular enclosures near the substation so the building holds only compute and the cooling that keeps it alive. Move it into the path —instead of a battery that watches and reacts, a system every electron runs through, all the time. There's nothing to detect and nothing to switch because nothing was ever routed around it. On paper, three straightforward upgrades. In practice, they rewrite every line item downstream. Making the change When thousands of GPUs spin up together, the system absorbs the swing and hands the grid a flat load profile. When a disturbance hits, the equipment behind it never notices. A difficult neighbor becomes a predictable one. And when the utility needs help, it becomes a useful one. Interconnection changes, too. The utility certifies one medium-voltage box instead of untangling every transformer, UPS, chiller, pump, and switchgear lineup behind it. Engineers swap chip generations without a fresh interconnection study. Months come off the permitting timeline. Inside the fence, UPS rooms become compute or cooling space. Density per construction dollar climbs. And the economics flip. Equipment that runs at medium voltage, sits outside, and stores its own energy can qualify for tax credits, and earn revenue in grid programs like peak shaving and demand response. Backup power stops being insurance and starts paying for itself. The architecture test In early 2026, we tested a full-scale system at the National Laboratory of the Rockies, a U.S. Department of Energy facility and the only place in the Western Hemisphere that can replicate real grid faults and AI-scale load swings concurrently in the same loop. We hit it from both directions: real AI load profiles hit the compute side at full medium voltage. Grid faults hit the utility side, including a full zero-voltage event. The compute side didn't flinch. Neither did the grid side. It cleared the large-load voltage ride-through requirements from the Electric Reliability Council of Texas (ERCOT), the grid operator, with room to spare. Those rules exist because operators no longer take facilities this size on faith, and more are coming. Most of the industry treats them as hurdles. A medium-voltage, inline system clears them out of the box. Compliance isn't an added feature. It's what the architecture does. The new layer Much of what looks like a grid problem in the AI buildout sits inside the fence, in equipment sized for a load that no longer exists. Move the right pieces up, out, and into the path, and a grid liability becomes a grid asset. Density goes up. Permitting time comes down. Backup power earns its keep. The engineering works—and the next wave of AI factories is being built on it. The industry hasn't named this layer yet. We call it the medium-voltage AI UPS. The name matters less than the choice: those factories can arrive as a strain on the grid or as strength for it. We already know how to build the second kind. This content was produced by ON.energy. It was not written by MIT Technology Review’s editorial staff. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven archive page AI is more likely than humans to form biases when hiring AI doesn’t just learn stereotypes from its training. It can cook up new ones, too. By Michelle Kim archive page Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. By Grace Huckins archive page Bill Gates says we’ve passed AI’s danger thresholds. Now what? In a new interview, the billionaire philanthropist sounds an alarm on the urgency of getting our AI policies in order. By Mat Honan archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.