메뉴
BL
The Decoder 8일 전

구글, 젬미니 아키텍처를 직접 새긴 '프로즌 v2' AI 칩 개발 중

IMP
8/10
핵심 요약

구글이 젬미니(Gemini) AI 모델의 아키텍처를 실리콘에 직접 하드코딩하여 AI 추론 효율을 극대화하는 신규 서버 칩 '프로즌 v2(Frozen v2)'를 개발하고 있습니다. 현재 TPU 대비 6~10배 높은 효율을 낼 것으로 예상되며, 2028년 배포를 목표로 하고 있습니다. 이는 구글이 막대한 AI 연산 비용을 절감하고, 오픈AI 및 앤스로픽과의 경쟁에서 가격 우위를 점하기 위한 핵심 전략으로 평가됩니다.

번역된 본문

구글은 젬미니(Gemini) AI 모델의 아키텍처를 실리콘에 직접 구워 넣은(하드웨어에 통합한) '프로즌 v2(Frozen v2)'라는 새로운 서버 칩을 제작하고 있습니다. 더 인포메이션(The Information)이 인용한 소식통에 따르면, 이 칩은 구글의 현재 TPU 칩들에 비해 AI 응답을 제공하는 데 있어 6~10배 더 효율적일 수 있다고 합니다. 구글은 2028년부터 이 칩을 배포할 계획이며, TPU 라인보다는 적은 생산량으로 프로즌 v2를 특화된 칩의 일종의 시범 운용(test run)으로 삼을 예정입니다.

여러 모델과 함께 작동하는 구글의 기존 TPU와 달리, 프로즌 v2는 젬미니 모델 구조의 일부를 하드웨어 자체에 내장하고 있습니다. 이 이름은 AI 모델에서 파라미터를 '고정(freezing)'하여 값이 더 이상 변하지 않게 만드는 것과 같은 논리를 따릅니다. 프로즌 v2의 경우 모델의 일부가 칩 자체에 영구적으로 고정되어 연산 단계를 줄이고 응답 속도를 높입니다.

이 아이디어는 원래 구글 딥마인드의 수석 과학자인 제프 딘(Jeff Dean)에게서 나온 것으로 알려졌습니다. 그의 초기 '프로즌' 디자인은 모델 가중치(weights, AI 모델이 쿼리에 어떻게 반응하는지 결정하는 구체적인 설정 값) 자체를 칩에 직접 내장하는 것을 요구했습니다. 하지만 구글은 이 방식을 폐기했는데, 특정 버전의 젬미니에서만 작동하고 너무 빨리 구식이 될 것이기 때문이었습니다.

반면 프로즌 v2는 튜닝된 파라미터가 아닌 기반이 되는 청사진, 즉 모델의 아키텍처를 내장함으로써 더 유연한 접근 방식을 취합니다. 새로운 가중치는 여전히 칩에 로드할 수 있습니다. 더 인포메이션에 따르면 아키텍처의 어느 정도가 실제로 하드코딩될지는 아직 결정되지 않았습니다.

AI 추론에서 더 나은 마진 확보 이 칩은 구글이 동일한 모델 아키텍처를 고수하는 동안에만 작동하기 때문에 외부 고객을 위한 제품이 되지는 않을 가능성이 높습니다. 구글은 이미 메타(Meta)에 TPU를 임대하고, 외부 클라우드 고객에게 제공하며, 엔비디아(Nvidia) 매출의 10%를 확보한다는 내부 목표로 'TPU@Premises' 프로그램을 통해 엔비디아의 대안으로 TPU를 포지셔닝하고 있습니다. 이와 대조적으로 프로즌 v2는 구글 내부의 AI 연산 capacity 부족 문제를 완화하기 위한 목적입니다.

하지만 이 칩이 기대에 부응한다면, 여전히 주요 경쟁 우위가 될 수 있습니다. AI 비즈니스에서 기업들이 추론 비용을 얼마나 잘 최적화하느냐가 점점 더 마진을 결정하는 요소가 되고 있습니다. 구글은 프로즌 v2를 활용해 강력한 모델을 더 낮은 가격에 구동하고, 오픈AI(OpenAI)와 앤스로픽(Anthropic)으로부터 시장 점유율을 빼앗아 올 수 있습니다.

원문 보기
원문 보기 (영어)
Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 20, 2026 GPT-Image-2 prompted by THE DECODER Ask about this article… Search Google is building a new server chip internally called "Frozen v2" that embeds the Gemini AI model's architecture directly into silicon. The chip could be 6 to 10 times more efficient at serving AI responses than Google's current TPU chips, according to sources cited by The Information . Google plans to deploy it starting in 2028 and sees Frozen v2 as a test run for specialized chips, with a smaller production volume than its TPU line. Unlike Google's TPUs, which work with many models, Frozen v2 has parts of Gemini's model structure built right into the hardware. The name follows the same logic as "freezing" parameters in AI models, where you lock values so they stop changing. With Frozen v2, a portion of the model gets permanently frozen into the chip itself, which cuts down on compute steps and speeds up responses. Ad The original idea reportedly came from Jeff Dean, Google Deepmind's chief scientist. His first Frozen design called for embedding the model weights directly into the chip. Weights are the specific settings that determine how an AI model responds to queries. Google scrapped that approach because the chip would have only worked with a single Gemini version and would have become outdated too quickly. Ad DEC_D_Incontent-1 Frozen v2 takes a more flexible path by embedding the model architecture instead of weights, meaning the underlying blueprint rather than the tuned parameters. New weights can still be loaded onto the chip. How much of the architecture will actually be hardcoded hasn't been decided yet, according to The Information. Squeezing better margins out of AI inference Because the chip only works as long as Google sticks with the same model architecture, it probably won't become a product for outside customers. Google already leases its TPUs to Meta , offers them to external cloud customers , and positions them through its "TPU@Premises" program as an alternative to Nvidia with an internal goal of capturing ten percent of Nvidia's annual revenue . Frozen v2, by contrast, is meant to ease Google's internal crunch on AI compute capacity . Ad If the chip delivers on its promise, it could still become a major competitive edge. In the AI business, how well companies optimize inference costs increasingly determines their margins. Google could use Frozen v2 to run powerful models at lower prices and take market share from OpenAI and Anthropic. Ad DEC_D_Incontent-2 AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: The Information