메뉴
HN
Hacker News • 50일 전

AMD, Taalas 인수... 모델을 실리콘에 각인해 추론 10배 끌어올린다

IMP
8/10
핵심 요약

AMD가 AI 칩 스타트업 Taalas를 인수하며, AI 모델의 가중치(weights)를 실리콘 자체에 직접 각인(etching)하여 추론 성능을 극대화하는 기술을 확보했습니다. 이 방식은 특정 모델에 최적화된 맞춤형 칩(MSIC)을 제작해 기존 GPU 대비 최대 48배 빠른 토큰 생성 속도를 보여주며, 코드 어시스턴트 등 고성능 AI 에이전트 추론 비용을 획기적으로 낮출 수 있습니다. 비록 모델이 변경될 경우 칩을 다시 설계해야 하는 치명적인 단점이 있지만, AMD는 자사 GPU와 결합하여 대규모 AI 워크로드를 처리하는 새로운 아키텍처를 구축할 계획입니다.

번역된 본문

AI 및 ML 분야: AMD가 모델 가중치를 실리콘에 각인하여 추론 성능을 높이기 위해 AI 칩 스타트업 Taalas를 인수했습니다. 초기 기술 데모에 따르면 모델 특화 주문형 반도체(MSIC)가 초당 최대 17,000개의 토큰을 쏟아낼 수 있는 것으로 나타났습니다.

토비아스 만 (Tobias Mann), 시스템 에디터 발행일: 2026년 8월 6일 목요일 // 21:05 UTC

AI 하드웨어 시장에서 엔비디아(Nvidia)의 지배력을 뒤집으려는 AMD의 최신 시도로, '젠의 가문(House of Zen)'이라 불리는 AMD가 AI 칩 회사인 Taalas를 인수했습니다. 이 회사는 모델 가중치를 실리콘에 직접 구워 넣는(baking) 방식을 통해 추론 성능을 기존 대비 10배 이상 끌어올릴 수 있다고 약속했습니다. 목요일 장 마감 후 발표된 이번 거래는 작년 12월 엔비디아가 Groq와 맺은 200억 달러 규모의 라이선스 계약과 맥락이 매우 비슷해 보입니다. 즉, 코드 어시스턴트와 같은 AI 에이전트를 위해 필수적인 고성능 '프리미엄' 추론 서비스를 더 빠르고 저렴하게 만드는 것입니다. AMD는 인수 조건을 공개하지 않았지만, 단순한 인재 영입(acquihire)이 아닌 실제 기업 인수로 파악됩니다.

2023년에 설립되어 토론토에 본사를 둔 Taalas의 추론 접근 방식은 기존의 GPU나 Groq의 LPU, Cerebras의 웨이퍼스케일 가속기를 뒷받침하는 데이터플로우 아키텍처와는 근본적으로 다릅니다.

[광고 자리] 모델 특화 주문형 반도체 [광고 자리]

이 스타트업의 칩은 모델 가중치를 저장하기 위해 HBM(고대역폭 메모리)에 의존하지 않고, 이를 실리콘에 직접 각인(etching)합니다. 어떤 의미에서 Taalas의 칩은 진정한 의미의 모델 특화 주문형 반도체(MSIC)라고 볼 수 있습니다. 더 중요한 점은 Taalas의 기술이 단순한 개념에 그치지 않는다는 것입니다. 지난 2월, 이 스타트업은 TSMC의 6nm 공정으로 제작된 첫 번째 테스트 칩인 'HC1'을 공개했습니다. 초기 벤치마크에서 이 칩은 메타의 Llama 3.1 8B 모델을 초당 16,960개의 토큰이라는 엄청난 속도로 서빙했습니다. 작년 2월에 발표되었을 때를 기준으로, 이는 엔비디아 GPU보다 48배 빠르고 Cerebras 가속기보다 8.5배 빠른 속도였습니다. 2024년 중반에 처음 등장해 오늘날 기준으로는 이미 구식이 되어버린 Llama 3.1 모델이긴 하지만, 이 레티클(reticle) 크기의 칩은 개념을 증명하기 위한 목적이었습니다.

Taalas는 자사 칩이 실제로 어떻게 작동하는지에 대해 극비를 유지하고 있지만, 이 프로세서가 두 가지 주요 영역으로 구성되어 있다는 것은 알려져 있습니다. 모델 가중치가 각인되는 '마스크-ROM 리콜 패브릭(Mask-ROM recall fabric)'과 KV 캐시 및 파인튜닝 어댑터가 저장되는 'SRAM 리콜 패브릭(SRAM recall fabric)'이 바로 그것입니다.

이번 여름에 출시될 예정인 2세대 HC2 칩의 경우, Taalas는 파라미터 수를 200억 개(20 billion)로 끌어올리는 것을 목표로 하고 있습니다. 그리 대단한 수치로 들리지 않을 수 있지만, 대형 모델을 위한 GPU처럼 파이프라인 병렬 처리(pipeline parallelism)를 사용하여 가중치를 여러 가속기에 분산시키면 됩니다. 칩당 200억 개의 파라미터를 처리할 수 있다면, 1조 개(1 trillion)의 파라미터를 가진 모델을 구동하는 데 단 50개의 가속기만 있으면 됩니다. 마침 AMD는 이를 여유롭게 수용할 수 있는 랙 규모의 컴퓨팅 플랫폼과 자체 시스템 디자인 팀을 보유하고 있습니다. 이는 동일한 모델을 서빙하기 위해 수십 개의 GPU와 최소 2,000개 이상의 Groq LPU가 필요한 엔비디아의 최신 LPX 시스템보다 공간과 전력 효율이 훨씬 뛰어난 수치입니다.

파악된 바에 따르면, AMD는 자사의 인스팅트(Instinct) 기반 헬리오스(Helios) 랙과 Taalas 기술을 기반으로 한 칩을 짝지일 계획입니다. 이는 컴퓨팅이 많이 요구되는 프롬프트 처리는 GPU에서 수행하고, 토큰 생성은 Taalas 기반 가속기로 오프로드하는 분리형 아키텍처를 의미합니다.

[광고 자리]

또한 AMD는 일종의 틱-톡(tick-tock) 개발 주기를 채택할 가능성도 있습니다. 즉, 고객이 처음에는 인스팅트 가속기에서 모델을 배포하고 검증한 뒤, 만족스러우면 Taalas 가속기로 전환하는 방식입니다. 현 시점에서는 추측만 할 수 있을 뿐이지만, AMD의 AI 부문 수석 부사장(SVP)인 밤시 보파나(Vamsi Boppana)는 준비된 성명서에서 이렇게 말했습니다. "AMD는 고객이 모든 AI 워크로드에 맞는 올바른 컴퓨팅 솔루션을 유연하게 배포할 수 있는 풀스택 AI 플랫폼을 구축하고 있습니다."

해당 모델을 정말 사랑해야만 합니다 이 기술은 매우 빠르지만, 아직 눈치채지 못했다면 상당히 큰 단점도 함께 가지고 있습니다. 칩이 일단 배포되고 나면 해당 모델에 갇히게(stuck) 됩니다. LoRA 어댑터와 같은 수준 이상의 변화가 생기면, 칩을 다시 설계하고 제조하는 리스핀(re-spin) 작업이 필요하게 됩니다.

원문 보기
원문 보기 (영어)
AI and ML AMD acquires AI chip startup Taalas to boost inference performance by etching models into silicon Early tech demos show model-specific integrated circuits churning out up to 17,000 tokens a second Tobias Mann Tobias Mann SYSTEMS EDITOR Published thu 6 Aug 2026 // 21:05 UTC In AMD’s latest bid to upset Nvidia's dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more. The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia’s $20 billion licensing deal with Groq last December: make high-performance “premium” inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn’t disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire. Founded in 2023 and based in Toronto, Taalas’ approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras' waferscale accelerators. REG AD A model-specific integrated circuit REG AD The startup’s chips don’t rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas’ chips are really model-specific integrated circuits or MSICs. Perhaps more importantly, Taalas’ tech isn’t just conceptual. In February, the startup revealed its first test chip fabbed on TSMC’s 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators. While Llama 3.1 is ancient by today’s standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept. Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored. For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism. At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that. That’s quite a bit more space and power efficient than Nvidia’s recently unveiled LPX systems, which would need a few dozen GPUs and at least 2,000 Groq LPUs to serve the same model. From what we understand, AMD intends to pair its Instinct-based Helios racks with chips based on Taalas’ tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators. REG AD It’s also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they’re satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here’s what AMD’s SVP of AI, Vamsi Boppana, had to say about it in a canned statement: “AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload." You better really love that model While the tech is blazing fast, if you hadn’t already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you’re stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming. Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas’ tech, AMD’s customers are going to have to be really sure about their choice of models, which will be easier for some than others. However, if the startup is to be believed, the situation isn’t quite as bad as it sounds. While new models will require a re-spin, it doesn’t require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming. MORE CONTEXT Elon pledges to give Nvidia a virtual monopoly over the stars AMD's AI eggs are in too few baskets, Wall Street worries China turns up the heat with open model blitz as US model makers panic A deep dive into Nvidia's Vera CPU and the Olympus cores that power it With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site The Next Platform in February, the company suggested that etching a model's weights into silicon is 100x less expensive than training a frontier model. AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn't be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators. REG AD The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to “think” for longer before responding. One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. If AMD’s Taalas buy can drive down the cost per token and boost output speeds by 10x or 20x, model devs may opt to extend the reasoning time even further. In any case, we may not have to wait long to see just how Taalas fits into AMD’s broader vision. Subject to regulatory approval, the deal is expected to close in the fourth quarter. ® gpu nvidia amd ai semiconductor systems cloud infrastructure month 2026