메뉴
BL
MIT Tech Review 40일 전

AI 스타트업, LLM 효율성 병목 현상 돌파 주장

IMP
8/10
핵심 요약

AI 스타트업 Subquadratic이 대형 언어 모델(LLM)의 연산량과 비용을 획기적으로 줄이는 새로운 아키텍처 'SubQ'를 공개했습니다. 기존 Transformer(트랜스포머) 구조의 비효율성을 극복하여 매우 빠른 처리 속도와 낮은 전력 소모를 자랑하며, 제3자 평가를 통해 그 성능이 입증되었습니다. 이는 AI 모델 구축 방식을 근본적으로 바꿀 수 있는 중요한 기술적 도약으로 평가됩니다.

번역된 본문

초록: 마이애미에 본사를 둔 AI 스타트업 Subquadratic(서브쿼드래틱)이 지난달 스텔스 모드(비공개 상태)를 벗어나며 엄청난 주장을 내놓았다. 이 회사는 지난 10년 동안 대형 언어 모델(LLM)의 발전을 가로막아온 수학적 병목 현상을 해결했다고 발표했다. 구체적인 내용이 부족해 많은 사람들이 반신반의했지만, Subquadratic은 최근 자사의 새로운 기술에 대한 독립적인 평가 결과를 공유하며 이를 뒷받침할 증거를 제시하기 시작했다. 그 결과, 업계가 이 회사의 주장에 주목할 만한 가치가 있다는 것을 보여주고 있다.

Subquadratic에 따르면, 이 회사는 시중의 어떤 모델보다 속도가 빠르고 비용이 저렴하며 에너지 소모가 훨씬 적은 'SubQ'라는 새로운 종류의 LLM을 개발했다. 회사 측은 또한 SubQ가 대부분의 다른 모델들보다 한 번에 최대 12배나 많은 텍스트를 처리할 수 있어 수백 개의 문서나 전체 코드베이스를 분석하는 것과 같은 다양한 데이터 집약적인 작업을 수행할 수 있다고 주장한다. 더 나아가 Subquadratic은 SubQ가 코딩과 같은 핵심 작업에서 구글 딥마인드, 오픈AI, 앤스로픽이 출시한 최고 수준의 모델들과 거의 맞먹는 성능을 낸다고 밝혔다.

문제는 처음에 이 회사가 자체적으로 게시한 몇 가지 테스트 점수 외에는 이러한 주장을 뒷받침할 증거를 거의 제공하지 않았다는 점이다. 그리고 사람들이 직접 사용해 볼 수 있도록 SubQ를 널리 공개하지도 않았다. 따라서 Subquadratic의 주장이 회의적인 시각을 받은 것은 당연한 일이었다. AI 엔지니어인 댄 맥아티어(Dan McAteer)는 X(옛 트위터)에서 전반적인 반응을 이렇게 요약했다. "SubQ는 트랜스포머 이후 가장 큰 돌파구이거나... 아니면 AI 업계의 테라노스(Theranos, 과장으로 몰락한 바이오테크 기업)가 될 것입니다."

한 달이 지난 지금, 이 회사는 서드파티 기업인 Appen(아펜)이 실시한 추가적인 독립 테스트 결과를 포함하여 모델에 대한 더 많은 정보를 공개했다. Subquadratic의 공동 창립자이자 최고 기술 책임자(CTO)인 알렉스 웨든(Alex Whedon)은 "건전한 회의론이 있을 것이라 예상했다"며, "뒤돌아보건대 초기 발표와 함께 서드파티 벤치마크 결과를 함께 공유했다면 이런 회의론의 상당 부분을 방지할 수 있었을 것입니다. 그렇기에 우리는 앞으로 어떤 결과를 발표하기 전에 충분한 시간을 들여 철저히 검증할 것입니다."라고 말했다.

Subquadratic은 타사의 모델을 평가하는 역할을 맡고 있는 Appen에 SubQ 테스트를 의뢰했다. 그 결과는 Subquadratic의 주장을 상당 부분 뒷받침하는 것으로 보인다. Appen의 제너레이티브 AI 연구 총괄 디렉터인 재닌 시난난-싱(Jeanine Sinanan-Singh)은 "저에게는 정말 흥미로운 일이었고 그들의 아키텍처가 입증된 셈이었다"며, "모델들이 속도와 비효율성으로 고통받고 있었기에 '와, 이것은 판도를 바꿀 수 있겠다'고 생각했습니다. 하지만 스스로 놀라운 결과를 말할 때는 그다지 신뢰성이 떨어지기 마련이죠."라고 덧붙였다.

SubQ가 기존의 모든 최고 수준 모델을 전면적으로 대체하지는 않겠지만, 특정 작업의 경우 일반적인 비용의 극히 일부만으로도 엄청난 속도 향상을 제공할 수 있다. 그럼에도 불구하고 Subquadratic은 장기적으로 자신들의 혁신이 LLM이 구축되는 방식을 바꿀 수 있다고 주장한다. 회사의 공동 창립자이자 최고 경영자(CEO)인 저스틴 단젤(Justin Dangel)는 "우리는 새로운 효율성의 시대를 열고 있다고 생각합니다."라며, "몇 년 안에는 아무도 트랜스포머를 기반으로 모델을 만들지 않을 것이라 생각합니다."라고 말했다.

'어텐션(Attention)!' Subquadratic의 주장이 왜 중요한지 이해하려면, 대부분의 LLM이 어떻게 작동하는지 파고들어야 한다. LLM 내부의 핵심 메커니즘은 '밀집 어텐션(dense attention)'이라는 과정을 실행하는 '트랜스포머(transformer)'라는 신경망이다. 오늘날의 LLM은 일반적으로 여러 개의 트랜스포머를 함께 연결한다. (2017년 구글 연구원들이 발표한 LLM 시대의 기념비적인 논문 제목은 'Attention Is All You Need', 즉 '당신에게 필요한 것은 어텐션뿐'이었다.)

밀집 어텐션은 이렇게 작동한다. 트랜스포머가 텍스트 덩어리를 처리할 때, 먼저 각 단어(또는 토큰(token)이라고 하는 단어의 일부)를 숫자로 인코딩한다. 전체 텍스트의 의미를 파악하기 위해, 그런 다음 해당 텍스트의 모든 다른 숫자와 각 숫자를 곱한다. 예를 들어, 10,000단어 길이의 텍스트는 약 5,000만 번의 개별 곱셈 연산을 유발한다. 이는 엄청난 양의 연산이며, LLM이 희대의 전력 소모광으로 악명이 높은 주된 이유이다. "『위대한 개츠비(The Great Gatsby)』를 요약하려면 첫 번째 단어와 마지막 단어를 함께 살펴야 하고, 그 다음에는 (중간 단어들도 모두 연관 지어) 살펴야 한다..."

원문 보기
원문 보기 (영어)
EXECUTIVE SUMMARY Miami-based AI startup Subquadratic came out of stealth mode last month with a huge claim. It announced that it had solved a mathematical bottleneck that had been holding back large language models for almost a decade. The details were thin, and many people were unconvinced. But Subquadratic has started to bring the receipts, sharing the results of an independent evaluation of its new tech. The results suggest that the company’s claims might be worth paying attention to. According to Subquadratic, it has developed a new kind of LLM, called SubQ, that is faster and cheaper and uses a lot less energy than any other model on the market. The company also claims that SubQ is able to process up to 12 times as much text at once than most other models, allowing it to carry out a range of data-heavy tasks, such as analyzing hundreds of documents or entire code bases. What’s more, Subquadratic says, SubQ does this while more or less matching the performance of the best models put out by Google DeepMind, OpenAI, and Anthropic on key tasks like coding. The problem was that the company at first provided little evidence for its claims beyond a handful of self-published test scores. And it has yet to make SubQ widely available for people to try out themselves. So it’s no surprise that Subquadratic’s claims were met with skepticism. Dan McAteer, an artificial intelligence engineer, captured the overall response on X : “SubQ is either the biggest breakthrough since the Transformer ... or it’s AI Theranos.” A month on, the company has published more information about its model , including the results of additional independent tests run by third-party firm Appen. “We expected healthy skepticism,” says Subquadratic cofounder and chief technology officer Alex Whedon. “In hindsight, releasing the third-party benchmarks alongside the initial announcement would have preempted much of the skepticism, which is why we’re taking the time to make sure any future results are fully verified before putting them out.” Subquadratic asked Appen, which evaluates other companies’ models, to run its tests on SubQ. The results seem to back up a lot of Subquadratic’s claims. “That was really exciting to me, it validated their architecture,” says Jeanine Sinanan-Singh, Appen’s director of generative AI research. “I was like, ‘Wow, this could be a game changer,’ because models struggle with speed and inefficiency,” she adds. “But when you have kind of shocking results, it’s really not as credible when you say it yourself.” SubQ won’t replace existing top models across the board, but it could offer huge increases in speed at a fraction of the typical cost for certain tasks. Subquadratic insists that in the long run, though, its breakthrough could change how LLMs are built. “We hope we’re kicking off a new age of efficiency,” says Justin Dangel, the firm’s cofounder and CEO. “We don’t think anybody will be building on transformers in a few years.” Attention! To understand why Subquadratic’s claims are a big deal, let’s dig into how most LLMs work. The key mechanism inside an LLM is a type of neural network called a transformer, which runs a process known as dense attention. Today’s LLMs typically chain together multiple transformers. (The foundational paper of the LLM era, published by researchers at Google in 2017, was titled “Attention Is All You Need.” ) Dense attention works like this: When a transformer processes a chunk of text , it first encodes each word (or part of a word, known as a token) with a number. To capture the meaning of the full text, it then multiplies each of those numbers with every other number for that text. For example, a piece of text 10,000 words long would kick off almost 50 million individual multiplications. That’s a lot of computation and the main reason that LLMs are notorious power hogs. “If you want to summarize The Great Gatsby , you have to look at the first word and the last word together, and then you have to look at every other combination,” says Dangel. As the length of the text increases, the number of computations skyrockets. That’s because each additional number must be multiplied by all other previous numbers. Double the number of words, and you roughly quadruple the number of computations, a rate of increase known as a quadratic expansion. (You can picture this yourself: Draw a circle and mark dots around its edge. Each dot is a token. Then draw lines between pairs of dots to represent the multiplication of those two tokens. A circle with five dots will have 10 lines crossing it. Make it 10 dots and you will have 45 lines, 20 dots and you will have 190 lines, and so on.) Slashing costs Subquadratic’s solution is to ditch dense attention, the core operation of a transformer, in favor of what’s known as sparse attention, which slashes the number of computations needed. Instead of multiplying the number assigned to each token by every other number, sparse attention selects just some of the numbers to multiply. The idea is that not all relationships between words in a piece of text matter. “Sparse attention says not all of those relationships are important, because they’re not,” says Whedon. “If you’re reading a book, you’re not going to look at the first and second words, first and third—that’s insane.” It’s a simple approach, and Subquadratic is not the first to try it. “Pretty much everything under the sun has been attempted,” says Will Depue, an independent AI researcher who previously worked at OpenAI. “It’s not impossible, but it’s akin to running a four-minute mile.” Previous techniques for selecting which numbers to multiply and which to ignore have not produced a mechanism that can capture the meaning of a document as well as dense attention can. Subquadratic claims to have cracked the problem at last. It pitches SubQ as the first sparse-attention LLM that rivals mainstream dense-attention models in performance. “Historically, most mechanisms have used fixed patterns, like always comparing the first word to the fifth,” says Whedon. “That’s pretty limiting. Language is too sophisticated for that. And so, one of the things that makes our mechanism unique is that we dynamically select which ones are important.” The firm won’t say exactly how SubQ chooses which words to focus on, but the selection is calculated on the fly and differs for each piece of text the model is given. “That’s kind of where the secret sauce is,” says Whedon. Testing, testing The upshot is that for certain tasks, SubQ may be faster and cheaper to run than most other models. Appen evaluated SubQ on a handful of standard tests. In a straight-up speed test, which sets a baseline for how fast a model can operate in theory rather than assess what a model can actually do, Appen found that SubQ was 56 times faster than models using FlashAttention, a previous sparse-attention technique. On LiveCodeBench, a test that looks at how well models perform on competitive coding problems taken from real contests, SubQ scored 89.7%, putting it in the same ballpark as other top coding models . “This model continues to provide frontier-level performance in coding,” says Appen’s Sinanan-Singh. Subquadratic's claims about cost are harder to verify because SubQ is not yet widely available. According to Dangel, it costs $2600 to run Anthropic's LLM Opus 4.6 through RULER 128, a test developed by Nvidia to assess a model's ability to retrieve information from large data sets. And SubQ? "It cost us eight dollars," he says. SubQ does seem to be able to handle very large data sets. The model has a context window (roughly akin to a working memory) up to 12 million tokens long. Most top models today have context windows one million tokens long. In a demo that Whedon ran for me, he asked SubQ to perform a task that required it to reason about information contained in 400 documents. It responded in seconds. When he gave Perplexity—a popular LLM-powered search engine—the same task, it fa