메뉴
BL
TechCrunch AI • 8일 전

PrismML, 초소형 LLM으로 AI 사용 방식 바꾸려 한다

IMP
7/10
핵심 요약

칼텍(Caltech) 연구진이 창업한 PrismML은 'Bonsai 2 27B'를 공개하며 알리바바 Qwen3.8 27B를 5.9GB로 압축, 원본 대비 98%의 벤치마크 성능을 유지하는 추론 모델을 PC와 스마트폰에서 구동 가능하게 만들었다. 3진(ternary) 가중치 기법으로 모델 크기를 9~10배 줄인 이 기술은 온디바이스 AI의 무료·프라이빗 실행을 가능하게 할 수 있어 업계 판도를 바꿀 잠재력이 있다.

번역된 본문

AI 연구소 PrismML이 아직 여러분의 관심 대상이 아니라면, 주목할 만하다. 거액의 투자를 유치해서가 아니라(아직은 2,225만 달러의 시드 라운드에 불과하다), 참여한 기술진의 역량과 업계를 바꿀 잠재력을 지닌 기술 때문이다. PrismML은 유능하고 고성능인 추론 대규모 언어모델(LLM)이 반드시 '대규모'일 필요는 없다고 베팅하고 있다. PC와 스마트폰에 들어갈 수 있을 만큼 작은 추론 모델을 만들고 있는 것이다. (애플과 협상 중이라는 소문도 있으나, CEO 바박 하시비(Babak Hassibi)는 테크크런치에 대해 이를 언급하기를 거부했다.)

목요일 PrismML은 모델 패밀리의 최신작인 'Bonsai 2 27B'를 공개했다. 이 모델은 알리바바의 널리 쓰이는 오픈소스 모델인 Qwen3.8 27B를 5.9GB로 압축한 것으로, PC는 물론 고성능 스마트폰에도 저장 가능한 크기다. 원본 대비 메모리 사용량이 9~10배 줄어든 것이다.

PrismML은 칼텍(Caltech) 연구진들이 창업했으며, 압축 기술 전문가이자 칼텍 교수인 하시비가 이끌고 있다. 이 스타트업의 자문으로는 아이온 스토이카(Ion Stoica)가 있다. 스토이카는 Databricks(외 다수 기업)의 공동 창업자이자 버클리의 유명한 Sky Computing Lab 소장으로, 이 연구소는 Letta부터 SGLang까지 수많은 기술과 스타트업을 배출했다. PrismML은 코슬라 벤처스, 케르베로스 캐피탈, 칼텍의 투자를 받고 있다.

LLM 압축 기술을 개발하는 회사는 PrismML만이 아니다. 스페인 도노스티아 국제물리센터의 저명한 교수가 창업한 Multiverse Computing도 있다. (그리고 Multiverse Computing은 거액의 자금을 조달했다.) 그러나 하시비는 PrismML의 압축 기술이 원본 대비 성능 손실이 사실상 없다는 점에서 독특하다고 말한다. Bonsai 2는 Qwen의 종합 벤치마크 점수의 98%를 달성한다. 이는 3월에 공개된 초대 Bonsai가 달성했던 95%에서 향상된 수치다. 회사에 따르면 초대 모델은 이미 1,100만 회 이상 다운로드되었으며, 더 작은 모델들도 추가로 260만 회 다운로드되었다. 이는 PrismML의 압축 결과가 릴리스를 거듭하며 개선되고 있음을 보여준다. 100% 벤치마크 성능 동등성에 도달할 수 있을지는 아직 미지수다. 하시비는 압축이 항상 어느 정도의 영향을 미칠 것이라고 말한다.

그래도 완벽한 벤치마크 동등성은 다소 이론적인 문제다. LLM은 비압축 상태에서도 그리 정확하지 않고, 벤치마크도 실제 작업을 완벽히 반영하지 못하기 때문에 2%의 성능 저하가 실제 사용에서 의미 있는 차이를 만들 가능성은 낮다. (게다가 모델을 둘러싼 소프트웨어, 즉 모델이 실행되는 '하니스'도 정확도에 큰 영향을 미친다.)

PrismML은 모델을 구성하는 '가중치(weights)'를 축소하는 방식으로 이를 달성한다고 설명한다. 가중치는 본질적으로 모델이 학습 중에 습득·저장하는 정보다. 일반적으로 각 가중치는 16비트를 필요로 하지만, PrismML의 '3진(ternary)' 가중치 접근 방식은 이를 +1, −1, 0 세 가지 값으로 단순화한다. 가중치당 저장해야 할 값이 훨씬 작아지면서 모델 크기가 극적으로 줄어드는 것이다. (압축 기법에 대한 더 깊은 내용은 프로젝트의 GitHub 페이지를 참고하라.)

이 스타트업의 다음 목표는 이 압축 기법을 더 큰 모델에 적용하는 것이다. 하시비는 테크크런치에 "앞으로 몇 달 내에 공개할 다음 모델들은 수천억 파라미터급이 될 것이며, 그 수준에서는 지능을 유지하기가 더 쉬울 것으로 기대한다"고 말했다. 모델 크기가 커질수록 "지능을 잃지 않고 압축할 여지가 더 많아진다. 그래서 일반적인 추세로서, 더 큰 모델일수록 100%에 도달하기가 더 쉽다"고 덧붙였다.

스토이카는 이 기술이 고급 모델을 사용자의 기기에서 실행할 수 있게 만들기 때문에 흥미진진하다고 말한다. "지성(intelligence)이 손끝에 있게 될 것이고, 이미 구매한 기기에서 실행되기 때문에 무료일 것이며, 데이터를 보내지 않기 때문에 프라이빗할 것이다."

원문 보기
원문 보기 (영어)
If AI lab PrismML isn't on your radar yet, it should be — not because it's raised gobs of money (it hasn't yet, just a $22.25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it's developing. PrismML is betting that capable, high-performing, reasoning large language models don't, in fact, have to be large. It is making reasoning models so small they can fit on PCs and smartphones. (It's even rumored to be in talks with Apple , though CEO Babak Hassibi declined to comment on that to TechCrunch.) On Thursday, PrismML released Bonsai 2 27B , its latest in a family of models, which compresses Qwen3.8 27B, a widely used open-source model from Alibaba, down to 5.9 GB. That's small enough to fit on a PC and, possibly, a high-end smartphone. It's a 9x to 10x reduction in memory versus the original. PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an advisor. Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley's famed Sky Computing Lab, which has birthed many technologies and startups, from Letta to SGLang . PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech. This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain's Donostia International Physics Center, is another. (And Multiverse Computing has raised gobs of cash .) But Hassibi says that PrismML's compression tech is unique because its LLMs have lost virtually no performance compared with the originals. Bonsai 2 matches 98% of Qwen's aggregate benchmark scores. That's up from the first Bonsai, released a couple of months ago in March, that matched 95%. That original model has already been downloaded over 11 million times, and PrismML's even smaller models have been downloaded another 2.6 million times, the company says. So this shows that PrismML's compression results have improved from one release to the next. Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have some impact, Hassibi says. Still, perfect benchmark parity is fairly academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not so perfectly reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use. (Plus, the surrounding software — the harness a model runs inside of — matters a lot when it comes to accuracy , too.) PrismML says it achieves this by shrinking the "weights" that make up a model — weights are, essentially, the information a model learns and stores during training. Normally, each weight requires 16 bits. PrismML's approach, called "ternary" weights, simplifies that down to three: +1, −1, or 0. With far smaller values to store for each weight, the model takes up dramatically less space. (For a deeper dive on the compression technique, here's the project's GitHub page .) The startup's next goal is to apply this compression technique to even bigger models. "The next models that we will release, hopefully in the next couple of months, will be in the several-hundred-billion-parameter range, and I expect it will be easier to retain the intelligence there," Hassibi told TechCrunch. As model size grows, he added, "There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it's easier to get to 100%." Stoica tells us that he's excited for this tech because it's making it possible for advanced models to run on users' devices. "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud." Topics AI , Startups , TC When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Julie Bort Venture Editor Julie Bort is the Startups/Venture Desk editor for TechCrunch. You can contact or verify outreach from Julie by emailing julie.bort@techcrunch.com or via @Julie188 on X. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant Salesforce and Nvidia's new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis Tesla says it will finally unveil the second-generation Roadster on October 1 Anthony Ha Revolut confirms customer data breach through fake government requests Jagmeet Singh Matt Mullenweg tells (trolls?) Automattic staff, saying he's back in control after CEO ouster Sarah Perez