메뉴
BL
Wired AI • 23일 전

러시아 수학자들이 '말 없는 AI 통신' 기술 개발

IMP
7/10
핵심 요약

러시아 스타트업 모스틱(Mostik)이 AI 모델들이 텍스트를 생성하지 않고 가중치의 수학적 값만으로 서로 소통하는 기술을 개발했다. 이를 통해 대형 모델의 능력을 소형 모델에 효율적으로 전달할 수 있으며, 중국 오픈 소스 모델 GLM과 Qwen을 결합해 비용 20분의 1로 중간 수준 성능을 달성했다. 이 기술이 확산되면 오픈 소스 모델의 가치가 높아져 폐쇄형 프런티어 모델과의 경쟁력이 커질 전망이다.

번역된 본문

최근 나는 인공지능 모델들이 일종의 '기계 텔레파시'와 비슷한 방식으로 통신할 수 있게 해주는 방법을 보여준 뛰어난 러시아 수학자들을 만났다. 이 수학자들은 '모스틱(Mostik)'이라는 스타트업에서 일하는데, 이는 러시아어로 '다리'를 뜻한다. 이 이름은 서로 다른 모델이 가중치(weights)에 담긴 수학적 값 — 프롬프트가 출력으로 변환되는 방식을 결정하는 요소 — 을 이용해 상호작용할 수 있게 하는 이 팀의 접근 방식을 상징한다.

실제로 이는 대형 모델의 능력을 소형 모델에 전달해 훨씬 효율적으로 지능을 끌어올릴 수 있음을 의미한다. 이 스타트업은 이 방식으로 악명 높은 AI 모델 대회인 ARC-AGI 3 상위권에 오른 모델을 구축했다. (대회에서 우승하고 싶어하기 때문에 더 자세한 내용은 알려주지 않았다.)

하지만 개념을 시연하기 위해 이들은 두 개의 중국 오픈 웨이트(open-weight) 모델 사이에 '다리'를 놓기도 했다. 7,530억 개 파라미터의 최대 버전 GLM-5.2와, 모바일 기기에서 실행 가능한 40억 파라미터 버전의 Qwen-3.5다. 그 결과 탄생한 하이브리드 시스템은 전체 GLM 모델 비용의 20분의 1이 들었고, 성능은 두 모델의 정확히 중간 지점이었다.

"모델의 앙상블이 개별 모델보다 더 나은 성능을 낸다는 것은 머신러닝 분야에서 잘 알려져 있습니다." 모스틱의 CEO인 사샤 말리셰바(Sasha Malysheva)가 커피를 마시며 나에게 말했다. 이 접근법을 개발한 말리셰바는 회사 내에서 유행하는 농담을 공유했다. AI의 미래는 돼지의 무게를 맞히는 것과 비슷하다는 것이다. 수학계에서는 여러 명의 무작위 사람들의 추측을 종합하고 평균내면 전문가 한 명보다 돼지의 무게를 더 정확하게 추정할 수 있다는 것이 잘 알려져 있다. 공동으로 돼지의 무게를 가늠하는 것과 마찬가지로, 여러 AI 모델의 출력을 결합하면 더 나은 결과를 얻는 경우가 많다.

일반적으로 이는 한 모델의 출력을 다른 모델에 입력하는 방식으로 이루어지는데, 상당한 시간과 비용이 든다. 하지만 모스틱 팀은 AI 모델들이 텍스트 출력을 생성하지 않고도 서로 대화할 수 있는 방법을 찾아냈다. 이 방식이 확산된다면 오픈 웨이트 모델의 가치를 높여 Anthropic이나 OpenAI 같은 프런티어 랩의 폐쇄형 독점 모델과 더 잘 경쟁할 수 있게 될 수 있다.

말리셰바는 다양한 모델을 결합하는 것이 AI를 발전시키는 더 나은 방법이 될 수 있다고 말한다. "저는 개인적으로 미래에 거대한 단일(monolithic) 모델이 존재하거나 모델의 능력이 스케일링(모델을 더 크게 만들고 더 많은 데이터를 학습시키는 전략)에서 나온다고 생각하지 않습니다."라고 그녀는 말했다.

"모스틱이 프런티어 모델을 도메인 특화 모델 — 생물학, 물리학 등 — 과 짝지을 수 있게 해준다면 훨씬 더 많은 전문 모델이 학습될 것입니다."라고 모스틱 팀을 잘 아는 AI 소프트웨어 회사 Lovable의 테크 리드 블라디미르 아루스타미안(Vladimir Arustamian)이 말했다. "이 팀은 불과 몇 달 만에 저라면 몇 년 뒤에야 가능하리라 예상했던 것을 이미 구동하고 있습니다."

모스틱의 기술은 "대형 모델이 전체 과정을 처리하지 않고도 대형 모델 수준의 품질에 접근할 수 있게 해주며, 소형 모델 하나만 옆에서 돌려도 상당한 개선을 얻을 수 있다"고 모스틱의 기술을 잘 아는 구글 딥마인드(Google DeepMind)의 전 컴퓨터 과학자 칼 튈스(Karl Tuyls)는 말했다. 튈스는 이 방법이 모델을 최대한 효율적으로 운영해야 하는 사람이라면 누구에게나 당연한 선택이라고 덧붙였다.

제네바 대학교 교수이자 2010년 필즈상 수상자인 스타니슬라프 스미르노프(Stanislav Smirnov)는 모스틱의 수석 과학자다. 그는 두 AI 모델 사이의 공통점을 찾는 것이 놀라울 정도로 어렵다고 말한다. "아직 적절한 수학적 언어가 존재하지 않는 것 같습니다."라고 그는 말했다. 그 과정에서 모스틱의 접근법은 문자 그대로 그 간극에 다리를 놓는 방법이다. 스미르노프는 모스틱의 연구가 AI 모델이 실제로 어떻게 작동하는지, 그리고 이것이 인간 두뇌의 작동 방식과 어떻게 비교되는지에 대한 새로운 사실을 밝혀낼 수도 있다고 말한다. 스미르노프는 더 깊은 수학적 분석을 통해 AI 모델과 인간이 추론하는 방식의 공통점이 밝혀질 수 있다고 말한다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story I recently met with some brilliant Russian mathematicians who showed me a way for artificial intelligence models to communicate via something akin to machine telepathy. The mathematicians work for a startup called Mostik—the Russian word for bridge. It’s a nod to the group’s approach, which allows different models to interact using the mathematical values found in their weights —the things that determine how a prompt gets turned into an output. In practice, this means the capabilities of a larger model can be fed to a smaller model to ramp up its intelligence much more efficiently. The startup used the approach to build a model that has rocketed to the top of ARC-AGI 3, a notoriously difficult competition for AI models. (They wouldn’t tell me more because they want to win the contest.) To demonstrate the idea, however, they also created a bridge between two Chinese open-weight models: the largest version of GLM-5.2, which has 753 billion parameters; and a 4-billion-parameter version of Qwen-3.5 that can run on a mobile device. The resulting hybrid system costs one-twentieth of the full GLM model, and its performance is exactly halfway between the two. “It’s well-known in machine learning that ensembles of models perform better than individual ones,” Sasha Malysheva, Mostik’s CEO, told me over coffee. Malysheva, who developed the approach, shared a running joke inside the company: The future of AI is similar to guessing the weight of a pig. In math circles, it’s well-known that a handful of random people can more accurately estimate a pig’s weight than an expert when their guesses are combined and averaged. Much like communally eyeballing porcine heft, combining the outputs of several AI models often nets better results. Typically, this involves feeding the output of one model into another, which takes a good chunk of time and money. The Mostik team, however, figured out a way for AI models to talk to one another without producing text output. If it takes off, it could increase the value of open-weight models, allowing them to better compete with the closed, proprietary models offered by frontier labs like Anthropic and OpenAI . Malysheva says that combining lots of different models may turn out to be a better way to advance AI. “I personally do not think we will have a monolithic model [in the future] or that the capabilities of models will come from scaling,” she told me, referring to the strategy of making models larger and feeding them more data. “If Mostik makes it possible to pair frontier models with domain-specific models—think biology, physics, and so on—many more specialized models would be trained,” says Vladimir Arustamian, the tech lead at the AI software company Lovable, who knows the Mostik team. “This team has been at it for a matter of months and already has something running that I would have guessed was years out.” The Mostik technique means “you can approach large-model quality without the large model handling the entire loop, giving you substantial improvements with just a smaller model running alongside,” says Karl Tuyls, a former computer scientist at Google DeepMind who is familiar with the company’s tech. The method is a no-brainer for anyone tasked with running models as efficiently as possible, Tuyls says. Stanislav Smirnov, a professor at the University of Geneva and a 2010 Fields Medalist, is Mostik’s chief scientist. He says finding common ground between two AI models is surprisingly difficult. "There seems to be no appropriate mathematical language yet," he says. In the interim, Mostik’s approach is a way to quite literally bridge the gap. Smirnov says Mostik’s work could also perhaps reveal new things about how AI models actually function and how this compares to the workings of the human brain. Smirnov says a deeper mathematical analysis may reveal a commonality in the way both AI models and human beings reason over difficult problems. During our coffee meeting, Malysheva told me that she only discovered a talent for math after her older brother told her she wouldn’t be able to solve the Math Olympiad problems he was studying. A few years later, she was studying at one of the top schools in St. Petersburg. More recently, some peers warned her that the bridge approach would be too difficult to pull off. “They said it might be too hard for a young girl,” she says. “I decided I need to prove them wrong.” This is an edition of Will Knight’s AI Lab newsletter . Read previous newsletters here.