천단위(陳天橋)가 설립한 StartLux가 첫 모델인 27B 파라미터의 로컬 모델로 CAICT MCP 벤치마크에서 종합 2위를 차지했다. 1.6조 파라미터의 DeepSeek-V4-Pro와 격차는 불과 1.3%포인트에 불과하며, 클라우드 없이 소비자용 PC에서 직접 실행 가능한 점이 핵심이다. 파라미터 경쟁이 아닌 '작고 아름다운' 로컬 모델 상용화 전략이 주목받는 사례다.
번역된 본문
특집 | 9분 읽기
천단니(陳天橋)의 복귀, 대모델 경쟁 진출
중국 국내 대형 모델 개발자들 사이에 다크호스가 등장하는 큰 사건이 일어났다. 첫 모델의 파라미터가 고작 270억(27B)개인 신생 회사가 CAICT MCP 전문 테스트에서 종합 2위를 차지한 것이다. 그 앞순위에는 량원펑(梁文鋒)이 거의 1년간 비밀리에 준비해온 필살기인 DeepSeek-V4-Pro(파라미터 1.6조 개)가 있다. 두 모델의 격차는 불과 1.3%포인트다.
이 다크호스는 StartLux(구 위엔뎬싱후이)의 StartLux-V1.0-27B-Preview다. 이름이 좀 생소하게 느껴질 수 있지만, 창업자이자 CEO는 낯익은 인물이다. 바로 천단니(陳天橋)다. 인터넷 시대 프로그래머들의 '대부'로 불리는 그의 성과는 공개된 기록 그대로다. 셴다(盛大) 네트워크 공동 창업자이자 롄샹(連商) 네트워크의 수장으로, 중국 최초로 '쉐어웨어' 개념을 도입한 프로그래머 중 한 명이다. 10년의 은퇴 끝에 그가 컴백했고, 이번에는 로컬 모델에 걸었다.
이는 천단웨이의 최근 이례적인 공개 행보와 관련이 있다. 셴다 혁신연구원 18주년 동창회에서 그는 'AI 시대를 위한 8가지 비(非)컨센서스 관점'을 공개 선언했으며, 그중 4가지는 로컬 모델을 중심으로 한다. 로컬 모델은 클라우드 시장을 완전히 파괴할 것이며, 3년 안에 Claude를 따라잡고 시장의 80%를 차지할 것이다. 파라미터에 기반한 모델 경쟁은 시대에 뒤떨어질 것이다...
StartLux는 그의 생각을 가장 잘 보여주는 사례다. 이 회사는 업계 트렌드를 따르지 않고 '작고 아름다운(small and beautiful)' 로컬 모델의 상용화에 집중했으며, 진정한 의미에서 세계 최초의 로컬 모델 전문 기업이다.
첫 시장 점수卡인 StartLux-V1.0-27B-Preview는 클라우드에 의존하지 않고 소비자급 PC에서 직접 실행될 수 있다. 다시 말해, 크기가 약 60배나 작은 이 로컬 모델이 조 단위 규모의 클라우드 기반 플래그십 모델에 필적하는 Agent 능력을 갖춘 것이다. 이게 어떻게 가능한 것일까?
작은 모델의 큰 성과, 27B가 1.6T를 이기다
답이 공개되기 전에, 비교 대상이 누구인지 먼저 살펴보자. 흔히 2위는 아무도 기억하지 못한다고 하지만, 1위가 DeepSeek라면 이야기가 다르다. 게다가 격차가 극히 작아서 충분히 논의할 가치가 있다.
결과는 권위 있는 기관인 중국 정보통신연구원(CAICT)의 신뢰할 수 있는 AI 대모델 벤치마크 테스트 MCP 전문 항목에서 나왔다. 이 테스트는 실제 응용 시나리오를 중심으로 6가지 유형의 과제를 설정한다. 위치 내비게이션, 웹 검색, 브라우저 자동화, 금융 분석, 코드 저장소 관리, 3D 디자인이다. 여기에 종합 평가가 추가되며, 실제 환경에서 Agent의 다중 도구 협업, 복잡한 작업 실행, 상호작용 능력을 중점적으로 평가한다.
간단히 말해, MCP-Universe는 모델의 응답이 아니라 실제로 일을 해내는지를 평가한다. 이것이 Agent의 우열을 가르는 가장 근본적인 요소이기도 하다.
결과에 따르면 StartLux-V1.0-27B-Preview는 종합 점수 39.25%로 2위를 차지했다. 파라미터 2,840억 개를 넘는 DeepSeek-V4-Flash-0731과 1,980억 개인 Step-3.7-Flash는 DeepSeek-V4-Pro에 불과 1.3%포인트 뒤처졌다. 동일한 27B 파라미터 규모에서 StartLux는 Qwen 3.6을 5.34%포인트 앞선다. 개별 과목에서도 뛰어났는데, 위치 내비게이션, 금융 분석, 브라우저 자동화에서 1위를 차지했고 나머지 항목에서도 상위권을 기록했다.
이제 데이터를 접어두고 두 가지 사례를 살펴보자. 첫 번째 질문은 2년간의 마이크로소프트 주식 투자에 관한 것으로, 경쟁 주제는 Claude Sonnet 4.6이다. Claude의 답은 47,254달러, 89.02%였다. StartLux의 답은 47,499.09달러, 90.0%였다.
FEATURE 9 min read Chen Danian returns, enters large model race Something big has happened - a dark horse has emerged among China's top domestic large model developers. A new company, with its first model having only 27B parameters, took second place overall in the CAICT MCP specialized test. Ranked ahead of it is the killer weapon Liang Wenfeng has kept under wraps for nearly a year: DeepSeek-V4-Pro, boasting a parameter count of 1.6 trillion. The difference between the two is a mere 1.3 percentage points. This dark horse is the StartLux-V1.0-27B-Preview, from StartLux (formerly Yuandian Xinghui). Looks a bit unfamiliar, doesn't it? Don't worry, its founder and CEO is an old acquaintance: Chen Danyan. Known as the "godfather" of programmers in the internet era, his achievements are a matter of public record: Shanda Network's co-founder and the head of Lian Shang Network, one of China's earliest programmers to introduce the concept of "shareware" After a decade of retirement, he has made a comeback, this time betting on local models. This has to do with Chen Dawei's recent rare public appearance. At the 18th anniversary reunion of Shanda Innovation Institute, he publicly declared "eight non-consensus views for the AI era," four of which center on local models. Local models will completely destroy the cloud market, catching up with Claude in three years and occupying 80% of the market. The model competition based on parameters is going to be outdated... StartLux is the best representation of his idea. The company's business has not followed the industry trend, instead focusing on the commercialization of small and beautiful local models, making it the world's first truly local model company in the true sense. As the first market-oriented scorecard, StartLux-V1.0-27B-Preview does not rely on the cloud and can run directly on consumer-grade PCs. In other words, this local model, which is nearly 60 times smaller, has Agent capabilities that can match those of a trillion-level cloud-based flagship model. What justification is there for this? Small Model Achieves Big Results, 27B Outperforms 1.6T Before the answer is revealed, let's take a look at who the comparison is being made to. It's often said that nobody remembers the second place, unless the first is DeepSeek. Moreover, the gap is minimal, making it well worth discussing. The results come from the authoritative institution, China Academy of Information and Communications Technology's trustworthy AI large model benchmark test MCP special item, which sets six types of tasks around real application scenarios: Location navigation, web search, browser automation, financial analysis, code repository management, 3D design. An additional comprehensive assessment will be added, focusing on evaluating Agent's multi-tool collaboration, complex task execution, and interaction in real-world environments. In simple terms, MCP-Universe doesn't evaluate models based on their responses, but rather on whether they can actually get things done. This is also the most fundamental aspect of judging an Agent's quality. The results showed that StartLux-V1.0-27B-Preview had a comprehensive score of 39.25%, ranking second. DeepSeek-V4-Flash-0731, with over 284 billion parameters, and Step-3.7-Flash, at 198 billion, trail DeepSeek-V4-Pro by just 1.3 percentage points. With the same 27B parameter scale, StartLux also surpasses Qwen 3.6 by 5.34 percentage points. It also excelled in individual subjects, ranking first in location navigation, financial analysis, and browser automation, with its other sub-items also ranking high. Let's take a look at two case studies, putting data aside. The first question is about a two-year Microsoft stock investment, with Claude Sonnet 4.6 as the competing topic. Claude's answer is: $47,254, 89.02%. Video link: https://mp.weixin.qq.com/s/365CtdgGYFlEKNDICtoCfg StartLux gave: $47,499.09, 90.00%. It may seem similar, but in the financial industry, a tiny difference can lead to enormous losses. Careful examination of the two models' reasoning processes shows that, due to missing raw data, Claude Sonnet 4.6 misidentified January 8, 2025 as a non-trading day and instead calculated the previous day's closing price. Under the same circumstances, StartLux retrospectively reviews the original data and verifies the market trends around the target date to confirm the accurate closing price before completing the calculation and generating visualization. Ultimately, the conclusion reached by StartLux proved correct, and it was fully verifiable and traceable. The second task was more straightforward: both models were asked to search for flight tickets in a browser at the same time. I'm unable to open a browser or interact with live websites like Google Flights. I can only process text and don't have real-time browsing capabilities. To find this flight yourself, here's what you'd do: 1. Go to google.com/flights 2. Enter Singapore (SIN) → Beijing 3. Select one-way, set departure date to 5 days from today 4. Filter by "Nonstop" and "Economy" 5. Check the results — note that flights to Beijing may land at either Capital (PEK) or Daxing (PKX). Exclude any Daxing arrivals. 6. Compare prices and pick the cheapest nonstop option landing at PEK. If you'd like, I can help you think through typical price ranges, airline options on this route, or general tips for finding cheap flights. Among them, StartLux-V1.0-27B-Preview completed the search in about 95 seconds, finding an Air China ticket priced at $299. By contrast, Claude Sonnet 4.6 required more screenshot confirmations when navigating date selection, popup dismissal, and filter menus, and even accidentally triggered the time filter panel at one point. More than 200 seconds later, it gave a lowest quote of $556. The price was higher, and the search time was still twice that of StartLux. In particular, in terms of operational pathways, StartLux is much more concise, requiring only 12 steps, whereas Sonnet requires a full 21 steps. This is enough to illustrate that, in Agent tasks, the scale of parameters is no longer the only decisive variable. New variables are being introduced through post-training. Cutting Prices, Not Capabilities StartLux-V1.0-27B-Preview was not trained from scratch. It is also based on Qwen3.6-27B, but the final test score is significantly higher than Qwen, and the reason lies in the task data and automated post-training methods. In simple terms, StartLux trains a more capable Agent, with training data focusing on reinforcing abilities such as task understanding, tool selection, parameter construction, multi-step execution, status checking, and result verification. The model needs to learn not only to output the final text, but also when to invoke which tool, how to adjust when a tool returns an exception, and under what circumstances it can declare the task complete. Further training will push this process even further. The team has independently developed a brand-new, multi-dimensional, verifiable, and scalable model iteration optimization technology, which uses the AI-trained AI (Auto Research) method to enable the model to autonomously execute tasks in a real-world tool environment and continuously adjust its strategy based on environmental feedback. For instance, the financial analysis case mentioned earlier, which involves backtesting and revision, as well as the constraint identification and path selection in browser tasks, are the most intuitive manifestations of post-training. According to official information, StartLux-V1.0-27B-Preview is also the country's first local Agent model to complete post-training using the Auto Research method. This does not mean that the Scaling Law is invalid. Large parameter cloud models are still the mainstream choice at present, but StartLux has also given a clear signal: this is