메뉴
HN
Hacker News • 8일 전

본사이 2 27B, 9배 작은 용량으로 거의 무손실 압축

IMP
7/10
핵심 요약

PrismML이 삼진법(trinary) 가중치 기반의 'Ternary Bonsai 2 27B'를 공개했습니다. Qwen3.8 27B를 기반으로 하며, 전체 원본 모델 대비 9배 이상 작은 5.9GB 크기로 벤치마크 성능의 98.2%를 유지합니다. 로컬 기기에서 코딩 에이전트, 컴퓨터 사용, 멀티모달 작업 등 실질적인 지식 노동이 가능해졌다는 점에서 주목할 만합니다.

번역된 본문

본사이 2 27B 소개: 9배 작은 크기로 거의 무손실 압축 (2026년 9월 17일, PrismML)

두 달 전, 우리는 첫 번째 Bonsai 27B 모델을 출시하며 27B급 멀티모달 모델을 로컬 기기에서 효율적으로 실행할 수 있을 만큼 압축할 수 있음을 보여주었습니다. 오늘 우리는 Bonsai 시리즈 중 가장 뛰어난 모델인 Ternary Bonsai 2 27B를 출시합니다. Qwen3.8 27B를 기반으로 하는 Ternary Bonsai 2 27B는 Bonsai 시리즈의 정체성인 획기적으로 작은 메모리 사용량, 높은 로컬 처리량, 우수한 에너지 효율이라는 배포 특성을 유지하면서 더 강력한 추론, 코딩, 비전, 에이전트 능력을 제공합니다.

Ternary Bonsai 2 27B는 {-1, 0, +1} 삼진법(trinary) 가중치와 FP16 그룹별 스케일링을 사용하여 가중치당 유효 비트수 1.76비트, 전체 모델 크기 5.9GB를 달성했습니다. 저비트 표현은 언어 모델 전체에 엔드투엔드로 적용되었습니다. 262K 토큰 컨텍스트 윈도우와 텍스트·이미지 멀티모달 입력을 지원하며, Apache 2.0 라이선스로 공개됩니다.

원본 고정밀(full-precision) 모델과 비교하여 Ternary Bonsai 2 27B는 9배 이상 작으면서도 종합 벤치마크 성능의 98.2%를 유지합니다. 이 수준의 성능 유지율에서 압축은 배포의 잠금 해제가 됩니다. 훨씬 더 많은 환경에서 실행 가능한 크기로 거의 동일한 능력을 제공하는 것입니다.

첫 번째 Bonsai 27B 출시와 비교한 변경점 첫 번째 Bonsai 27B는 로컬 기기에서 27B급 지능을 실행하는 실용적인 방법을 제시한 중요한 이정표였습니다. Bonsai 2 27B는 그 다음 단계에 집중합니다. 즉, 실제 로컬 애플리케이션에 필요한 모델 품질과 런타임 성능을 개선하는 것입니다.

이전 세대 대비 Bonsai 2 27B의 개선점:

  • 더 강력한 기반 모델인 Qwen3.8 27B 적용
  • 고정밀 원본 모델 대비 종합 능력 유지율 98.2%로 향상
  • 추론, 코딩, 비전, 장기 에이전트 성능 개선

동일한 배포 조건에서 더 높은 능력 추론, 수학, 코딩, 지시 이행, 비전, 에이전트 도구 사용을 아우르는 벤치마크 스위트에서 Ternary Bonsai 2 27B는 83.9점을 기록하며 Qwen3.8 27B 종합 성능의 98.2%를 유지했습니다.

주요 결과는 종합 점수 자체가 아니라 능력이 어디에서 유지되는가입니다. 코딩 에이전트, 도구 사용 시스템, 멀티모달 워크플로, 장기 작업은 작은 오류가 여러 단계에 걸쳐 누적될 수 있어 모델 성능 저하에 특히 민감합니다. Bonsai 2 27B는 바로 이러한 영역에서 메모리 사용량의 극히 일부만으로 고정밀 모델 성능의 대부분을 보존합니다.

고정밀 모델 및 다른 저비트 대안들과 비교했을 때, Bonsai 2 27B는 '지능 밀도(intelligence density)' 측면에서 두드러진 이례적 사례입니다. 많은 저비트 대안들은 코딩, 비전, 에이전트 도구 사용에서 상당한 능력을 포기해야만 배포 가능해집니다. Bonsai 2 27B는 더 높은 능력과 더 낮은 메모리 사용량 양쪽으로 프론티어를 밀어붙입니다.

데모 I: NVIDIA GeForce RTX 5090에서 Ternary Bonsai 2 27B로 구동되는 Cline 기반 코딩 에이전트 데모 II: NVIDIA GeForce RTX 5090에서 Ternary Bonsai 2 27B 모델로 구동되는 컴퓨터 사용

Bonsai 2 27B를 통해 로컬 모델은 이제 실질적인 지식 노동을 수행할 수 있습니다. 코딩 에이전트 루프, 컴퓨터 사용 워크플로, 비공개 문서 분석, 멀티모달 디버깅, 그리고 로컬 모드와 결합된 하이브리드 오케스트레이션 등이 가능해집니다.

원문 보기
원문 보기 (영어)
About Blog News Careers Docs LAUNCH 1001011100 11010 1 001 Back to all posts Introducing Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint September 17, 2026 • PrismML Two months ago, we released our first Bonsai 27B models and showed that a 27B-class multimodal model could be compressed enough to run efficiently on a local device. Today, we’re releasing Ternary Bonsai 2 27B, our most capable model yet. Based on Qwen3.8 27B, Ternary Bonsai 2 27B brings stronger reasoning, coding, vision, and agentic capability to the Bonsai series while preserving the deployment profile that defines it: a dramatically smaller memory footprint, high local throughput, and better energy efficiency. Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight and a total model footprint of 5.9GB . The low-bit representation is applied end to end across the language model. It supports a 262K-token context window , multimodal text-and-image input, and is released under the Apache 2.0 license. Against its full-precision counterpart, Ternary Bonsai 2 27B is more than 9x smaller while retaining 98.2% of aggregate benchmark performance . At this level of retention, compression becomes a deployment unlock: nearly the same capability, in a footprint that can run in far more places. What changed from the first Bonsai 27B release Our first Bonsai 27B release was an important milestone, offering a practical way to run 27B-class intelligence on local devices. Bonsai 2 27B focuses on the next step: improving the model quality and runtime performance needed for real-world local applications. Compared with the previous Bonsai 27B generation, Bonsai 2 27B brings: a stronger base model, Qwen3.8 27B higher aggregate capability retention of 98.2% against the full-precision model improved reasoning, coding, vision, and long-horizon agentic performance Higher capability at the same deployment point Across a benchmark suite spanning reasoning, math, coding, instruction following, vision, and agentic tool use, Ternary Bonsai 2 27B scores 83.9 , retaining 98.2% of Qwen3.8 27B’s aggregate performance. Capability Ternary Bonsai 2 27B Qwen3.8 27B Qwen3.6 27B Agentic & Tool Calling τ²-bench, BFCLv3 77.57 79.74 80.05 Coding HumanEval+, LiveCodeBench v6, MBPP+, BigCodeBench 81.58 82.17 82.57 Instruction Following IFBench, IFEval 82.66 81.25 74.53 Knowledge & Reasoning MMLU-Redux, GPQA Diamond, AA-LCR 83.95 86.66 84.71 Math AIME 2026, AIME 2025, GSM8K, MATH-500 96.57 97.06 94.64 Vision CharXiv, A-OKVQA, OmniDocBench v1.6, RealWorldQA, OCRBench v2 78.59 81.64 79.82 Overall 83.9 85.4 83.6 ‍ Figure I: Benchmark scores of Ternary Bonsai 2 27B (thinking mode) compared with the full-precision Qwen3.8 27B and Qwen3.6 27B baselines. Full per-benchmark results are in the whitepaper . ‍ The key result is not only the aggregate score, but where the capability is retained. Coding agents, tool-use systems, multimodal workflows, and long-horizon tasks are particularly sensitive to model degradation because small errors can compound over many steps. Bonsai 2 27B preserves much of the full-precision model’s performance in exactly these areas while operating at a fraction of the memory footprint. Compared with the full-precision model and other low-bit alternatives, Bonsai 2 27B stands out as an outlier on intelligence density. Many low-bit alternatives become deployable only by giving up meaningful capability in coding, vision, or agentic tool use. Bonsai 2 27B pushes the frontier toward both higher capability and lower memory usage. Demo I: Coding agents with Cline, powered by Ternary Bonsai 2 27B on NVIDIA GeForce RTX 5090. Demo II: Computer use powered by Ternary Bonsai 2 27B model on NVIDIA GeForce RTX 5090. ‹ › With Bonsai 2 27B, local models can start to take on real knowledge work: coding-agent loops, computer-use workflows, private document analysis, multimodal debugging, and hybrid orchestration where local models handle sensitive or high-frequency tasks while escalating selectively to the cloud. Throughput and energy efficiency Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision. For coding assistants, higher throughput means faster edit-debug loops. For multimodal agents, it means quicker iterations over screenshots, documents, and tool calls. For private local workflows, better energy efficiency means more useful inference on the same device, longer battery life, and a more realistic path to assistants that can stay available in the background without constantly calling the cloud. Why this release matters Compared to Ternary Bonsai 27B, the new Ternary Bonsai 2 27B has closed the retention gap between the full precision model from 95% to over 98%. This is a significant improvement that makes the current release practically “lossless”. It further cements the notion that low-bit models can be the best way to deploy AI. That has implications well beyond local inference. Low-bit models can change the economics and architecture of AI systems across devices, workstations, and datacenters: fitting larger models into the same memory envelope, serving more users on the same hardware, reducing energy per inference, and enabling hybrid systems that dynamically decide what should run locally and what should run in the cloud. The question will increasingly be not just how capable a model is, but how much useful intelligence can be delivered within a given memory, compute, and power budget. If capability can continue to scale while those requirements fall dramatically, the deployment envelope for future models expands across the stack: from personal devices to large-scale datacenters. Platform Coverage Bonsai 2 27B runs on NVIDIA GPUs via CUDA and on Apple devices (Mac, iPhone, iPad) via MLX, through custom low-bit kernels. Model weights are available today under the Apache 2.0 License. Full technical details of our compression, evaluation, and benchmarking processes are available in our whitepaper . Work with Us We work with teams to tailor Bonsai models to their applications, from post-training on domain-specific data to optimizing inference for target hardware. If you’re building AI products with tight memory, latency, or power requirements, we’d love to explore how Bonsai can help. Reach out at contact@prismml.com . Join Us Prism ML emerged from a team of Caltech researchers and was founded with support from Khosla Ventures, Cerberus, and Google, with continuing support from Samsung. We've spent years tackling one of the field's hardest problems: compressing neural networks without sacrificing their reasoning ability. If you want to help build the next generation of state-of-the-art AI, we'd love to hear from you. Check out our careers page . Back to all posts Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone July 14, 2026 Today we're announcing Bonsai 27B, our multimodal flagship: ternary at 5.9GB for laptops, 1-bit at 3.9GB for an iPhone 17 Pro, with a 262K-token context. PrismML Launches Bonsai 2 27B, Its Most Capable Model Yet September 17, 2026 New flagship model brings 27B-class reasoning, coding, vision, and agentic capability into a dramatically smaller, faster deployment footprint Thanks, we’ll keep you posted! Something went wrong. Resources Demo Whitepaper Docs Models Hugging Face GitHub Follow X Discord LinkedIn Contact us Have a question, partnership idea, or a project that needs efficient intelligence? Reach out—we’d love to hear from you. Thanks, your message has been received! Oops! Something went wrong while submitting the form.