메뉴
HN
Hacker News 58일 전

로컬 기기용 초경량 이미지 생성 모델

IMP
9/10
핵심 요약

PrismML이 노트북과 스마트폰 같은 로컬 기기에서 고품질 이미지 생성을 가능하게 하는 40억 파라미터(4B) 모델 'Bonsai Image 4B'를 공개했습니다. 이 모델은 가중치를 1비트(1-bit) 또는 삼진법(Ternary) 형태로 압축하여, 기존 풀 정밀도(FP16) 모델 대비 메모리 사용량을 약 6~8배 획기적으로 줄였습니다. 특히 이 파라미터 클래스의 이미지 모델 중 최초로 아이폰에서 직접 구동될 수 있어, 온디바이스 AI 생성 기술의 새로운 지평을 열었다는 데 큰 의미가 있습니다.

번역된 본문

LAUNCH 1001011100 001011 11010 모든 게시물로 돌아가기

1비트 및 삼진법(Ternary) Bonsai Image 4B 소개: 로컬 기기를 위한 이미지 생성 2026년 5월 26일 • PrismML

오늘 우리는 노트북부터 스마트폰에 이르기까지 로컬 하드웨어에서 고품질 디퓨전 추론을 실행하도록 설계된 초소형 이미지 생성 모델 패밀리인 Bonsai Image 4B를 출시합니다.

Bonsai Image 4B는 두 가지 변형으로 제공됩니다:

1비트(1-bit) Bonsai Image 4B는 FP16 그룹별 스케일링 팩터가 적용된 이진법 {−1, +1} 트랜스포머 가중치를 사용하여 가중치당 1.125비트의 유효 비트를 가집니다. 이는 최대 압축을 목표로 하며, 메모리 압력, 대역폭 및 배포 공간이 주요 제약 조건일 때 적합한 선택입니다.

삼진법(Ternary) Bonsai Image 4B는 FP16 그룹별 스케일링 팩터가 적용된 {−1, 0, +1} 트랜스포머 가중치를 사용하여 가중치당 1.71비트의 유효 비트를 가집니다. 추가된 제로(0) 상태는 모델에 더 많은 표현의 유연성을 제공하여 매우 콤팩트함을 유지하면서도 시각적 품질과 프롬프트 충실도를 향상시킵니다.

그 결과 이미지 생성을 위한 새로운 배포 체제가 마련되었습니다. 이전에는 이 클래스의 모델이 접근할 수 없었던 기기에서도 우수한 결과물, 개방형 가중치(Open weights)를 통한 실용적인 로컬 추론이 가능해졌습니다. 우리가 아는 한, Bonsai Image 4B는 해당 파라미터 클래스에서 아이폰에서 직접 실행되는 최초의 이미지 모델입니다.

로컬 생성을 위해 설계 로컬 이미지 생성은 모델이 기기의 메모리 예산 내에 들어가야 한다는 엄격한 제약 조건에서 시작됩니다. 4B(40억 파라미터) 클래스 이미지 모델의 경우, 디퓨전 트랜스포머가 모델에서 가장 큰 부분을 차지하며 생성 중에 반복적으로 실행되는 부분입니다. 각 노이즈 제거(denoising) 단계마다 트랜스포머가 다시 호출되므로, 트랜스포머 크기는 메모리 압력, 대역폭 수요 및 로컬 추론 속도에 직접적인 영향을 미칩니다.

Bonsai Image 4B는 FLUX.2 Klein 4B를 기반으로 구축되었습니다. 아키텍처는 그대로 유지하면서 트랜스포머 가중치가 표현되는 방식만 변경했습니다. 가중치를 이진 및 삼진 형식으로 변환함으로써 Bonsai는 로컬 배포에 있어 가장 중요한 이미지 파이프라인의 크기를 줄였습니다.

모델 디퓨전 트랜스포머 크기 FP16 대비 감소 비율
FLUX.2 Klein 4B 7.75 GB 1.0x
1비트 Bonsai Image 4B 0.93 GB 8.3x
삼진법 Bonsai Image 4B 1.21 GB 6.4x

표 1: 모델별 디퓨전 트랜스포머 공간 크기

이진 레이어는 전체 정밀도 트랜스포머 가중치에 비해 약 14배의 크기 감소를 제공합니다. 프로젝션 레이어라고 불리는 정밀도에 민감한 소규모 지원 텐서(약 5%)는 FP16으로 유지되므로, 최종 1비트 Bonsai Image 4B 트랜스포머의 크기는 0.93 GB로 전체 정밀도의 FLUX.2 Klein 4B(7.75 GB) 대비 8.3배 감소했습니다.

삼진법 변형도 동일한 구조를 따릅니다. 삼진 레이어는 약 10배의 크기 감소를 제공하며, 최종 삼진법 Bonsai Image 4B 트랜스포머의 크기는 1.21 GB로 전체 정밀도 트랜스포머 대비 6.4배 감소했습니다. 1비트 모델보다 약간 크지만, 추가된 제로(0) 상태 덕분에 시각적 품질과 프롬프트 충실도가 향상됩니다.

압축된 텍스트 인코더와 FP16 VAE를 포함할 경우, Apple Silicon 기기에서의 배포 페이로드는 1비트 Bonsai Image 4B의 경우 3.42 GB, 삼진법 Bonsai Image 4B의 경우 3.88 GB입니다. 비교를 위해, 전체 정밀도 FLUX.2 Klein 4B는 15.97 GB의 배포 페이로드를 필요로 합니다.

런타임 시 프롬프트 인코딩이 완료된 후에는 텍스트 인코더가 오프로드되므로 평균 메모리 사용량은 전체 페이로드보다 적습니다. 512x512 이미지를 생성할 때 평균 활성 메모리는 이진 및 삼진 모델의 경우 각각 1.5 GB 및 1.96 GB로, 기존 FLUX.2 Klein 4B의 11.74 GB와 비교해 각각 7.8배, 6.0배 감소했습니다.

1024x1024 이미지의 경우, 평균 활성 메모리는 이진 및 삼진 모델에서 각각 1.95 GB 및 2.38 GB로, 기존 FLUX.2 Klein 4B의 14.39 GB와 비교해 각각 7.4배, 6.0배 감소했습니다.

이러한 메모리 공간의 감소는 모델이 실행될 수 있는 환경을 완전히 바꿔놓았습니다. 당사의 배포 스택은 Apple 실리콘이 탑재된 아이폰, 아이패드, 맥과 CUDA GPU를 지원하며, Apple 하드웨어에서는 MLX 저비트 경로를, CUDA에서는 Gemlite 저비트 GEMM 커널을 사용합니다.

아이폰 17 프로 맥스(iPhone 17 Pro Max)에서는 전체 정밀도 FLUX.2 Klein 4B 파이프라인이 기기의 메모리 예산에 맞지 않지만, 두 가지 Bonsai 변형 모두 원활하게 구동됩니다.

원문 보기
원문 보기 (영어)
LAUNCH 1001011100 001011 11010 Back to all posts Introducing 1-bit and Ternary Bonsai Image 4B: Image Generation for Local Devices May 26, 2026 • PrismML Today we’re releasing Bonsai Image 4B , a family of compact image-generation models designed to run high-quality diffusion inference on local hardware: from laptops to phones. Bonsai Image 4B comes in two variants: 1-bit Bonsai Image 4B uses binary {−1, +1} transformer weights with an FP16 group-wise scaling factor, giving 1.125 effective bits per weight. It targets maximum compression and is the right fit when memory pressure, bandwidth, and the deployment footprint are the primary constraints. Ternary Bonsai Image 4B uses {−1, 0, +1} transformer weights with an FP16 group-wise scaling factor, giving 1.71 effective bits per weight. The additional zero state gives the model more representational flexibility, improving visual quality and prompt fidelity while remaining extremely compact. The result is a new deployment regime for image generation: capable outputs, open weights, and practical local inference on devices that were previously out of reach for this class of model. To our knowledge, Bonsai Image 4B is the first image model in its parameter class to run directly on an iPhone . Built for local generation Local image generation starts with a hard constraint: the model has to fit within the device’s memory budget. For a 4B-class image model, the diffusion transformer is the largest part of the model and the part that runs repeatedly during generation. Each denoising step invokes the transformer again, so transformer size directly shapes memory pressure, bandwidth demand, and local inference speed. Bonsai Image 4B is built from the FLUX.2 Klein 4B. It keeps the architecture intact but changes how the transformer weights are represented. By moving those weights into binary and ternary form, Bonsai reduces the part of the image pipeline that matters most for local deployment. Model Diffusion Transformer Reduction vs FP16 FLUX.2 Klein 4B 7.75 GB 1.0x 1-bit Bonsai Image 4B 0.93 GB 8.3x Ternary Bonsai Image 4B 1.21 GB 6.4x Table I: Diffusion transformer footprint for models. The binary layers provide roughly a 14x reduction relative to full-precision transformer weights. A small set of precision-sensitive supporting tensors (~5%), called the projection layers, remains in FP16 so the final 1-bit Bonsai Image 4B transformer is 0.93 GB : an 8.3x reduction from the 7.75 GB full-precision FLUX.2 Klein 4B. The ternary variant follows the same structure. Its ternary layers provide roughly a 10x reduction and the final Ternary Bonsai Image 4B transformer is 1.21 GB , a 6.4x reduction from the full-precision transformer. It is slightly larger than the 1-bit model, but the additional zero state improves visual quality and prompt fidelity. Including the compressed text encoder and FP16 VAE, the Apple Silicon deployment payload is 3.42 GB for 1-bit Bonsai Image 4B and 3.88 GB for Ternary Bonsai Image 4B. For comparison, the full precision FLUX.2 Klein 4B requires a deployment payload of 15.97 GB. Since, at runtime, the text encoder is offloaded after prompt encoding, the mean memory usage is smaller than the total payload. When generating a 512x512 image, the mean-active memory is 1.5 GB and 1.96 GB, for the binary and ternary models, compared to 11.74 GB for the original FLUX.2 Klein 4B (a reduction of 7.8x and 6.0x, respectively). For a 1024x1024 image, the mean-active memory is 1.95 GB and 2.38 GB, for the binary and ternary models, compared to 14.39 GB for the original FLUX.2 Klein 4B (a reduction of 7.4x and 6.0x, respectively). This reduction in memory footprint changes where the model can run. Our deployment stack supports Apple Silicon iPhones, iPads and Macs and CUDA GPUs, using MLX low-bit paths on Apple hardware and Gemlite low-bit GEMM kernels on CUDA. On iPhone 17 Pro Max, the full-precision FLUX.2 Klein 4B pipeline does not fit within the device memory budget, while both Bonsai Image variants run on-device. Video I: Image generation on Bonsai Studio In practice, Bonsai Image 4B generates a 512x512 image in 9.4 seconds on an iPhone 17 Pro Max and about 6 seconds on Mac M4 Pro. On Mac M4 Pro, Bonsai Image 4B is up to 5.6x faster than the stock full-precision MFLUX pipeline. Benchmarking performance Compression only matters if the model remains useful. We evaluated Bonsai Image 4B across three complementary benchmarks: GenEval for object composition and attribute binding; HPSv3 human preference and aesthetic quality; DPG-Bench dense prompt following and semantic faithfulness. Model Diffusion Transformer Footprint (GB) GenEval HPSv3 DPG-Bench Size reduction relative to FLUX.2 Klein 4B Performance relative to FLUX.2 Klein 4B 1-bit Bonsai Image 4B 0.93 0.671 11.15 0.822 8.3x 88% Ternary Bonsai Image 4B 1.21 0.723 12.22 0.851 6.4x 95% FLUX.2 Klein 4B 7.75 0.819 12.84 0.853 1x 100% SDXL 5.14 0.3 10.05 0.74 1.5x 67% BK-SDM-Small 0.98 0.297 3.05 0.559 7.9x 42% Stable Diffusion 1.5 1.72 0.396 4.2 0.601 4.5x 51% PixArt-Σ XL 2 1.2 0.541 11.93 0.769 6.4x 83% Table II: Image quality benchmark comparison across Ternary Bonsai Image 4B and other models. Ternary Bonsai Image 4B is the quality-oriented variant. At 1.21 GB, it retains 95% of the FLUX.2 Klein 4B accuracy across GenEval, HPSv3, and DPG-Bench, while reducing the diffusion transformer footprint by 6.4x. 1-bit Bonsai Image 4B is the footprint-oriented variant. It brings the diffusion transformer below 1 GB, an 8.3x reduction, while still delivering strong benchmark scores across the same three evaluations (it retains 88% of the accuracy of FLUX.2 Klein 4B). Together, the two variants move the quality–footprint frontier. Bonsai Image remains competitive with modern 4B-class image models while using a fraction of their diffusion-transformer footprint. At the same time, it substantially outperforms smaller models with similar memory footprints. That is the same Pareto shift we have seen in our prior Bonsai language models. Bonsai Image brings modern diffusion-transformer behavior into a memory range that previously belonged to much smaller, lower-capability models. Why this is important Image generation is not only a model-quality problem. It is also a deployment problem. Cloud APIs will continue to be the right choice for many products. But cloud-only generation imposes certain product constraints: every prompt is a remote request, every iteration carries marginal serving cost, and every interaction adds round-trip latency. That matters because image generation is naturally iterative. Users rarely stop at one image. They revise prompts, compare outputs, generate variations, discard failures, and try again. When each attempt is a server-side job, the creative loop becomes something users have to meter and wait for. Local inference changes that. Once the model fits on the device, generation can sit directly inside the product experience. It becomes cheaper to run, faster to iterate on, and easier to use in environments where prompts, and generated assets should remain private. Bonsai Image 4B is a step toward that deployment regime: capable image generation running closer to the user, on hardware they already own. Availability Both 1-bit and Ternary Bonsai Image 4B will be released with open weights and code under the Apache 2.0 license . With this launch, we are also launching Bonsai Studio, its iOS app for trying Bonsai Image 4B directly on iPhone. Join Us PrismML emerged from a team of Caltech researchers and was founded with support from Khosla Ventures, Cerberus and Google. We’ve spent years tackling one of the field’s hardest problems: compressing neural networks without sacrificing their reasoning ability. If you want to help build the next generation of state-of-the-art AI, we’d love to hear from you. Check out our careers page . Resources Whitepaper Hugging Face WebGPU demo Bonsai Studio for iPhone GitHub Back to all posts Announcing 1-bit Bonsai: T