메뉴
BL
MarkTechPost • 23일 전

퍼플렉시티, 애플 실리콘용 추론 엔진 '릴리' 오픈소스화

IMP
7/10
핵심 요약

퍼플렉시티가 자사 Perplexity Computer의 Hybrid Compute를 뒷받침하는 로컬 추론 엔진 'Lily(릴리)'를 오픈소스로 공개했습니다. Rust로 작성되고 애플 실리콘용 커스텀 Metal 커널을 포함해, 40코어 128GB M5 Max에서 MLX-LM 대비 프리필 1.23배, 디코드 1.35배의 처리량을 보입니다.

번역된 본문

퍼플렉시티(Perplexity)는 Perplexity Computer의 Hybrid Compute를 뒷받침하는 로컬 추론 엔진인 'Lily(릴리)'를 오픈소스로 공개했습니다. Rust로 작성되었으며, 단 하나의 모델(Qwen3.6-35B-A3B)과 하나의 칩 패밀리(애플 실리콘)를 위해 커스텀 Metal 커널을 구현한 것이 특징입니다. 40코어 128GB M5 Max 환경에서 MLX-LM 대비 프리필(prefill) 처리량 평균 1.23배, 디코드(decode) 처리량 평균 1.35배의 성능을 기록했습니다.

원문 보기
원문 보기 (영어)
Perplexity has open sourced Lily, the local inference engine behind Hybrid Compute in Perplexity Computer. Built in Rust with custom Metal kernels for one model on one chip family, it averages 1.23x MLX-LM's prefill throughput and 1.35x its decode throughput on a 40-core, 128 GB M5 Max. The post Perplexity Open Sources Lily: A Rust + Metal Inference Engine for Qwen3.6-35B-A3B on Apple Silicon appeared first on MarkTechPost.