메뉴
HN
Hacker News • 38일 전

Shoehorn – 어떤 모델이든 내 컴퓨터에 맞게 양자화해 실행

IMP
6/10
핵심 요약

Shoehorn은 사용자의 하드웨어 메모리 예산에 맞춰 허깅페이스(Hugging Face) 인기 모델들을 양자화(quantize)하여 로컬에서 실행할 수 있게 해주는 오픈소스 도구입니다. 브라우저에서 내 장비에 맞는 모델을 스캔해 품질 순으로 보여주고, llama.cpp를 백엔드로 사용해 원 클릭으로 채팅까지 가능합니다. 맥(macOS), 리눅스, 윈도우를 지원하며 Homebrew나 바이너리 다운로드로 설치할 수 있습니다.

번역된 본문

다운로드 전에 확인하세요. 내 장비에는 무엇이 맞을까요? 하드웨어를 선택하면 이 페이지가 허깅페이스(Hugging Face)에서 가장 많이 다운로드된 모델들을 스캔하여, shoehorn이 예산(메모리)에 맞출 수 있는 모델들을 찾아줍니다 — 메모리가 감당할 수 있는 품질 기준으로 순위가 매겨집니다. 모든 과정은 브라우저에서 실행됩니다.

내 장비 Mac — 8 GB / Mac — 16 GB / Mac — 24 GB / Mac — 32 GB / Mac — 48 GB / Mac — 64 GB / Mac — 96 GB / Mac — 128 GB GPU — 8 GB VRAM / GPU — 12 GB VRAM / GPU — 16 GB VRAM / GPU — 24 GB VRAM / GPU — 32 GB VRAM / 사용자 지정…

대화 컨텍스트 용량 4k 토큰 — 짧은 채팅 8k 토큰 — 일상적 사용 16k 토큰 — 긴 문서 32k 토큰 — 중편 소설 전체

사용 가능한 GPU 메모리(GiB) 모델 찾기

Shoehorn 설치 shoehorn은 추론 백엔드로 llama.cpp가 PATH에 설치되어 있어야 합니다(Homebrew 설치 시 자동으로 함께 설치됩니다). 그 다음 'shoehorn ui'를 실행하면 로컬 앱이 열립니다 — 모델을 고르고, 버튼 한 번 누르면 채팅이 시작됩니다.

brew install notactuallytreyanastasio/shoehorn/shoehorn (복사)

macOS Apple Silicon: 다운로드 → Linux x86-64 · NVIDIA 또는 AMD: 다운로드 → Windows x86-64 · NVIDIA: 다운로드 →

또는 소스에서 설치: 저장소를 클론한 후 'cargo install --path .' 실행. 전체 릴리스 보기.

이 앱 버튼 하나, 나의 전체 예산 로컬 웹 앱은 사용자의 장비 사양을 측정하고, 적합한 모델을 스트리밍으로 보여주며, 예산을 줄자(tape measure) 형태로 시각화하고, 양자화로 인한 품질 손실을 퍼플렉시티(perplexity) 수치로 표시한 뒤, 마지막에 채팅 버튼으로 마무리됩니다.

원문 보기
원문 보기 (영어)
Before you download What fits your machine? Pick your hardware and this page scans Hugging Face's most-downloaded models for ones shoehorn can fit to your budget — ranked by the quality your memory affords. Runs entirely in your browser. Your machine Mac — 8 GB Mac — 16 GB Mac — 24 GB Mac — 32 GB Mac — 48 GB Mac — 64 GB Mac — 96 GB Mac — 128 GB GPU — 8 GB VRAM GPU — 12 GB VRAM GPU — 16 GB VRAM GPU — 24 GB VRAM GPU — 32 GB VRAM Custom… Conversation room 4k tokens — short chats 8k tokens — everyday use 16k tokens — long documents 32k tokens — the whole novella Usable GPU memory, GiB Find models Get shoehorn Install shoehorn needs llama.cpp on your PATH as the inference backend (the Homebrew install pulls it in for you). Then shoehorn ui opens the local app — pick a model, press one button, chat. brew install notactuallytreyanastasio/shoehorn/shoehorn Copy macOS Apple Silicon Download → Linux x86-64 · NVIDIA or AMD Download → Windows x86-64 · NVIDIA Download → Or from source: cargo install --path . after cloning the repo . All releases . The app One button, your whole budget The local web app measures your machine, streams the fit, renders the budget as a tape measure, puts a perplexity number on what the fit cost, and ends at a Chat button.