메뉴
HN
Hacker News • 22일 전

Qwen 3.8 27B, Cerebras에서 초당 1500토큰으로 제공

IMP
7/10
핵심 요약

Cerebras의 공개 API 엔드포인트에서 Qwen 3.8 27B 모델을 초당 약 1500토큰의 속도로 사용할 수 있게 되었습니다. 무료 체험 및 종량제 요금제에서 이용 가능하며, 컨텍스트는 무료 64k/유료 128k입니다. Cerebras는 가지치기(Pruning) 없이 원본 모델을 제공하며, 저장 시에만 선택적 양자화를 사용해 품질을 보존한다고 강조했습니다.

번역된 본문

Cerebras 공개 엔드포인트의 모델은 무료 체험 및 종량제 요금제에서 이용할 수 있으며, 요율 제한과 가격 정책이 적용됩니다. 추가 모델 패밀리, 예약 용량, 더 높은 처리량 및 프로덕션 SLA가 필요한 경우 전용 엔드포인트(Dedicated Endpoints)를 참고하세요. 처음 이용한다면 퀵스타트를 따라 첫 API 호출을 해보세요. 용도에 맞는 모델 선택은 모델 선택 가이드를 참고하시고, 아래 모델 이름을 클릭하면 전체 사양, 기능, 요금제별 제한을 확인할 수 있습니다.

사용 가능한 모델

  • 모델명: OpenAI GPT OSS / 모델 ID: gpt-oss-120b / 파라미터: 1,200억 / 컨텍스트(무료/유료): 65k / 131k / 속도: 약 3000 tokens/s
  • 모델명: Qwen 3.8 27B / 모델 ID: qwen-3.8-27b / 파라미터: 270억 / 컨텍스트(무료/유료): 64k / 128k / 속도: 약 1500 tokens/s

더 많은 모델이 필요하다면? 다양한 추가 모델 패밀리가 전용 엔드포인트를 통해 제공됩니다.

모델 압축 이 섹션은 플랫폼에서 제공되는 각 모델의 압축 상태에 대한 투명성을 제공합니다. 저희는 커뮤니티의 다양한 오픈소스 모델을 호스팅하며, 현재 공개 엔드포인트에서 가지치기(Pruning)된 모델은 호스팅하지 않습니다. 공개 엔드포인트를 통해 제공되는 모든 모델은 원본, 즉 가지치기되지 않은 버전입니다. REAP(라우터 가중 전문가 활성화 가지치기, Router-weighted Expert Activation Pruning)와 같은 가지치기 기술을 연구하고 있지만, 이러한 모델은 Hugging Face에서 연구 커뮤니티와 공유되며 공유 API를 통해 제공되지 않습니다. REAP에 대한 자세한 내용은 연구 블로그에서 확인할 수 있습니다. 모든 공개 모델은 가지치기되지 않은 원본입니다.

Cerebras는 최대한의 품질 보존을 위해 저장 시에만 선택적 가중치 전용 양자화(weight-only quantization)를 사용합니다. 즉, 가중치는 업계 표준에 따라 16비트/8비트/4비트 혼합 형태로 저장됩니다. 품질을 위해 민감한 레이어는 전체 정밀도로 저장되고 사용 시 즉시 역양자화되어 고정밀 연산이 이루어집니다. 활성화값(activations), 어텐션(attention), KV 캐시는 전체 정밀도를 유지하며 양자화되지 않습니다.

자주 묻는 질문

Q: 사전 통지 없이 모델 아키텍처를 변경하나요? A: 아니요. 저희는 기존 모든 엔드포인트에서 수정 없이 원본 모델을 제공하기로 약속합니다. 호스팅 포트폴리오에서 가지치기를 통해 모델 아키텍처를 변경하지 않습니다. 향후 추가 압축 기법(가지치기 등)을 탐색하더라도, 가지치기 전용 이름이 붙은 별도 엔드포인트로 제공하여 완전한 투명성을 보장하고 사용자가 자신에게 맞는 버전을 선택할 수 있도록 합니다.

Q: REAP 가지치기 모델은 어디서 찾을 수 있나요? A: REAP 가지치기 모델은 연구 및 실험 목적으로 Hugging Face에서 이용할 수 있습니다(Cerebras REAP 컬렉션). 이 모델들은 저희의 가지치기 연구를 보여주지만 프로덕션 API를 통해 제공되지는 않습니다.

Q: 압축, 양자화, 가지치기란 무엇인가요? A: 압축은 모델 크기나 연산 요구량을 줄이는 기법을 총칭하는 용어입니다. 대표적인 압축 기법은 다음과 같습니다:

  • 양자화(Quantization): 모델 가중치를 표현하는 숫자의 정밀도를 낮추는 것(예: FP16을 FP8로 변환). 아키텍처를 변경하지 않고 메모리 사용량을 줄입니다.
  • 가지치기(Pruning): 모델 크기를 줄이기 위해 레이어나 전문가(expert) 등 모델의 일부를 영구적으로 제거하는 것. 모델 아키텍처가 변경되어 사실상 다른 모델이 됩니다.
원문 보기
원문 보기 (영어)
Copy page Copy MCP Server View as Markdown Models on Cerebras public endpoints are available on the free trial and pay-as-you-go tiers, subject to rate limits and pricing . For additional model families, reserved capacity, higher throughput, and production SLAs, see Dedicated Endpoints . New here? Follow the Quickstart to make your first API call. To pick a model by use case, see the model selection guide . Select any model name below for full specs, capabilities, and per-tier limits. ​ Available Models Model Name Model ID Parameters Context (free / paid) Speed (tokens/s) OpenAI GPT OSS gpt-oss-120b 120 billion 65k / 131k ~3000 Qwen 3.8 27B qwen-3.8-27b 27 billion 64k / 128k ~1500 Looking for more models? Many additional model families are available through Dedicated Endpoints . ​ Model Compression This section provides transparency about the compression state of each model available on our platform. We host a variety of open-source models from the community. We do not currently host pruned models on our public endpoints. All models served through our public endpoints are the original, unpruned versions. While we conduct research on pruning techniques like REAP (Router-weighted Expert Activation Pruning), these pruned models are shared with the research community on Hugging Face but are not available through our shared API. You can read more about REAP in our research blog . All of our public models are unpruned. Cerebras uses selective weight-only quantization only during storage to preserve maximal quality. This means that the weights are stored in partial 16-bit / 8-bit / 4-bit, in-line with industry standards. For quality, sensitive layers are stored at full precision with dequantization on the fly, so operations are done in high precision. The activations, attention, and kv cache remain in full precision and unquantized. ​ Frequently Asked Questions Will you change a model's architecture without notice? No. We are committed to serving the original models for all existing endpoints, without modification. We do not alter model architectures via pruning on our hosted portfolio. If we explore additional compression techniques (like pruning) in the future, these would be offered as separate endpoints with pruning-specific names, ensuring complete transparency and allowing you to choose which version best fits your needs. Where can I find your REAP pruned models? Our REAP pruned models are available on Hugging Face for research and experimentation purposes: Cerebras REAP Collection . These models demonstrate our pruning research but are not served through our production API. What are compression, quantization, and pruning? Compression is an umbrella term for techniques that reduce model size or computational requirements. Common compression techniques include: Quantization : Reducing the precision of numbers used to represent model weights (e.g., converting from FP16 to FP8). This reduces memory usage without changing the model’s architecture. Pruning : Permanently removing parts of a model, like layers or experts, to reduce model size. This changes the model’s architecture and creates a different model. Was this page helpful? Yes No ⌘ I
관련 소식