메뉴
HN
Hacker News • 43일 전

세레브라스, 오픈AI GPT-5.6 Sol 초고속 모드 공개

IMP
8/10
핵심 요약

세레브라스(Cerebras)와 오픈AI(OpenAI)가 초당 최대 750 토큰을 처리하는 '울트라패스트 모드(Ultrafast Mode)'를 선보였습니다. 이를 통해 GPT-5.6 Sol 모델의 품질을 유지하면서도 경쟁 모델 대비 최대 11배 빠른 추론 속도를 자랑하며, 시간이 생명인 금융, 보안, 코딩 등 실무 환경에서 AI 작업의 효율을 극대화할 수 있게 되었습니다.

번역된 본문

2026년 8월 13일 — 세레브라스(Cerebras)와 오픈AI(OpenAI), GPT-5.6 Sol 초고속 모드 가속화 (작성자: Joyce Er)

오늘 세레브라스와 오픈AI는 세레브라스의 기술력을 바탕으로 오픈AI API에서 최초로 출시되는 새로운 서비스 계층인 '울트라패스트 모드(Ultrafast Mode)'를 일찍 공개하게 되었습니다. 울트라패스트 모드는 우선 일부 선별된 고객 그룹에게 제공되며, 점진적으로 액세스 권한이 확대될 예정입니다. 세레브라스가 지원하는 울트라패스트 모드의 GPT-5.6 Sol은 초당 최대 750개의 출력 토큰(Token)을 처리하면서도 품질 저하가 전혀 없습니다. 이를 통해 솔 울트라패스트(Sol Ultrafast)는 시간이 가장 중요하고 미션 크리티컬한 업무를 가속화할 수 있습니다.

전례 없는 속도의 최첨단 인공지능

AI 개발자들은 항상 속도와 지능 사이에서 하나를 선택해야 했습니다. 모델의 크기와 지능이 커질수록 더 많은 컴퓨팅 및 데이터 이동 비용이 발생하여 응답 속도가 느려집니다. 사용자는 고품질의 결과를 기다리거나, 짧은 시간 안에 낮은 수준의 결과를 타협해서 받아들여야 했습니다. GPT-5.6 Sol 울트라패스트 모드는 이러한 트레이드오프를 해결하여 1초가 아쉬운 제품 및 워크플로우에 최첨단 지능을 제공합니다.

인공지능 분석 플랫폼 Artificial Analysis가 보고한 출력 속도와 비교했을 때, 울트라패스트 모드의 GPT-5.6 Sol은 페이블 5(Fable 5)보다 11배 빠르며, 패스트 모드(Fast mode)의 오퍼스 4.8(Opus 4.8)보다 5배 빠릅니다. 세레브라스는 인류의 마지막 시험(Humanity's Last Exam, HLE)에서 인기 있는 모델들과 정면 대결하여 울트라패스트 모드를 테스트했습니다. HLE는 화학, 경제학, 문학 등 여러 분야에서 박사학위를 소지한 사람만 답할 수 있는 2,500개의 질문으로 구성된 까다로운 벤치마크입니다.

평가 결과, 울트라패스트 모드의 GPT-5.6 Sol은 11시간 11분 만에 2,500개의 HLE 질문에 모두 답변했습니다. 반면 클로드 페이블 5(Claude Fable 5)는 동일한 결론에 도달하기 위해 78시간 27분, 즉 3일 이상의 연속 컴퓨팅 시간이 필요했습니다. 다시 말해, 울트라패스트 모드는 단 하루 근무 시간 안에 인류 지식의 최전선을 탐구했으며, 비슷한 정확도를 거의 7배 더 빠르게 달성했습니다.

인류의 마지막 시험(Humanity's Last Exam) 벤치마크: 벤치마크는 세레브라스가 수행했으며, 7월 10일에 Codex 기반 '매우 높은 추론(xhigh reasoning)' 설정의 GPT 5.6 Sol 울트라패스트 모델과 7월 13~15일에 Claude Code 기반 '매우 높은 추론' 설정의 Claude Fable 5를 사용했습니다.

모델의 성능이 계속 발전함에 따라 빠른 추론을 위한 애플리케이션의 범위도 확장되고 있습니다. GPT-5.6 Sol은 법적 변론, 재무 모델 및 엔지니어링 보고서를 작성하는 데 있어 오픈AI의 역대 최고의 모델입니다. 경제적으로 가치 있는 지식 작업을 측정하는 벤치마크인 GDP-Val에서 울트라패스트 모드는 품질 저하 없이 엔드투엔드 속도를 5.6배 향상시켰으며, 이는 더 빠른 추론이 경제적으로 가치 있는 작업을 어떻게 가속화할 수 있는지 보여줍니다.

벤치마크는 세레브라스가 2026년 7월 31일 Codex 내에서 '중간 추론(medium reasoning)' 설정을 사용하여 GPT 5.6 Sol 및 GPT 5.6 Sol 울트라패스트로 수행되었습니다.

고속 인공지능, 고위험업무의 핵심 동력으로

더 빠른 인텔리전스는 개인과 조직에게 가능성의 지평을 넓혀줍니다. 울트라패스트 모드를 통해 이제 여러분은 1초가 급급한 문제의 핵심 경로에 AI 에이전트를 투입할 수 있습니다.

"GPT-5.6 Sol 울트라패스트를 통해 세레브라스는 사용자가 생각하고, 코드를 작성하고, 협업하는 속도를 따라잡는 AI를 구현합니다. 우리는 울트라패스트 추론을 통해 워크플로우와 애플리케이션이 어떻게 혁신되는지 기대하고 있습니다." - Rohan Varma, OpenAI 제품 총괄

울트라패스트 모드는 최첨단 AI를 사용하여 들어오는 정보에 신속하게 대응하려는 조직에 지속적인 경쟁 우위를 제공합니다. 웹 서비스를 운영하는 기업은 울트라패스트 모드를 활용하여 서비스 중단의 근본적인 원인을 파악하고 해결함으로써 고객의 신뢰를 유지하고, 수익 손실을 방지하며, SLA(서비스 수준 계약)에 따른 다운타임을 줄일 수 있습니다. 또한 고위험 사이버 공격 상황에서 울트라패스트는 막대한 피해를 막기 위해 악의적인 위협을 신속하게 탐지하고 대응해야 하는 보안 팀에게 매우 귀중한 도구입니다.

더 나아가 울트라패스트 모드는 에이전트를 활용하는 완전히 새로운 작업 방식을 가능하게 합니다. 에이전트의 성능을 최대한 끌어내기 위해 여러 병렬 세션 간에 컨텍스트를 전환할 필요 없이 실시간 통찰력과 업데이트를 제공합니다.

"과거에는 작업이 완료될 때까지 몇 분을 기다려야 했지만, 이제는 내가 컨텍스트를 전환할 틈도 없이 작업이 완료됩니다. 이 점이 제 생산성을 비약적으로 높여주었습니다." - Jeffrey Wang, OpenAI 연구원

원문 보기
원문 보기 (영어)
Aug 13 2026 Accelerating GPT-5.6 Sol Ultrafast Joyce Er Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode , a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding over time. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise, allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work. Frontier Intelligence at Unprecedented Speed AI builders have always needed to choose between speed and intelligence. As models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Users often need to wait for high-quality results or accept inferior results within a shorter timeframe. GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows where every second matters. Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode. At Cerebras, we put Ultrafast to the test by running it head-to-head with popular models on Humanity's Last Exam. HLE is a challenging model benchmark that consists of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature. In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster. Humanity's Last Exam Benchmark Benchmarking was performed by Cerebras using GPT 5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10 and Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15. As model capabilities continue to advance, the range of applications for fast inference expands. GPT-5.6 Sol is OpenAI’s best model yet for legal briefs, financial models, and engineering reports. On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, showing how faster inference can accelerate economically valuable work. Benchmarking was performed by Cerebras on July 31 2026 using GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium reasoning within Codex. High-Speed Intelligence Powers High-Stakes Work Faster intelligence changes what’s possible for individuals and organizations. With Ultrafast, you can now put agents on the critical path of problems where every second counts. " With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference. " Rohan Varma Product at OpenAI Ultrafast is a persistent edge for organizations using frontier AI to quickly respond to incoming information. Companies operating web services can leverage Ultrafast to root-cause and address production outages, preserving customer trust, preventing lost revenue, and saving downtime minutes against their SLAs. And in adversarial, high stakes cyberattacks, Ultrafast is an invaluable tool for security teams who must quickly detect and respond to bad actors to contain catastrophic losses. More broadly, Ultrafast enables entirely new modes of working with agents, it delivers real-time insights and updates, so you don’t have to context-switch across multiple parallel sessions to get the most out of your agents. " Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive. " Jeffrey Wang OpenAI Researcher With Ultrafast, researchers and engineers can reserve their attention for going deep on select problems that matter most, while continuing to use Standard processing for parallelizing commodity tasks. Cerebras is excited to power the next wave of AI innovation, raising the ceiling for what individuals and organizations can accomplish with responsive AI. Breakneck Speed is Enabled by Breakthrough Innovation GPT-5.6 Sol on Ultrafast mode is powered by Cerebras’ revolutionary Wafer-Scale Engine architecture, purpose-built for frontier AI workloads. Fast frontier inference is a data movement problem: on GPUs, inference on large models is bottlenecked by memory bandwidth, as model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens within a model response. Cerebras takes a contrarian approach to eliminating this inefficient data movement: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This technical approach scales smoothly with model size, paving the way for a continued speed advantage on future frontier models. Ultrafast: Now in Limited Preview GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. Access will expand as capacity grows. Sign up for updates.
관련 소식