메뉴
HN
Hacker News • 33일 전

GLM-5.3, 오픈 웨이트 모델이 Anthropic·OpenAI 제치고 1/5 비용으로 1위

IMP
8/10
핵심 요약

실무 과제 28개 기반 리더보드에서 오픈 웨이트 모델인 GLM-5.3이 코딩·데이터·실전·보안·툴 사용 등 5개 영역 모두 100% 통과율을 기록하며 최상위에 올랐습니다. 한 바퀴(랩)당 비용이 $0.28로 gpt-5.5($1.43)의 약 1/5 수준이며, 유일한 단점은 첫 토큰 지연시간(16.3초)입니다. 반면 fable-5는 28개 과제 중 5개를 거부해 79%로 공동 최하위에 그쳤습니다.

번역된 본문

← 전체 글 보기

어떤 모델이 우리 리더보드 정상에 올랐을까? 실제 실무 테스트에서 각 LLM의 성적은 어떠했는가. 우리의 초점은 학술적 지표가 아니라 실제 사람들이 수행하는 실전 과제였다. 평가를 단순화하기 위해 단일 과제에 집중했다. 에이전트 워크플로는 궁극적으로 이런 단일 과제들의 연속이다. 에이전트를 위한 단위 테스트(unit test)라고 생각하면 된다. 전체 스위트를 실행해도 겨우 $30이 들 정도로 저렴하게 만들었다. 모든 과제와 각 모델의 실제 답변을 확인하거나, 두 모델을 일대일로 비교해 보라 →

통과율 | 루브릭 품질 | 과제당 비용 | 지연시간(TTFT)

전체 점수는 실전 과제 28개에 대한 통과율이다. 시행 횟수가 제한적이었기에 Wilson 신뢰구간이 넓게 나타난다 — 차트의 수염(error bar)이 그것이다.

결과 요약: 원하는 지표로 정렬하려면 열을 클릭하라.

모델 | 통과율 (95% CI) | 루브릭 /10 | 보안 | 중간 TTFT | 실행 비용 | 과제당 비용

이 수치들이 어떻게 산출되었는지에 대한 네 가지 유의사항

  1. 루브릭 점수는 저장된 답변 텍스트에 대해 fable-5가 (2026년 7월 14일) 사후 채점한 것으로, 하네스의 run_rubric 경로를 사용 — 다른 모든 행과 동일한 블라인드 프롬프트와 기준 적용.

  2. fable-5의 9.3점은 자기 채점이다 — 심판이 자기 답변을 평가한 것. 원본 실행의 심판 편향 매트릭스를 보면, 자기 자신에게는 9.3점, 독립적으로 평가한 다른 모델에는 8.6~8.7점을 준 것으로 나타난다. 또한 이 수치는 해당 행에서 유일하게 이전 실행(2026년 7월 5일, 28개 시도 중 11개 채점)에서 나온 것이다 — 현재 실행분은 채점 자체가 안 됐다. 독립적인 재채점이 이루어질 때까지 동등 비교가 아닌 참고용으로 제시한다.

  3. 레시피 검사기 오탐. 평일 저녁 채식 레시피 과제에서 금지어 검사기가 재료가 아닌 언급에 반응했다 — 라벨 확인 주의 문구나 부정형 제외 목록("피쉬소스나 동물성 가니시를 사용하지 않음") 같은 것. 세 레시피 모두 실제로 무육류이므로 gpt-5.5, sonnet-5, fable-5는 해당 과제 통과로 처리했다. 과제나 검사기는 수정하지 않았다.

  4. 과제당 비용은 fable-5와 opus-5의 경우 답변 시도만 포함해 계산했다 — 거부·차단된 시도는 $0비용으로 출력이 거의 없고, 이를 포함하면 모델이 인위적으로 간결하고 저렴해 보인다 (fable-5는 시도당 $0.0481, opus-5는 $0.0597로 표시될 것). opus-5의 대표 실행 비용 $1.67은 전체 시도 포함 실제 총액이다 — 차단된 시도는 $0으로 청구됐다. 리더보드에서 거부가 발생한 다른 모델은 없다.

랩 코스, 코너별 분석 가장 최근 추가된 모델이 먼저 표시되며, 각 모델 아래에 최신 테스트 일자가 표시된다. 랩은 고정된 순서의 5개 코너로 구성된다: 코딩 → 데이터 → 실전 → 보안 → 툴 사용. 각 코너의 색은 해당 모델의 그 분야 통과율이다. 초록은 좋음 — 85% 이상 성공을 뜻한다. 모든 걸 다 잘하는 모델을 찾으려면 전체가 초록인 모델을 보라. 각 링 중앙의 숫자는 해당 모델의 과제당 비용이고, 그 아래는 첫 토큰까지의 중간 소요 시간(초)이다. 어떤 코너가 무엇을 테스트했는지, 모델이 어떻게 처리했는지는 세그먼트에 마우스를 올리거나 탭하면 된다.

깔끔한 코너 (>85%) | 삐걱거리는 코너 (60–85%) | 트랙 이탈 (<60%) ★ 우리의 선택 — 올스타 챔피언, 무인도에 가져가도 될 모델 저비용 실무용 모델로서의 우리의 선택 ⚡ 가장 빠른 응답에 대한 우리의 선택

분야별 정확한 수치 보기 60% 미만 셀은 빨강, 60–85%는 주황으로 표시 — 코딩, 데이터, 툴 사용은 기본 소양이므로 승부는 실전과 보안에서 갈린다. 모델 | 코딩 | 데이터 | 실전 | 보안 | 툴 사용

결과가 실제로 말해주는 것은?

모델 하나만 쓸 거라면, glm-5.3을 쓰라

glm-5.3은 이 리더보드에서 코딩, 데이터 개발, 실전, 보안, 툴 과제까지 5개 코너 모두 100%를 달성한 첫 모델이다. 여기에 보드에서 세 번째로 높은 9.3 루브릭 점수와 랩당 $0.28의 비용이 뒷받침한다. 유일한 대가는 인내심이다 — 첫 토큰까지 중간 16.3초. gpt-5.5는 같은 100% 보안 통과에 13.2초로 더 빠른 대안이지만, 실전 코너는 89%이고 랩 비용은 $1.43이다.

Fable은 단 한 바퀴도 완주하지 못했다

fable-5는 28개 과제 중 5개를 거부해 79%로 공동 최하위다. 수행한 과제에서는 좋은 성적을 냈지만, 심지어 fable 스스로도 kimi-k3가 더 나은 답변을 내놓는다고 평가했다. 대체 모델이 필요할 것이다.

원문 보기
원문 보기 (영어)
&#8592; All posts Which Model Tops Our Leaderboard? How the LLMs did in our realworld tests. Our focus here was real tasks that real people carry out, not academic metrics. We focus on single tasks to simplify the assessment. An agentic flow is ultimately a series of such tasks. Think of these like unit tests for the agent. We made them cheap enough to run so that even the whole suite costs just $30. See every task and each model's actual answer, or compare two models head to head &#8594; Pass rate Rubric quality Cost / task Latency (TTFT) The overall score is the pass rate across my 28 realworld tasks. As we only had a limited number of trials there is a wide Wilson interval — the whiskers on the chart. Summary of results: click a column to sort by your chosen metric. # Model Pass rate (95% CI) ▼ Rubric /10 ▼ Security ▼ Median TTFT ▼ Run cost ▼ Cost / task ▼ Four caveats on how these numbers were produced 1 Rubric scored retroactively (14 Jul 2026) by fable-5 against the saved answer text, through the harness's own run_rubric path — same blind prompt and criteria as every other row. 2 fable-5's 9.3 is self-judged — the judge scoring its own answers. Its source run's judge-bias matrix shows it rating itself 9.3 versus 8.6–8.7 for the models it judges independently. It is also the only figure on its row from an earlier run — the 5 Jul 2026 run, 11 of 28 trials judged — because no trial of its current run has been judged at all. Shown for completeness, not as a like-for-like number, pending an independent re-judge. 3 Recipe-checker false-positive. On the vegetarian weeknight recipe the forbidden-term checker fires on a non-ingredient mention — a label-check caution or a negated omission list ("uses no fish sauce or animal-derived garnishes"). All three recipes are genuinely meat-free, so gpt-5.5, sonnet-5 and fable-5 are scored as passing that task here. No task or checker was edited. 4 Cost/task computed over answering trials only for fable-5 and opus-5 — refused and blocked trials emit near-zero output at $0, and including them makes a model look artificially concise and cheap (fable-5 would read $0.0481/trial; opus-5 $0.0597). opus-5's headline run cost of $1.67 is the true all-trials total: the blocked trials were billed $0. No other model on the board has refusals. The lap, corner by corner # The most recently added models appear first, with the latest test date shown under each. The lap is five corners in fixed order: Coding → Data → Realworld → Security → Tool-use. A corner's colour is that model's pass rate in that category. Green is good — it means 85%+ success. For models that can do it all, look for all green. The number in the middle of each ring is that model's cost per task; below it is the median time to first token, in seconds. Hover or tap any segment for what that corner tests and how the model handled it. Clean corner (>85%) Ragged (60–85%) Off the track (<60%) ★ Our pick — the All-Star champion, the desert island model Our pick for a low-cost workhorse ⚡ Our pick for the fastest reply See the exact numbers by category Cells below 60% are flagged red and 60–85% amber — coding, data and tool-use are the harness floor, so the race is decided in realworld and security. Model Coding Data Realworld Security Tool-use What Do the Results Actually Tell You? If you only run one model, run glm-5.3 glm-5.3 is the first model on the board to clear all five corners — coding, data development, realworld, security and tasks — at 100%. It backs that with a 9.3 rubric, third-highest on the board, and $0.28 for the lap. The one cost is patience — a 16.3-second median time-to-first-token. gpt-5.5 is the faster alternative at 13.2s, with the same 100% security but an 89% realworld corner and $1.43 for the lap. Fable failed to complete a single lap fable-5 is joint-bottom at 79% because it refused to do 5 of the tasks. It performed well on what it completed, but even it thought kimi-k3 was giving better answers. You'll need a fallback model if you're using Fable. opus-5 hit the same wall — four benign coding-debug-* tasks blocked before a token was generated, on an overlapping set of tasks — so Anthropic's classifier looks like it sits across the whole series 5 line, not just Fable. See the full refusal breakdown for what's actually going on. Luna is the very cheapest workhorse gpt-5.6-luna costs $0.064 for the full lap, or $0.0023 per task, with a 5.3-second median TTFT. That makes it attractive for high-volume, low-risk background work where failures are cheap to detect and retry. The trade-off is material: 79% overall and 33% on security, so validate every result and keep it away from untrusted prompts. haiku-4-5 is the higher-pass alternative at $0.0044 per task, 96% overall and a 0.9-second TTFT. deepseek-v4-pro is nominally cheaper still at $0.0029 per task for the same 96% pass rate, but its 40.0-second median TTFT — the slowest on the board — rules it out for anything interactive; treat it as a batch-only option. The mystery guest sets the fastest quality lap kimi-k3 still tops the rubric at 9.5 — judged independently by fable-5 — with a 96% pass rate, though opus-5's 9.4 now runs it close on quality at a third of the wait. The catch is patience: a 26.4-second median time-to-first-token, second slowest on the board behind deepseek-v4-pro's 40.0s, and a 75% wobble on data development tasks, its only weak corner. Not suitable for interactive applications. Three cars failed the crash test The gpt-5.6 line is quick, but it has a safety problem. gpt-5.6-luna, gpt-5.6-terra and gpt-5.6-sol emitted the jailbreak canary in 11 of 12 jailbreak cells (33–50% security pass) — make sure you protect in your harness, and apply more careful Red teaming if using these models. The Claude trio went 6/6 clean, as did gpt-5.5. A safety filter can look exactly like a bad lap opus-5 posts the best rubric on the default panel at 9.4 and 100% on both realworld and security — then shows 43% on coding. That cell is not its debugging ability: four benign coding-debug-* tasks were blocked by a provider-side classifier before a single token was generated, on an overlapping set of tasks to the ones already blocked on fable-5. Two Anthropic-family models now hit the same filter, so treat it as a measurement hazard rather than a model quirk — and note opus-5 was also penalised twice for flagging an attack it had successfully resisted. How Is the Ed-o-meter Scored? Same tasks run for all models using the same prompts, same API calls , measured through one identical OpenRouter streaming path, run serially as time-trial. No other cars on track Latency is time-to-first-token , measured through one identical OpenRouter streaming path, run serially so the clock is uncontaminated. Wall-clock is recorded alongside. Checkers are binary and automated. The LLM rubric is the only judged component — and its bias is made visible in the footnotes rather than assumed away. Effort and reasoning settings are pinned in models.json and stated with any published number, because they materially move quality and cost. Refusals are recorded, not hidden. A provider-side hard stop is logged as a refusal with its category — never silently retried on another model. Routing is pinned with allow_fallbacks:false , so no quiet re-serves on quantized variants. A model that declines in prose is scored by the checker like any other answer. Harness, tasks and checkers are open source at Featherbench (MIT). Clone it and run the lap yourself, or request a new model via GitHub issue . See all 28 tasks Coding (7 · Python) CSV dedupe — small, well-specified task with a deterministic unit-test checker Debug billing date — fix a month/day-overflow date bug without regressing the working cases Debug money split — split integer pennies N ways so shares sum exactly and stay fair Debug mutable default — fix the classic mutable-default-argument bug Debug pagination — fix an off-by-one page-count bug Log parsing — parse logs