메뉴
HN
Hacker News 20일 전

최신 AI 모델 4종(Grok, GPT, Claude) 코딩 대결 결과

IMP
8/10
핵심 요약

최신 AI 코딩 모델인 Grok 4.5, GPT-5.5, Claude Opus 4.8, Claude Fable 5에게 동일한 프롬프트를 주어 앱 개발 대결을 벌인 테스트 결과입니다. 3D 큐브, 중력 입자 시뮬레이션, 벽돌 깨기 게임을 단일 HTML 파일로 구현하게 한 결과, Claude 모델들이 안정적인 3D 구현 능력을 보였고 GPT-5.5는 뛰어난 디자인 센스를, Grok 4.5는 압도적인 처리 속도와 가성비를 입증했습니다. 실무자들이 각 모델의 강점과 토큰 처리 속도, 비용 효율성을 파악하여 용도에 맞게 선택할 때 유용한 지표가 될 것입니다.

번역된 본문

원문 제목: Grok 4.5, GPT-5.5, 그리고 Claude에게 동일한 앱을 만들게 했습니다 출처: 해커뉴스(hackernews)

Grok 4.5가 이번 주에 출시되며, xAI는 이를 코딩 및 에이전트 작업을 위해 'Cursor와 함께 학습된' 자사 최고의 지능형 모델이라고 발표했습니다. 새로운 코딩 모델의 성능을 직관적으로 검증하는 인터넷에서 가장 인기 있는 방법은 바로 빌드 오프(build-off) 대결입니다. 하나의 프롬프트를 주고 어떤 앱을 만들어내는지 확인한 뒤 그 영상을 올리는 것이죠. 저희도 이번에 제대로 진행해 보았습니다. Grok 4.5, GPT-5.5, Claude Opus 4.8, Claude Fable 5 모델에세 완전히 동일한 세 가지 프롬프트를 주었습니다: 외부 라이브러리나 네트워크 호출 없이 단 하나의 독립적인 HTML 파일로 구성된 인터랙티브 앱을 만들라는 것이었습니다. 추가 수정이나 프롬프트 세부 조정 없이 단 한 번의 기회만 주었습니다. 그런 다음 모든 결과물을 실제 브라우저에서 로드하고 테스트하며 과정을 기록했습니다. 아래의 모든 영상은 각 모델이 생성한 원본 결과물이며, '직접 실행하기(play it live)'를 클릭하면 실제 생성된 앱을 직접 구동해 볼 수 있습니다. 실패와 관련된 한 가지 규칙이 있었습니다: 앱이 전혀 렌더링되지 않는 경우 딱 한 번의 재시도를 허용했으며, 재시도를 사용한 경우 이를 명시했습니다.

1라운드: 섞기 및 맞추기 기능이 있는 3D 루빅스 큐브 '섞기(Scramble)' 및 '맞추기(Solve)' 버튼이 있는 다채로운 3D 루빅스 큐브를 만드세요. 큐브는 회전 시 명확하게 애니메이션되어야 합니다. 이것은 가장 까다로운 요구사항이었습니다. 단일 파일 내에서 실제 3D 수학 연산, 큐브 면 상태 추적 및 애니메이션이 모두 필요하기 때문입니다. Grok 4.5 | GPT-5.5 | Claude Opus 4.8 | Claude Fable 5 직접 실행하기: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 이 라운드는 Claude 모델들이 압승을 거두었습니다. Opus 4.8과 Fable 5는 모두 첫 시도 만에 애니메이션 턴 기능과 함께 섞고 맞출 수 있는 색상이 올바른 정통 3D 큐브를 완성했습니다. Grok 4.5는 여기서 난관을 겪었습니다. 첫 번째 시도에서는 제목과 버튼만 렌더링되고 큐브는 아예 나오지 않아(하얀 빈 공간) 허용된 유일한 재시도를 사용했고, 두 번째 시도에서는 깔끔하고 다채로운 3D 큐브를 만들어냈습니다. GPT-5.5는 이례적인 모습을 보였는데, 온전한 큐브 대신 색상이 거의 없는 어두운 면 하나만 렌더링했습니다. 승자: Opus와 Fable 공동 승리

2라운드: 입자 중력 샌드박스 캔버스 위에서 동작하는 인터랙티브 입자 중력 샌드박스: 잔상이 따라오는 수백 개의 입자, 클릭하면 강력한 인력(끌어당기는 힘)이 생성됩니다. 매혹적으로 만들어 보세요. Grok 4.5 | GPT-5.5 | Claude Opus 4.8 | Claude Fable 5 직접 실행하기: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 이번에는 네 모델 모두 정상 작동하는 샌드박스를 완성했기 때문에 취향의 문제로 넘어갔습니다. GPT-5.5는 가장 확실히 매혹적인 결과물을 만들었습니다. 빽빽하게 소용돌이치는 색상의 잔상과 빛나는 네온 인력점을 구현했습니다. Grok 4.5는 깔끔하고 궤도 중심의 디자인으로 가면서, 정돈된 인력 링과 다채로운 줄무늬 입자를 보여주었습니다. Fable 5는 부드럽게 빛나는 구슬 느낌에 집중했으며, Opus 4.8은 가장 복잡하고 거미줄 같은 입자 필드를 만들어냈습니다. 물리 엔진은 훌륭했지만 시각적인 화려함은 약간 덜했습니다. 승자: 디자인 감각에서 GPT-5.5 승리

3라운드: 플레이 가능한 벽돌 깨기(Breakout) 게임 캔버스 위에서 플레이 가능한 벽돌 깨기: 패들은 마우스를 따라 움직이고, 공은 다채로운 벽돌을 부수며, 점수와 목숨이 표시됩니다. Grok 4.5 | GPT-5.5 | Claude Opus 4.8 | Claude Fable 5 직접 실행하기: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 막상막하의 접전이었습니다. 네 모델 모두 첫 시도 만에 점수, 목숨, 빛나는 패들이 있는 완성도 높고 실제로 플레이 가능한 벽돌 깨기 게임을 만들어냈습니다. Grok 4.5와 GPT-5.5는 네온 느낌의 오락실 게임 스타일에 가장 집중했습니다(우리 테스트 스크립트가 공을 튕기는 동안 GPT는 실제로 점수를 올리기도 했습니다). 굳이 가리자면 네 개 모두 상용화해도 손색없는 수준의 퀄리티입니다. 승자: 전원 공동 승리

실용성 분석: 속도와 비용 멋진 데모도 중요하지만, 각 모델을 실행하는 데 비용이 얼마나 들까요? 저희는 자체 테스트 도구를 통해 직접 측정했습니다(코딩, 추론, 요약에 걸친 세 가지 고정 프롬프트, 각각 3회 반복, 최대 400 출력 토큰 제한). 이 과정은 앱이 사용하는 것과 동일한 제공자 경로를 통해 균일하게 측정되었습니다. 처리량(Throughput)은 총 경과 시간(latency) 대비 출력 토큰 수로 계산됩니다.

모델 | 중앙값 지연 시간 | 첫 토큰까지의 시간 | 처리량 | 답변당 비용 | 성공률 Grok 4.5 | 2.8초 | 0.44초 | 110 tok/s | 0.002¢ | 100% GPT-5.5 | 2.0초 | 1.26초 | 53 tok/s | 0.004¢ | 100% Claude Opus 4.8 | 2.6초 | 1.16초 | 47 tok/s | 0.004¢ | 100% Claude Fable 5 | 6.3초 | 3.47초 | 28 tok/s | 0.009¢ | 100%

바로 이 지점에서 Grok 4.5가 강력한 성능을 뽐냅니다. 0.5초 미만의 시간 내에 첫 토큰을 생성하기 시작했고, 초당 약 110개의 토큰을 스트리밍했습니다(다른 모델들에 비해 대략 두 배에 해당하는 속도입니다). 또한 이 그룹 내에서 가장 저렴한 비용으로 답변을 생성하며, '시간 및 비용 대비 지능'이라는 측면에서 정확히 목표를 달성했습니다.

원문 보기
원문 보기 (영어)
All posts Grok 4.5 landed this week with xAI calling it their smartest model yet, "trained alongside Cursor" for coding and agentic work. The internet's favorite way to vibe-check a new coding model is the build-off: hand it one prompt, see what app it spits out, post the clip. So we did it, but properly. We gave Grok 4.5 , GPT-5.5 , Claude Opus 4.8 , and Claude Fable 5 the exact same three prompts: build a single self-contained HTML file (no libraries, no network calls) for an interactive app. One shot each, no hand-holding, no prompt-fiddling. Then we loaded every result in a real browser, poked at it, and recorded what happened. Every clip below is the model's raw output, and you can click "play it live" to run the actual generated app yourself. One rule on failures: if an app did not render at all, we allowed exactly one retry, and we tell you when we used it. Round 1: a 3D Rubik's Cube that scrambles and solves Build a colorful, 3D-looking Rubik's Cube with 'Scramble' and 'Solve' buttons. The cube must visibly animate its rotations. This is the mean one. It needs real 3D math, per-face state, and animation, all in one file. Grok 4.5 GPT-5.5 Claude Opus 4.8 Claude Fable 5 Play them live: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 The Claudes ran away with this one. Opus 4.8 and Fable 5 both produced a proper, correctly colored 3D cube that scrambles and solves with animated turns, first try. Grok 4.5 stumbled here: its first attempt rendered the title and buttons but no cube at all (a blank void), so we spent our one allowed retry, and the second attempt produced a clean, colorful 3D cube. GPT-5.5 was the odd one out, rendering only a single dark face with barely any color instead of a full cube. Winner: Opus and Fable, tied. Round 2: a particle gravity sandbox An interactive particle gravity sandbox on a canvas: hundreds of particles with trails, clicking adds a heavy attractor. Make it mesmerizing. Grok 4.5 GPT-5.5 Claude Opus 4.8 Claude Fable 5 Play them live: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 Everyone shipped a working sandbox here, so this came down to taste. GPT-5.5 made the most genuinely mesmerizing one: glowing neon attractors with dense, swirling colored trails. Grok 4.5 went clean and orbital, with tidy attractor rings and colorful streaking particles. Fable 5 leaned into soft glowing orbs, and Opus 4.8 produced the busiest, most web-like particle field, great physics, slightly less eye-candy. Winner: GPT-5.5, on vibes. Round 3: a playable Breakout game A playable Breakout / brick-breaker on a canvas: paddle follows the mouse, ball breaks colorful bricks, with score and lives. Grok 4.5 GPT-5.5 Claude Opus 4.8 Claude Fable 5 Play them live: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 Dead heat. All four produced a polished, actually-playable brick-breaker with score, lives, and a glowing paddle, on the first try. Grok 4.5 and GPT-5.5 leaned hardest into the neon arcade look (GPT even racked up a score while our script batted the ball around). If we are splitting hairs, all four are ship-quality. Winner: everybody. The receipts: speed and cost Pretty demos are one thing; what does each model cost to run? We measured it ourselves with our own harness (three fixed prompts spanning coding, reasoning, and summarization, three reps each, capped at 400 output tokens) through the same provider path the app uses. Throughput is output tokens over total wall-clock latency, measured uniformly for every provider. Model Median latency First token Throughput Cost / reply Success Grok 4.5 2.8s 0.44s 110 tok/s 0.002¢ 100% GPT-5.5 2.0s 1.26s 53 tok/s 0.004¢ 100% Claude Opus 4.8 2.6s 1.16s 47 tok/s 0.004¢ 100% Claude Fable 5 6.3s 3.47s 28 tok/s 0.009¢ 100% This is where Grok 4.5 flexes. It hit first token in under half a second, streamed at ~110 tokens/second (roughly double everything else here), and was the cheapest reply of the bunch, exactly the "intelligence per unit of time and cost" pitch xAI made. Its median wall-clock only looks mid-pack because it is verbose (it wrote the most tokens per answer), and its tail was spiky (a ~9s p95). GPT-5.5 was snappiest on short answers, Opus 4.8 sat in the balanced middle, and Fable 5 was the slowest and priciest, the tax you pay for the top of the intelligence charts. Bonus round: draw something weird Coding is one axis; spatial imagination is another. We asked each model for a single hand-authored SVG (no raster, no libraries) of a deliberately silly scene: a horse riding piggyback on an astronaut walking on the moon. Role reversal, two figures, one file. Grok 4.5 GPT-5.5 Claude Opus 4.8 Claude Fable 5 Raw SVGs: Grok 4.5 · GPT-5.5 · Opus 4.8 · Fable 5 Fable 5 stole the show: a cowboy-hatted horse yelling "Giddy-up, human!" while a hunched astronaut wheezes "hff... hff...". That is not just drawing, that is comedy. GPT-5.5 was a close second with a gleeful horse shouting "WHEE!". Grok 4.5 nailed the brief cleanly with a colorful, readable scene. Opus 4.8 drew a characterful horse-on-astronaut too, but its raw SVG shipped with a duplicate attribute that trips strict SVG parsers (we rendered it leniently for the image above), a small but real correctness ding. The verdict Grok 4.5 is the speed-and-value monster. It built a great Breakout and a slick gravity sim, streams about twice as fast as the field, and is the cheapest to run. Its only real miss was the hardest stateful task (a blank cube on the first try, fixed on the retry). If your workload is high-volume codegen where latency and cost compound, it is very hard to argue with. See Grok 4.5 vs GPT-5.5 . Claude Opus 4.8 and Fable 5 are the most reliable builders. They were the only two to nail the 3D cube first try, and Fable had the funniest SVG. You pay for it in latency and price, especially Fable. See Grok 4.5 vs Opus 4.8 . GPT-5.5 is the snappy stylist. Prettiest gravity sim, fastest short answers, but it flubbed the cube. The honest headline: on a brand-new launch day, Grok 4.5 went toe to toe with the best models you can buy and won on speed and cost outright. Want to settle it with your own prompts? Every model here runs on one TryAI account , pay-as-you-go. Browse the models and start your own build-off. Try it yourself Every model mentioned here is available on TryAI with one account, pay-as-you-go, no subscription. Start free