메뉴
HN
Hacker News 7일 전

최신 AI 비전 모델 4종, '모나리자' 그려보기 결과

IMP
7/10
핵심 요약

GPT-5.6, Claude Fable 5, Grok 4.5, Gemini 3.6 Flash 모델에 색연필 도구를 제공해 그림을 그리게 하는 테스트를 진행했습니다. AI가 스스로 붓질, 혼합, 지우기 등을 반복하며 목표 이미지를 모방하는 과정은 복잡한 에이전트 작업 수행 능력과 실제 비용을 시각적으로 보여줍니다. 단순 벤치마크 점수를 넘어, 최신 폐쇄형 모델들이 실제 도구를 활용해 장기적인 작업을 수행할 때 어떤 성능 차이를 보이는지 확인할 수 있는 중요한 실험입니다.

번역된 본문

우리는 '그림 그리기 경기장(drawing arena)'을 구축했습니다. 모델에게 하얀 빈 캔버스와 색연필 도구 세트를 쥐여주고 방해하지 않고 관망합니다. 모델은 색상, 붓 끝 너비, 압력을 설정하고, 여러 묶음의 선을 그리고, 섞기 위해 번지게 한 뒤(smudge), 지우고, view_canvas를 호출하여 자신의 작업 결과를 보고 무엇을 수정할지 결정합니다. 모델은 타겟 이미지를 재현하거나 텍스트 프롬프트를 기반으로 그림을 그립니다.

우리는 4개의 비전 모델(GPT-5.6 Sol, Claude Fable 5, Grok 4.5, Gemini 3.6 Flash)을 대상으로 두 개의 타겟(모나리자와 반 고흐의 별이 빛나는 밤, 둘 다 객관적으로 점수가 매겨짐)과 5개의 개방형 프롬프트를 테스트하여 총 28개의 그림을 그리게 했습니다. 도구 사용, 비용, 결과물, 그리고 모델들이 실제로 작업을 개선했는지에 대해 다룰 것이며, 마지막에는 우리의 의견을 밝힐 것입니다.

왜 이 테스트를 하는가 우리가 이 테스트를 왜 하는지에 대한 간단한 설명입니다. 왜냐하면 지난번 테스트(모델이 뮤직비디오를 만드는 것)에서 모델의 창의성이나 예술적 능력을 홍보하는 것으로 읽는 사람들이 있어 좋은 토론이 촉발되었기 때문입니다. 이것들은 모델의 능력을 보여주는 객관적인 테스트가 아닙니다. 의도적으로 개방형이고 모호한 작업들입니다.

개인적으로 이 테스트를 계속 진행하는 이유는 다음과 같습니다:

재미있고 진정으로 유익합니다. 모델이 느슨하고 개방된 작업을 해결하는 것을 지켜보는 것은 또 다른 벤치마크 숫자보다 더 흥미로운 능력의 시각적 지표입니다. 최신 프론티어 모델과 오픈소스 모델 간의 격차를 보여줍니다. 저렴한 오픈웨이트(Open-weight) 모델은 많은 실행 작업에서 최신 프론티어 모델을 대체할 수 있으며, 우리는 계속해서 이에 대해 보고할 것입니다. 하지만 세상에는 여전히 벤치마크 점수에만 집착하는(benchmaxxing) 경우가 많으며, 이와 같은 작업은 그 허상을 깨부숴 줍니다. 이 테스트는 프론티어 모델의 능력을 나머지 모델들과 진정으로 구분합니다. 곧 보시겠지만, Grok 4.5는 이 '기본적인' 그리기 작업에서 매우 서툴렀고, 우리가 시도한 오픈웨이트 모델들은 아예 사용할 수도 없었습니다. (일부는 그저 빈 캔버스만 반환했습니다. 완전히 오픈소스화되면 보고할 Kimi K3에서는 이것이 바뀔 수도 있습니다.) 오래 걸리는 작업의 실제 비용을 보여줍니다. Claude Fable 5는 거의 항상 다른 모델들보다 훨씬 더 오래 걸리고 훨씬 더 많은 돈이 들었으며, 여기서는 더 나쁜 결과물을 생성했습니다. Fable은 뮤직비디오 챌린지를 포함한 여러 작업에서 우리에게 GPT-5.6을 이겼지만, 이번 작업에서는 더 많은 오버헤드를 들이고도 더 나쁜 결과를 냈습니다. 사용 사례에 따라 이러한 트레이드오프는 중요합니다. 토론을 사랑합니다. 토론은 진심으로 흥미로웠으며, 설정이나 새로운 작업에 대한 제안이 있다면 반영하겠습니다. 이 테스트들을 실행하는 것은 정말 즐거웠고, 결과물을 보는 것은 매우 신났습니다.

도구들 모든 모델은 완전히 동일한 색연필 도구 세트로 작업했습니다:

plan: 단계 사이에 생각하고 계획하기 위한 작업 없는 스크래치패드(메모장). view_target: 타겟 이미지를 다시 봅니다. view_canvas: 현재 페이지를 렌더링하고 봅니다. 이때 타겟에 대한 그림의 점수도 매겨집니다. set_color / set_brush / set_pressure: 현재 연필 색상, 끝 너비 및 압력(0에서 1 사이의 불투명도. 낮은 압력은 더 부드러운 톤을 위함). draw: 한 번의 호출로 여러 마크(선, 획, 모양 윤곽선, 점)를 그립니다. 단색 칠하기 기능은 없으며, 실제 연필처럼 덧칠하여 톤과 색상을 만듭니다. smudge: 경직 도구(blending stump)처럼 사각형 영역을 혼합/부드럽게 합니다. erase / clear_canvas: 영역을 다시 하얗게 되돌리거나 / 전체 페이지를 리셋합니다. 전체 테스트 하네스는 github.com/hershalb/canvas-arena 에 오픈소스로 공개되어 있습니다. 원하는 타겟 이미지나 텍스트 프롬프트를 가리키고 직접 실행해 보세요.

두 가지 타겟 재현 모델은 내내 타겟을 볼 수 있었고, 타겟과의 구조적 유사도(SSIM)를 기준으로 점수가 매겨졌습니다. SSIM 점수는 아래에 표시됩니다.

모나리자 (타겟) GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash 전체 기록(Transcripts): GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

별이 빛나는 밤 (타겟) GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash 전체 기록(Transcripts): GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash

다섯 가지 프롬프트 그림 타겟 이미지도, 점수도 없습니다. 오직 텍스트 간단 설명과 빈 페이지뿐입니다. "늙은 어부의 주름진 얼굴, 따뜻한 늦은 오후의 햇살, 깊은 주름과 덥수룩한 수염" GPT-5.6 Sol Claude Fabl

원문 보기
원문 보기 (영어)
All posts We built a drawing arena: hand a model a blank white canvas and a set of colored-pencil tools, then get out of the way. The model sets a color, tip width, and pressure, lays down batches of strokes, smudges to blend, erases, and calls view_canvas to see its own work and decide what to fix. It either reproduces a target image or draws from a text prompt. We ran four vision models, GPT-5.6 Sol , Claude Fable 5 , Grok 4.5 , and Gemini 3.6 Flash , across two targets (the Mona Lisa and Van Gogh's Starry Night, both scored objectively) and five open-ended prompts, for 28 drawings total. We'll cover tool use, cost, output, and whether the models actually improved their work, with our opinion at the end. Why we run these A quick note on why we're doing this, because our last one of these (models making music videos) kicked off a good discussion where some folks read it as us promoting the creativity or artistic ability of models. These are not objective tests of model capability. They are deliberately open-ended, fuzzy tasks. Here is why we personally keep running them: They are fun, and genuinely informative. Watching a model tackle a loose, open-ended task is a more interesting visual indicator of capability than yet another benchmark number. They expose the real frontier-vs-open gap. Cheaper open-weight models can absolutely replace the frontier for a lot of execution work, and we will keep reporting on that. But there is also a lot of benchmaxxing out there, and tasks like this cut through it. They genuinely separate frontier capability from the rest. Grok 4.5, as you will see, was rough at this "basic" drawing task, and the open-weight models we tried were not even usable, several just returned a blank canvas. (That may change with Kimi K3 which we will report on once it is fully open-sourced) They show what long-running tasks actually cost. Claude Fable 5 almost always took much longer than the others, for far more money, and here produced worse output. Fable has beaten GPT-5.6 for us on plenty of tasks, including the music-video challenge , but on this one it was worse for a lot more overhead. Depending on your use case, that trade-off matters. We're loving the discussions. The discussions have genuinely been interesting, and if you have suggestions for the setup or new tasks, we will take them into account. These have been a blast to run, and the outputs are super exciting to see. The tools Every model worked with the exact same colored-pencil toolset: plan : a no-op scratchpad for thinking and planning between steps. view_target : look at the target image again. view_canvas : render the current page and see it. This is also when we score the drawing against the target. set_color / set_brush / set_pressure : the current pencil color, tip width, and pressure (0 to 1 opacity, low pressure for softer tone). draw : lay down a batch of marks (strokes, lines, shape outlines, dots) in one call. There are no solid fills, tone and color are built by layering, like a real pencil. smudge : blend/soften a rectangular region, like a blending stump. erase / clear_canvas : lift a region back to white / reset the whole page. The whole harness is open source at github.com/hershalb/canvas-arena . Point it at any target image or a text prompt and run it yourself. The two target reproductions The models could see the target the whole time and were scored on structural similarity (SSIM) to it. We show the SSIM scores below. Mona Lisa (target) GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash Starry Night (target) GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash The five prompt drawings No target, no score, just a text brief and a blank page. "An elderly fisherman's weathered face, warm late-afternoon sun, deep wrinkles and stubble" GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash "A beautiful sunset over a calm ocean, orange-to-violet sky, silhouetted horizon" GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash "A single red rose in a glass vase against a dark background, dramatic Rembrandt lighting" GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash "A tabby cat curled asleep on a windowsill in afternoon sun" GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash "A cozy cabin interior with a lit fireplace, warm amber light and soft shadows" GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash Full transcripts: GPT-5.6 Sol · Claude Fable 5 · Grok 4.5 · Gemini 3.6 Flash The headline numbers Averages and totals across all seven drawings per model. Model Drawings ↕ Avg steps ↕ Avg time ↕ Total tokens ↕ Est. token cost ↕ Avg self-reviews ↕ Total draw calls ↕ GPT-5.6 Sol 7 29 6.2 min 3.4M $7.74 6.0 116 Claude Fable 5 7 54 12.5 min 14.6M $160.58 15.6 207 Grok 4.5 7 99 4.8 min 34.0M $9.21 4.3 354 Gemini 3.6 Flash 7 73 6.9 min 27.7M $12.87 22.6 190 Tool calling The four models used the exact same toolset in completely different ways. GPT-5.6 Sol Claude Fable 5 Grok 4.5 Gemini 3.6 Flash GPT-5.6 Sol never called set_color, set_brush, or set_pressure once , it set those inline on each draw call instead, so its calls are almost all draw, smudge, and view_canvas. Grok 4.5 did the opposite : 65% of its 1,349 tool calls were set_color / set_brush / set_pressure, which is why it averaged 99 steps per drawing. Claude Fable 5 leaned on smudge (123 calls) and reviewed constantly. Gemini 3.6 Flash was the most obsessive reviewer of all, nearly a third of its calls were view_canvas (about 23 self-reviews per drawing), and like Sol it never touched the set_* tools. Cost and tokens Claude Fable 5 cost an estimated $160 for its seven drawings, roughly 20x GPT-5.6 Sol ($7.74), Grok 4.5 ($9.21), or Gemini 3.6 Flash ($12.87). This shouldn't really be a surprise by now. Grok used by far the most tokens (34M) but stayed cheap because ~98% were cached reads at a fraction of the input rate, and Grok's rates are the lowest of the four. Gemini also leaned heavily on cached reads, so even at 27.7M tokens it stayed modest. Did the models actually improve? Every view_canvas re-scored the canvas against the target, so we can watch progress over a run. Two patterns held across both targets: First, they plateau early. Claude Fable 5 reviewed its Mona Lisa 27 times, but its similarity was essentially flat after about the fifth review. Second, in all eight target runs, the final drawing scored below the best the model reached mid-run. GPT-5.6 Sol's Mona Lisa peaked at 0.352 SSIM and ended at 0.325; its Starry Night peaked at 0.179 and ended at 0.130. Gemini 3.6 Flash reviewed the most (about 23 times per drawing) and reached the highest peak of the whole batch, 0.449 on the Mona Lisa, then slid all the way back to 0.337 by the end. More reviewing did not translate into a better finish. The models kept editing past their own best performance. Objective similarity (target runs) SSIM is 0 to 1, higher is closer to the target. RMSE is 0 to 255, lower is closer to the target. Model Target ↕ Steps ↕ Time ↕ Self-reviews ↕ Final SSIM ↕ Best SSIM ↕ Final RMSE ↕ GPT-5.6 Sol Mona Lisa 44 8m41s 7 0.325 0.352 54.2 Claude Fable 5 Mona Lisa 86 18m58s 27 0.286 0.289 56.9 Grok 4.5 Mona Lisa 106 4m30s 5 0.151 0.152 95.6 Gemini 3.6 Flash Mona Lisa 60 5m31s 22 0.337 0.449 51.1 GPT-5.6 Sol Starry Night 23 5m47s 4 0.130 0.179 37.8 Claude Fable 5 Starry Night 43 11m30s 12 0.153 0.176 40.9 Grok 4.5 Starry Night 58 5m51s 3 0.039 0.039 104.9 Gemini 3.6 Flash Starry Night 74 6m36s 24 0.172 0.190 42.8 On raw SSIM, Gemini