메뉴
HN
Hacker News • 58일 전

물리 AI 분야 최강 모델은? GPT-5.6 vs Claude Fable 5 성능 비교

IMP
8/10
핵심 요약

에이전트 AI가 작성한 코드는 문제없이 실행되더라도 실제 물리 법칙과 어긋나는 치명적인 오류를 범할 수 있습니다. 이에 JuliaHub은 모델링 및 시뮬레이션 에이전트인 Dyad를 활용해 최신 AI 모델들의 '물리적 정확성'을 엄격하게 테스트했습니다. 그 결과, Anthropic의 Claude Fable 5가 가장 우수한 성능을 보였으나 비용이 높았고, OpenAI의 GPT-5.6 계열 모델들은 더 낮은 비용과 빠른 속도를 기록했습니다.

번역된 본문

홈 / 블로그 / 물리 AI를 위한 GPT-5.6 vs Claude Fable 5, 어느 모델이 가장 성능이 뛰어난가? ‹ › 연구 및 혁신 물리 AI를 위한 GPT-5.6 vs Claude Fable 5, 어느 모델이 가장 성능이 뛰어난가? 연구 및 혁신 물리 AI를 위한 GPT-5.6 vs Claude Fable 5, 어느 모델이 가장 성능이 뛰어난가? 발행일: 2026년 7월 20일 기여자 공유 발행일: 2026년 7월 20일 기여자 공유

우리는 당사의 에이전트에 최신 프론티어 모델인 gpt-5.6-terra를 적용하여 난이도가 다양한 5가지 모델링 및 시뮬레이션 문제를 테스트했습니다. 다음은 그 결과입니다:

claude-fable-5: 가중 평점 0.889 · 테스트당 $9.60 · 16.1분 소요 gpt-5.6-sol: 가중 평점 0.814 · 테스트당 $1.74 · 13.4분 소요 gpt-5.6-terra: 가중 평점 0.786 · 테스트당 $1.25 · 12.6분 소요 gpt-5.6-luna: 가중 평점 0.727 · 테스트당 $3.26 · 25.0분 소요

물리 AI(Physical AI)의 생사는 모델링된 물리 법칙이 정확한지에 달려 있습니다. 항공기, 증류탑, 대전 입사체의 모델은 물리적으로 불가능한 원리를 담고 있더라도 오류 없이 컴파일되고 실행될 수 있습니다. 에이전트 AI(Agentic AI)는 이러한 실패 모드를 악화시킵니다. 에이전트는 테스트의 피드백을 통해 방향을 잡는데, 이 테스트는 종종 에이전트 본인이 작성하기 때문입니다. 웹 개발이나 컴파일러 같은 도메인에서는 올바른 동작이 명확하게 통제되고 검증 가능하므로 이러한 루프가 잘 작동합니다. 하지만 엔지니어링 분야에서는 "이것이 현실 세계와 일치하는가?"가 진짜 문제이며, 이를 해결하기는 훨씬 어렵습니다. 에이전트는 자신이 작성한 모든 테스트를 통과할 수 있지만, 그 테스트들은 실제 모델이 사용될 환경에서는 성립하지 않는 단순화된 가정에 기반할 수 있습니다.

또한 기존에 발표된 수치들에는 신뢰성 문제가 있습니다. 아무도 모델 제공업체가 자체적으로 보고한 벤치마크를 곧이곧대로 믿지 않습니다. 점수를 공개하는 제공업체가 그 점수를 높이기 위해 벤치마크에 맞춰 튜닝할 강력한 동기를 가진 당사자이기 때문입니다. 반면 우리의 동기는 다릅니다. JuliaHub은 다양한 공급업체의 멀티 에이전트 백엔드와 함께 Dyad를 제공합니다. 우리는 어떤 연구소의 모델이든 상관없이 사용자가 최고의 Dyad 경험을 할 때 이익을 얻습니다. 특정 모델이 물리적 모델링에 더 뛰어나다면, 이를 파악하고 추천하는 것이 우리의 이익입니다. 그래서 우리는 조사했습니다: 과연 어느 모델이 가장 뛰어날까?

어떤 프론티어 모델이 이 작업을 가장 잘 처리하는지 측정하기 위해 우리는 다른 모든 변수를 고정했습니다. 모델링 및 시뮬레이션 워크플로우를 위해 특별히 제작된 Dyad AI 에이전트 하네스(harness)는 문제, 추론 노력(xhigh), 컨텍스트 창(1M), 토큰 예산(128k)과 함께 모든 테스트에서 동일하게 유지되었습니다. 유일한 변수는 모델입니다: OpenAI의 GPT 5.6 패밀리(terra, sol, luna) 대 Anthropic의 claude-fable-5입니다. 4가지 핵심 문제에 대해 모델당 3번의 테스트를 진행했고, 5번째 문제에 대해서는 장기적인(long-horizon) 테스트를 각각 1번씩 진행하여 총 52번의 평가를 마쳤습니다.

01 · 문제들 (물리 AI 평가의 작동 방식) 우리는 내부 평가의 일부를 바탕으로 탐구를 진행했습니다: 모델링 및 시뮬레이션 실무에서 뽑아낸 5가지 비공개 문제입니다. 이를 난이도별로 배열했으며, 난이도는 정확한 솔루션이 거쳐야 할 단계, 사소한 실수가 치명적인 세심한 디테일, 쉽게 접근할 수 있는 지름길, 사양서 분석부터 하네스(harness) 수리까지의 모델링 자체에 대한 엔지니어링 작업 등을 기준으로 정의했습니다. 각 문제를 선정한 이유는 동일합니다. 즉, 물리적으로는 틀렸음에도 코드가 컴파일되고 정상적으로 실행되는 해결책이 나올 수 있기 때문입니다. 따라서 평가는 코드를 무시하고, 비공개 정답(Ground truth)에 대해 전체 시뮬레이션 궤적을 채점합니다.

5가지 중 가장 어려운 문제는 NASA의 HL-20 비행 기체입니다. 이 문제의 쉬운 버전은 이 영상에서 자세히 설명하고 있으며, 이 연구에서는 더 어렵고 비공개된 변형 문제를 실행합니다.

P1 구성 일관성 (Constitutive consistency) P2 제약 일관성 (Constrained consistency) P3 상대론적 역학 (Relativistic dynamics) P4 정상 상태 선형화 (Steady-state linearization) P5 장기 비행 기체 (Long-horizon flight vehicle) 이전의 모든 문제를 한 번에 포함 엔지니어링 사양서 및 공기 역학 데이터를 읽고, 6자유도 운동을 도출하며, 시뮬레이션, 검증 및 반복합니다.

원문 보기
원문 보기 (영어)
Home / Blog / GPT-5.6 vs Claude Fable 5 for Physical AI, which performs best? ‹ › Research & Innovation GPT-5.6 vs Claude Fable 5 for Physical AI, which performs best? Research & Innovation GPT-5.6 vs Claude Fable 5 for Physical AI, which performs best? Date Published Jul 20, 2026 Contributors Share Date Published Jul 20, 2026 Contributors Share We tested the latest frontier models gpt-5.6-terra in our agent, on five modeling and simulation problems of varying difficulty. Here are the results: claude-fable-5 0.889 weighted score $9.60 / trial · 16.1 min gpt-5.6-sol 0.814 weighted score $1.74 / trial · 13.4 min gpt-5.6-terra 0.786 weighted score $1.25 / trial · 12.6 min gpt-5.6-luna 0.727 weighted score $3.26 / trial · 25.0 min Physical AI lives or dies on whether the modeled physics is correct. A model of an aircraft, a separation column or a charged particle can compile and run cleanly while the physics it encodes is impossible. Agentic AI makes this failure mode worse, because agents steer by feedback from tests, and the tests are often written by the same agent. In domains like web development or compilers that loop works well enough, since correct behavior is contained and checkable. In engineering, the real question is “does this match the real world?”, and that is much harder to close: the agent can pass every test it wrote while those tests rest on simplifications that don’t hold in the regime where the model will actually be used. There is also a trust problem with the numbers that already exist. Nobody takes the model providers’ self-reported benchmarks at face value, and for good reason: the provider publishing the score is the same party with every incentive to tune for it. Our incentives point the other way. JuliaHub ships Dyad with multiple agent backends across vendors, and we win when our users get the best possible Dyad experience, whichever lab’s model delivers it. If one model is better at physical modeling, it is in our interest to find that out and recommend it. So we investigated: which one actually is? To measure which frontier model handles that work best, we hold everything else still. The Dyad AI agent harness - purpose-built for modeling & simulation workflows - is pinned across every trial, along with the problems, the reasoning effort (xhigh), the context window (1M) and the token budget (128k). The only variable is the model: OpenAI’s GPT 5.6 family ( terra , sol , luna ) against Anthropic’s claude-fable-5 . Three trials per model on each of the four core problems, one long-horizon trial each on the fifth: 52 graded runs in all. 01 · The problems How Physical AI Evaluation Works. We ground our exploration in a subset of our internal evals: five sealed problems drawn from the daily work of modeling & simulation. We order them by difficulty, defined in terms of the stages a correct solution must chain through, the careful details where a small slip is fatal, the easy shortcuts within reach, and the engineering work around the modeling itself - from parsing specs to repairing harnesses. We selected each for the same property: they admit solutions that compile and run green while being physically wrong. Grading therefore ignores the code and scores the full simulated trajectory against sealed ground truth. The hardest of the five is NASA’s HL-20 flight vehicle. An easier version of the problem is walked through in detail in this video ; the study runs a harder, sealed variant. P1 Constitutive consistency P2 Constrained consistency P3 Relativistic dynamics P4 Steady-state linearization P5 Long-horizon flight vehicle Every previous problem at once Read the engineering spec and aerodynamic data, derive six-degree-of-freedom motion, simulate, verify, and iterate. P1 Constitutive consistency P2 Constrained consistency P3 Relativistic dynamics P4 Steady-state linearization P5 Long-horizon flight vehicle Every previous problem at once Read the engineering spec and aerodynamic data, derive six-degree-of-freedom motion, simulate, verify, and iterate. P1 Constitutive consistency P2 Constrained consistency P3 Relativistic dynamics P4 Steady-state linearization P5 Long-horizon flight vehicle Every previous problem at once Read the engineering spec and aerodynamic data, derive six-degree-of-freedom motion, simulate, verify, and iterate. 02 · The scoreboard What they scored . We run every model through the full suite - three trials on each of P1–P4, one on the long-horizon P5 - and grade every trial against sealed ground truth. Per-problem scores fold into a single number through a difficulty-weighted average, the hardest weighing the most. Cost and time carry no weighting: average cost per trial and average wall-clock per trial each get a tab: The single number above has anatomy. Per problem, averaged across trials, the same three lenses: 03 · Work-style fingerprints Each model has its own approach . We analyze the transcripts behind every trial. Each tool call is bucketed by what the agent was doing at that moment - orienting in the problem, deriving and testing physics, editing the model, compiling, or hunting for something to copy - then counted per trial and averaged per model. The result is a work-style fingerprint: The bars say how the work divides; the transcripts show the temperament behind it. Four portraits, drawn from where each model was tested hardest: Fable verifies by constructing examples that fail. In one trial on the relativistic dynamics problem it swept solver tolerances across three decades, watching the mass-shell invariant u·u drift, before committing the run at 1e-10. It checked its fields against an independent reimplementation at 200 random points - then deliberately sign-flipped a field component to prove the check could fail (residual 2.0 wrong, 0.066 right). It refused to treat the time component of the initial four-velocity as a free input, deriving it from the mass-shell constraint instead. With twenty-eight calls per trial, it is a model that writes only once because it has already tried to break what it is about to write. The bill for that conviction arrives later, in tokens. Sol specifies more than the task asks. The habit cuts both ways. On the constitutive problem it built one of only two genuinely independent oracles in the cohort, closing a quadrature check of the constitutive identity to 3.7e-12 under the session title “Falsify the update against the independent bulk-modulus definition.” On the relativistic problem the same instinct manufactured its slip: the problem pins the z-component of the initial four-velocity at -0.62, and sol renormalized the state so the coordinate velocity came out at -0.62 instead, handing the particle 27% excess longitudinal momentum. Its own consistency check then blessed the misreading: mass-shell exactly one, velocity exactly as read, every downstream check self-consistent and wrong. Precision is not the same thing as fidelity to the spec. Luna iterates where it should investigate. It carries the least configuration into a run and reads more of the problem than anyone, which serves it well on the easier problems. When the physics pushed back on the constrained-consistency problem, the same temperament turned into churn: 68 tool calls in a single trial, and a committed solution still 30% off. That trial was also the most-validated in the cohort - 22 checks, every one verifying its own equations against themselves - and the pattern repeats: on the constitutive problem, luna was the only model that never tested against anything independent of its committed equations. Validation without an external reference isn’t validation. Terra economizes. It spends the smallest share of its clock on derivation and testing and the largest on editing - more trial and error against the file - and it broadly gets away with it: correct formulations, the cheapest and fastest trials in the study. Its slip on the steady-state linearization problem is the study’s narrowest miss, and its stra
관련 소식