메뉴
HN
Hacker News • 21일 전

코드 리뷰에서 확인된 GPT-6 Astra: 성능 향상, 프라이버시, 비용

IMP
7/10
핵심 요약

한 평가에서 GPT-6 Astra는 GPT-5.6 Sol 대비 약 4%, Opus 5 대비 22% 더 많은 버그를 발견했으며, 어려운 교차 파일 리뷰에서는 Sol 대비 20%, Opus 5 대비 33% 더 높은 성능을 보였습니다. 다만 API 비용은 백만 입력 토큰당 10달러, 출력 토큰당 50달러로 Sol의 2.5배에 달해, 강력한 추론 능력이 필요한 작업에 선택적으로 사용하는 전략이 중요합니다.

번역된 본문

코드 리뷰에서 가장 어려운 작업 중 상당수는 변경된 줄 바깥에서 일어납니다. 변경 사항이 개별적으로는 올바르게 보여도 시스템의 다른 곳에 있는 코드를 망가뜨릴 수 있기 때문입니다. 바로 이 점이 OpenAI의 GPT-6 Astra에 대한 우리의 초기 결과가 흥미로운 이유입니다. 우리의 평가에서 Astra는 실행 가능한 발견(findings)을 통해 GPT-5.6 Sol보다 약 4% 더 많은 라벨링된 버그를 발견했고, Opus 5보다는 22% 더 많이 발견했습니다. 가장 큰 향상은 더 어려운 교차 파일 리뷰에서 나타나며, 이 영역에서 Astra는 Sol 대비 20%, Opus 5 대비 33%의 성능 향상을 기록했습니다. 고객 규모에서 이 능력을 활용한다는 것은 고객 데이터를 보호하고 모델의 공개 API 가격을 평가해야 한다는 의미이기도 합니다.

Astra가 코드 리뷰에 더해준 것

여기서 우리의 측정 지표는 '실행 가능한 버그 커버리지', 즉 개발자가 조치할 수 있는 발견을 통해 모델이 얼마나 많은 라벨링된 버그를 잡아내는지입니다. 이 첫 평가 지표에서 GPT-5.6 Sol 대비 전체적인 향상은 비교적 미미해 보입니다. 이 평가에는 더 단순한 리뷰도 포함되어 있어, 더 강력한 모델이 차별화될 여지가 적을 수 있습니다. Astra의 더 큰 우위는 더 어려운 교차 파일 하위 집합에서 나타납니다. 이는 초기의 방향성 있는 결과입니다.

더 어려운 교차 파일 비교는 더 고무적입니다. Astra의 상대적 우위는 Sol 대비 20%, Opus 5 대비 33%로 커집니다. 이는 변경의 의도를 코드베이스 전반에 분산된 결과와 연결하는 데 가치가 있음을 시사합니다. 이러한 결과는 리뷰 성능의 한 측면만을 설명합니다. 리뷰 품질의 전체 순위를 확정하거나, 팀의 결함률을 예측하거나, 모든 풀 리퀘스트에서 동일한 향상을 약속하는 것은 아닙니다.

컨텍스트를 활용하기

큰 컨텍스트 윈도우는 정보를 담을 공간을 만듭니다. 유용한 추론이 되려면 모델이 어떤 정보가 중요한지 식별하고, 이를 연결하고, 증거로 뒷받침되는 결론에 도달해야 합니다. 우리의 해석은 Astra의 가장 흥미로운 발전이 '올바른 정보를 연결하는 능력'에 있다는 것입니다. 더 어려운 교차 파일 리뷰에서의 더 큰 향상은 관련 정보가 분산되어 있는 작업에서의 진전을 시사합니다. 하지만 이 결과는 그 진전의 원인을 분리해내지 않으며, 컨텍스트가 더 크다는 것만으로 모델 성능이 향상된다는 것을 증명하지도 않습니다.

OpenAI는 Astra를 코드, 브라우저, 전문 소프트웨어에 걸친 다단계 작업용으로 포지셔닝하고 있습니다. 모델로 제품을 만드는 팀에게 유용한 질문은 '추가적인 추론이 비용을 정당화할 만큼 결과를 충분히 바꾸는 곳이 어디인가'입니다. 증거가 흩어져 있는 어려운 작업이, 이미 저렴한 모델이 안정적으로 처리하는 일상적인 작업보다 더 유력한 검토 대상입니다. 그렇다고 해서 모든 세션을 Astra로 전환하고, 추론 노력을 최대로 올리고, 마음껏 돌리라는 뜻은 아닙니다.

강력한 추론에는 프리미엄 가격이 따릅니다

Astra의 공개된 가격에 따르면 표준 API 요금은 백만 입력 토큰당 10달러, 백만 출력 토큰당 50달러입니다. Fable 5.1은 기본 입력 및 출력 요금이 동일하지만 캐싱 가격은 다릅니다. Anthropic의 가격 문서에서 전체 내역을 확인할 수 있습니다.

이 가격을 맥락에서 이해하기 위해, 추론 토큰을 포함해 캐시되지 않은 입력 10만 토큰과 과금 대상 출력 1만 토큰을 사용하는 예시 작업을 생각해 보겠습니다. 토큰 사용량을 고정하면 공개된 요금을 비교하기 쉽습니다. 실제 작업 비용은 사용량에 따라 달라집니다.

| 모델 | 입력 / 100만 토큰 | 출력 / 100만 토큰 | 예시 작업 비용 | | GPT-5.6 Luna | $0.20 | $1.20 | $0.032 | | GPT-5.6 Terra | $2.00 | $12.00 | $0.32 | | GPT-5.6 Sol | $4.00 | $20.00 | $0.60 | | GPT-6 Astra | $10.00 | $50.00 | $1.50 | | Claude Fable 5.1 | $10.00 | $50.00 | $1.50 |

이 고정 사용량 기준으로 Astra는 Sol의 2.5배, Terra의 약 4.7배, Luna의 약 47배 비용입니다. 이는 결코 무시할 수 없는 프리미엄입니다. 다만 이것은 완료된 작업당 비용 차이를 예측하는 것은 아닙니다. 더 적은 토큰이나 더 적은 시도로 작업을 끝내는 모델이라면 격차를 좁힐 수 있습니다. OpenAI는 자체 평가 일부에서 토큰 가격이 높음에도 불구하고 Astra의 예상 작업 비용이 더 낮다고 보고합니다. 따라서 토큰 가격이나 능력 점수 하나로 결정을 내리기보다는, 성공적인 결과 하나당 총비용을 자신의 작업에서 직접 측정해볼 가치가 있습니다. OpenAI의 효율성 가이드를 참고하세요.

코드 리뷰 너머로 이전될 수 있는 것

이전 가능한 핵심 아이디어는 분산된 정보에 대한 추론입니다.

원문 보기
원문 보기 (영어)
Some of the hardest work in code review happens outside the changed lines. A change can look correct in isolation and still break code elsewhere in the system. That is what makes our early results for OpenAI's GPT-6 Astra most interesting. In our evaluation, Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol , and 22% more than Opus 5 . The biggest jump comes on harder cross-file reviews, where Astra's gains reach 20% over Sol and 33% over Opus 5 . Using that capability at customer scale also means protecting customer data and assessing the model’s public API pricing. What Astra added to code review Our measure here is actionable bug coverage, meaning how many labeled bugs a model catches through findings a developer can act on. The overall gain over GPT-5.6 Sol appears modest in this first evaluation measure. This evaluation also includes simpler reviews, where there may be less room for a stronger model to differentiate itself; Astra's larger advantage appears in the harder cross-file subset. It is an early, directional result. The harder cross-file comparison is more encouraging. Astra's relative advantage grows to 20% over Sol and 33% over Opus 5. That suggests value in connecting a change's intent to consequences distributed across a codebase. These results describe one part of review performance. They do not establish an overall ranking of review quality, predict a team's defect rate, or promise the same gain on every pull request. Putting context to work A large context window creates room for information. Useful reasoning requires the model to identify which pieces matter, connect them, and reach a conclusion supported by evidence. Our interpretation is that Astra's most interesting advance lies in connecting the right information. Its larger gains on harder cross-file reviews suggest progress on work where relevant information is distributed. They do not isolate the cause of that progress or prove that more context alone improves a model's performance. OpenAI positions Astra for multistep work across code, browsers, and professional software . For teams building with models, the useful question is where the extra reasoning changes the outcome enough to justify its cost. A difficult task with scattered evidence is a stronger candidate to investigate than a routine task that a less expensive model already handles reliably. That doesn't mean switching every session to Astra, maxing out reasoning effort, and letting it rip. Stronger reasoning comes at a premium Astra’s standard API rates are $10 per million input tokens and $50 per million output tokens , according to Astra’s published pricing . Fable 5.1 has the same base input and output rates, although caching prices differ. Anthropic’s pricing documentation provides the full breakdown. To put those prices in context, consider an illustrative task using 100,000 uncached input tokens and 10,000 billable output tokens, including reasoning tokens. Holding token usage constant makes the published rates easier to compare; actual task costs vary with usage. Model Input / 1M tokens Output / 1M tokens Illustrative task cost GPT-5.6 Luna $0.20 $1.20 $0.032 GPT-5.6 Terra $2.00 $12.00 $0.32 GPT-5.6 Sol $4.00 $20.00 $0.60 GPT-6 Astra $10.00 $50.00 $1.50 Claude Fable 5.1 $10.00 $50.00 $1.50 At that fixed usage, Astra costs 2.5 times Sol, about 4.7 times Terra, and about 47 times Luna . Those are meaningful premiums. They are not predictions of the difference in cost per completed task. A model that needs fewer tokens or fewer attempts could narrow the gap. OpenAI reports lower estimated task costs for Astra in some of its own evaluations despite higher token prices. That makes total cost per successful outcome worth measuring on your work, rather than assuming that either token price or a capability score settles the decision. Read OpenAI's efficiency guidance . What might transfer beyond code review The transferable idea is reasoning over relationships among separate sources. Our findings suggest several uses worth evaluating; we have not measured Astra on these tasks: Research synthesis: reconcile conflicting reports, connect claims to their evidence, and identify gaps that a summary of each document would miss. Operational investigation: assemble a coherent explanation from logs, incident notes, and runbooks, while separating observations from hypotheses. Requirements and policy analysis: trace a proposed change across specifications, internal policies, and implementation plans to flag inconsistencies for expert review. Document and spreadsheet work: check whether assumptions, formulas, and narrative conclusions agree across a report and its supporting materials. The common structure is scattered evidence with dependencies between its parts. Start with bounded work whose answer can be checked. The easiest way to test whether Astra is right for your workflows or products is to run it alongside your current model on the same tasks, then compare answer quality, verification time, and total cost. Building and balancing NIGHTSHIFT We also used Astra to build a whole game: NIGHTSHIFT , an action RPG made with Godot and GDScript. Its hardest problem was balancing the interactions between systems, then revisiting that balance as the game changed. That meant reasoning across seven character classes, a 988-node passive skill tree inspired by the legendary Path of Exile passive node tree , active skills, runes and socketable upgrades, skill evolutions, and co-op. Across 40 zones in 10 acts, the campaign introduces increasingly complex enemy swarms and combinations. Changing one class can also change which upgrades are useful, how a skill develops, and what a party can handle. We made fundamental changes to core systems during development and asked Astra to work through the consequences and rebalance them. "Balance" in a game like this is the most difficult creative challenge for compelling gameplay. How can you make sweeping changes or introduce entirely new game systems and mechanics while keeping progression, combat, and difficulty coherent? We also wanted players to discover overpowered endgame builds through clever combinations of classes, stats, items, passives, and active skills. The challenge was shaping a progression where reaching even mid-game was uncertain, but creativity and experimentation could pay off in spectacular ways. The reward for finding those combinations was working up to a build that could satisfyingly melt enemy swarms. The process meant returning to the game between other work, giving Astra feedback, and giving it full autonomy to use its judgement in refining that balance. Astra also built native PS5 and Xbox controller support, native macOS, web, and Linux builds, and co-op. Co-op was particularly interesting, because Astra was able to build it out for playing on the same compute and LAN, which since we wanted to distribute to coworkers running macOS, it required generating a new Xcode project, App Store Connect account, establish multiple certificates and entitlements, and notarization. Astra handled it all autonomously, only pausing to occasionally ask for authority and permissions it didn't already have. As a fun bonus, the game was also built for agents themselves to be players, which was somewhat surreal to be able to play live co-op sessions with Astra as a teammate. Several of us found it hard to put down. "I'm working on model evals" became a useful explanation for having a game open when a manager stopped by. The game gave us a creative setting to explore the same capability that stood out in code review. Astra had to reason about how each change affected the rest of the system. What we want to see next The next exciting advance would make this depth of reasoning dependable enough to use more often. That means consistent gains on difficult work, conclusions people can verify, and lower total cost for a successful outcome. That calls
관련 소식