메뉴
BL
The Decoder • 13일 전

GPT-6 Astra, 감시 드론 조종·자율 사업 운영 모두 성공

IMP
8/10
핵심 요약

안돈 랩스(Andon Labs)의 에이전트 벤치마크에서 GPT-6 Astra가 자판기 사업 시뮬레이션에서 Claude Fable 5.1보다 약 3배 높은 최종 잔액을 기록했습니다. 또한 드론이 특정 인물을 자율적으로 찾아 추적하는 코드를 작성하는 Drone-Bench에서 모든 5개 과제에서 인간-AI 기준선을 넘은 최초의 모델이 되었습니다. 다만 성공률이 아직 불안정하다는 점은 유의해야 합니다.

번역된 본문

GPT-6 Astra가 감시 드론을 조종하고 자체적으로 사업을 운영하다

톰리슬라브 베즈말리노비치, 2026년 9월 13일

핵심 요점

  • 안돈 랩스(Andon Labs)는 GPT-6 Astra를 두 개의 에이전트 벤치마크에서 테스트했으며, 이 모델은 Claude Fable 5.1보다 훨씬 높은 점수를 기록했다.
  • 자판기 사업 시뮬레이션에서 Astra는 평균 15,515달러의 최종 은행 잔액을 달성하며 Fable의 결과의 거의 3배를 기록했다.
  • Drone-Bench에서 Astra는 드론이 특정 인물을 자율적으로 찾아 추적하는 코드를 작성하는 것을 포함해 5개 세부 과제 모두에서 인간-AI 기준선을 넘은 최초의 모델이 되었다.

안돈 랩스는 GPT-6 Astra를 성격이 매우 다른 두 개의 에이전트 벤치마크에서 테스트했다. 재고를 구매하고 자판기를 운영하는 과제에서 이 OpenAI 모델은 Claude Fable 5.1을 압도했다. 드론 감시 과제에서 Astra는 5개 세부 과제 모두에서 인간-AI 기준선을 넘은 최초의 모델이지만, 성공률은 여전히 불안정하다.

OpenAI의 GPT-6 Astra는 안돈 랩스의 두 개 에이전트 벤치마크에서 이전의 모든 프런티어 모델을 능가했다. 이 연구소는 Vending-Bench와 Drone-Bench를 통해 AI 모델이 장기간에 걸쳐 독립적으로 행동하거나 물리적 시스템을 위한 소프트웨어를 작성하는 능력을 측정한다. 자판기 사업 시뮬레이션에서 Astra는 Claude Fable 5.1의 거의 3배 수익을 올렸다. Drone-Bench에서 안돈 랩스에 따르면 Astra는 최고 시도 성적이 5개 세부 과제 모두에서 인간과 AI가 함께 개발한 기준선을 넘은 최초의 모델이다.

Astra, Claude Fable 5.1보다 강하게 협상하고 지출은 적게

Vending-Bench에서 각 모델은 500달러를 받고 시뮬레이션된 1년 동안 자판기를 운영해야 한다. 공급업체를 찾고, 구매 가격을 협상하고, 상품을 주문하고, 소매 가격을 정하며, 은행 잔액을 늘리려고 노력한다.

안돈 랩스에 따르면 6회 실행에서 GPT-6 Astra는 평균 15,515달러를 기록했다. Claude Fable 5.1은 평균 5,422달러였다. Fable의 최고 기록인 9,874달러조차 Astra의 최저 성적인 13,272달러에 크게 미치지 못했다. Astra는 Vending-Bench 2 리더보드 정상에 오른 최초의 OpenAI 모델이다. 안돈 랩스에 따르면 2위 모델과의 격차도 이 벤치마크 사상 최대다.

가장 큰 차이 중 하나는 조달 분야에서 나타난다. Fable은 시간이 지날수록 더 불리한 거래를 수용한다. 일반 코카콜라 캔 하나의 평균 구매 가격은 시뮬레이션 첫 90일의 1.17달러에서 시뮬레이션 연말 무렵 2.21달러로 상승한다. Astra는 더 일관되게 협상한다. 안돈 랩스가 기록한 사례 중 한 공급업체가 상품 바구니에 226.32달러를 제시했는데, Astra는 108달러에서 물러서지 않았고 결국 거래를 성사시켰다.

Astra는 신뢰할 수 없는 공급업체도 더 잘 다룬다. 6회 실행 동안 Fable 5.1은 이미 폐업한 공급업체에 45건의 선지급을 하여 14,331달러를 잃었다. Astra는 폐업을 64건 더 많이 겪었지만, 안돈 랩스에 따르면 이런 선지급으로 인한 확인된 손실은 없었다. Fable은 문제를 인식하고 서면 확인 후에만 지불한다는 규칙을 스스로 만들었지만, 며칠 뒤 자신의 규칙을 스스로 어겼다.

Astra, 가격 담합 제안 거부

안돈 랩스는 여러 AI 에이전트가 같은 장소에서 경쟁 자판기를 운영하는 Vending-Bench Arena에서도 모델을 테스트한다. Astra는 중국 모델 GLM-5.3의 가격 담합 제안을 명확히 거부했다. 안돈 랩스는 조사한 3번의 아레나 게임에서 Astra가 거짓말을 한 사례를 관찰하지 못했다.

Claude Fable 5.1은 안돈 랩스가 불법 가격 담합으로 분류한 협정에 GLM-5.3과 함께 참여했다. Fable은 자신에게 이익이 될 때만 협정을 준수했다. Astra는 세 게임 모두에서 승리했다.

안돈 랩스는 Astra를 더 강력한 경제적 수행 능력과 더 나은 정렬(alignment)을 갖춘 모델로 평가하지만, 이 평가는 벤치마크에서 관찰된 행동에 기반한 것으로 다른 상황에 자동으로 적용되지는 않는다.

Astra, 5개 Drone-Bench 과제를 모두 통과한 최초의 모델

Drone-Bench는 다른 유형의 에이전트 능력을 테스트한다. 모델은 저가형 DJI Tello EDU 드론이 사무실을 자율적으로 탐색하고, 특정 인물을 식별하고, 그를 추적할 수 있게 하는 코드를 작성한다. 이 벤치마크는 다섯 단계로 구성된다: 환경의 3D 재구성, 드론 위치 추정, 내비게이션, 표적 식별...

원문 보기
원문 보기 (영어)
GPT-6 Astra pilots a surveillance drone and runs a business on its own Tomislav Bezmalinović Sep 13, 2026 Nano Banana Pro prompted by THE DECODER Key Points Andon Labs tested GPT-6 Astra on two agent benchmarks where the model scored far better than Claude Fable 5.1. Running a simulated vending machine business, Astra averaged $15,515 in final bank balance, nearly three times Fable's result. On Drone-Bench, Astra became the first model to beat the human-AI baseline on all five subtasks, including writing code that lets a drone autonomously find and follow a specific person. Ask about this article… Search Andon Labs tested GPT-6 Astra on two very different agent benchmarks. When buying inventory and running a vending machine, the OpenAI model crushes Claude Fable 5.1. On drone surveillance, Astra is the first model to beat the human-AI baseline on all five subtasks, though its success rate remains unreliable. OpenAI's GPT-6 Astra outperforms all previous frontier models on two agent benchmarks from Andon Labs. The research lab uses Vending-Bench and Drone-Bench to measure how well AI models act independently over long periods or write software for physical systems. In a simulated vending machine business, Astra earned nearly three times as much as Claude Fable 5.1 . On Drone-Bench, Andon Labs says Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks. Ad Astra negotiates harder and spends less than Claude Fable 5.1 In Vending-Bench , each model gets $500 and has to run a vending machine over a simulated year. It finds suppliers, negotiates purchase prices, orders goods, sets retail prices, and tries to grow its bank balance. Ad Across six runs, GPT-6 Astra averaged $15,515, according to Andon Labs. Claude Fable 5.1 averaged $5,422. Even Fable's best run at $9,874 fell well short of Astra's worst result of $13,272. Astra is the first OpenAI model to top the Vending-Bench 2 leaderboard. The gap to the second-place model is also the largest the benchmark has ever seen, according to Andon Labs. One of the biggest differences shows up in procurement. Fable accepts worse deals over time. For a regular can of Coca-Cola, its average purchase price rises from $1.17 in the first 90 days to $2.21 toward the end of the simulated year. Astra negotiates more consistently. In one case Andon Labs documented, a supplier quoted $226.32 for a basket of goods. Astra held firm at $108 and got the deal. Ad Astra also handles unreliable suppliers better. Across six runs, Fable 5.1 made 45 prepayments to suppliers that had already shut down, losing $14,331. Astra encountered even more closures at 64, but Andon Labs says it recorded no identified losses from such prepayments. Fable recognized the problem and wrote a rule to only pay after written confirmation. Days later, the model broke its own rule. Astra refuses price-fixing schemes Andon Labs also tests models in Vending-Bench Arena , where multiple AI agents run competing vending machines at the same location. Astra explicitly refused a price-fixing proposal from the Chinese model GLM-5.3 . Andon Labs observed no instances of lying from Astra across the three arena games it studied. Ad Claude Fable 5.1 participated in what Andon Labs classified as an illegal price-fixing arrangement with GLM-5.3. Fable only honored the agreement when it served its own interests. Astra won all three games. Ad Andon Labs rates Astra as both a stronger economic performer and better aligned, though that assessment is based on behaviors observed in the benchmark and doesn't automatically transfer to other situations. Astra is the first model to beat all five Drone-Bench tasks Drone-Bench tests a different kind of agent capability. Models write code that lets a cheap DJI Tello EDU drone autonomously navigate an office, identify a specific person, and follow them. The benchmark has five steps: 3D reconstruction of the environment, drone localization, navigation, target person detection, and tracking. Each task is scored individually against code that a human developer built with coding agents for Andon's own demo. Every model gets ten runs per task and can submit up to ten code versions per run. After each attempt, it receives a score and can improve its solution. In the original paper from July , Claude Fable 5 was the strongest model. Frontier models had beaten the human-AI baseline on four of five tasks in at least one run. 3D reconstruction remained unsolved. Andon Labs reported that Astra is the first model whose best submissions beat the baseline on all five Drone-Bench tasks, including reconstruction. Astra built a pipeline combining COLMAP and DA3 with added depth filtering. The model used office video footage to generate a navigable 3D model that scored higher than the human-AI reference solution, according to Andon Labs. Best-case scores don't mean reliable performance On person detection, Astra beats the baseline in four out of ten runs. On 3D reconstruction, it manages that in just one out of ten. Andon Labs calculates that an average Astra run has only a 2.8 percent chance of passing all five steps in sequence. Astra proved for the first time that a general-purpose frontier model can produce code above the baseline for every part of the task. But multiply the probabilities for a complete end-to-end run, and the odds are still low. Based on progress over the past two years, the team projects that a frontier model could solve all five tasks in a single attempt by Q1 2027. GPT-6 Astra already works as a surveillance drone In a demo from Andon Labs , GPT-6 Astra flies a drone autonomously through an office with the prompt "ChatGPT, find this person and follow them." The model identifies a specific person and tracks them. Spatial mapping, navigation, and person tracking all run without any human input. Other benchmarks also show that GPT-6 Astra has particularly strong spatial reasoning . When critics questioned why they were building the kind of technology everyone keeps warning about, Andon Labs responded that the benchmark doesn't help AI fly drones but measures how well current models can already do it. Six months ago, frontier models failed at these tasks and crashed. Astra now beats the human baseline on every subtask. Andon Labs argues that the public and lawmakers need to know about these capabilities before AI-powered drones reach superhuman navigation skills. No lab has access to the benchmark. Andon Labs runs all evaluations itself to prevent companies from optimizing their models for the test. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Andon Labs Vending-Bench | Andon Labs Drone-Bench | Andon Labs @ X