메뉴
BL
The Decoder • 26일 전

AI 에이전트는 시간 감각이 없고 그 사실조차 인식하지 못한다

IMP
7/10
핵심 요약

MATS 연구 프로그램 소속 연구진 두 명이 Anthropic의 Claude Code와 OpenAI의 Codex를 테스트한 결과, 이 코딩 에이전트들은 작업 소요 시간을 예측하지 못하고 이미 얼마나 작업했는지도 신뢰성 있게 판단하지 못하는 것으로 나타났습니다. 시간 경과를 알려주는 도구를 제공하면 거의 정확히 맞히므로, 장시간 실행 작업의 제어 가능성 측면에서 중요한 시사점을 줍니다.

번역된 본문

AI 에이전트는 시간 감각이 없고 그 사실조차 인식하지 못한다 막시밀리안 슈라이너 (Maximilian Schreiner), 2026년 8월 30일, THE DECODER

새로운 연구에 따르면 널리 쓰이는 코딩 어시스턴트는 작업에 걸릴 시간을 예측하지 못하고, 이미 얼마나 오래 작업했는지도 신뢰성 있게 판단하지 못하는 것으로 나타났다. 이는 오래 실행되는 작업에서 문제가 된다.

AI 어시스턴트가 작업을 수행할 때, 시간이 얼마나 흐르는지 종종 전혀 알지 못한다. 이것이 두 명의 독립적인 AI 연구자가 MATS 연구 프로그램의 일환으로 수행한 연구의 핵심 결론이다. 이들은 널리 사용되는 두 코딩 어시스턴트인 Anthropic의 Claude Code와 OpenAI의 Codex를 대상으로 시간 감각을 테스트했다.

각 코딩 작업 전에 에이전트는 필요한 시간을 추정해야 했다. 그런 다음 작업을 수행한 뒤, 되돌아보며 얼마나 시간이 흘렀는지 보고했다. 테스트 소재는 ProgramBench라는 컬렉션의 200개 작업과 연구진이 직접 만든 18개 벤치마크 모음이었다.

테스트에서 에이전트들은 필요한 시간을 일관되게 과대평가했다. ProgramBench에서 두 모델 모두 난이도와 관계없이 대부분 약 90분 정도로 추정했다. 두 번째 라운드에서 Claude는 평균 3배, Codex는 6~10배 벗어났다. 짧은 작업에서 추정이 가장 부정확했으며, 몇 시간 단위의 작업에서야 일부 예측이 실제에 근접했다.

같은 AI라도 실행 환경에 따라 완전히 다르게 행동한다 결과는 모델이 실행되는 소프트웨어 환경에 따라 달라진다. Claude Code는 작업이 끝났다고 생각할 때까지 계속 작업하며 중앙값 약 90분을 소요한다. 반면 Codex는 작업과 거의 무관하게 약 30분 후에 멈춘다. 연구에 따르면 같은 언어 모델이 Claude Code에서 Codex보다 평균 2.5배 더 많은 단계를 거친다. 따라서 실행 시간은 모델과, 그리고 '하니스(harness)'라 불리는 주변 소프트웨어에 크게 의존한다.

에이전트는 자신의 작업 품질을 판단하는 데에도 마찬가지로 신뢰할 수 없다. 이전 세대 모델인 Opus 4.8과 GPT-5.5는 자신의 결과를 평균 20점 과대평가했으며, 실패한 작업에서도 스스로 높은 점수를 매겼다. 한 사례에서 두 모델 모두 자신의 작업을 약 70% 성공했다고 평가했지만, 실제 점수는 각각 7%와 14.5%였다.

연구진은 이러한 자기 평가 능력이 중요하다고 말한다. 에이전트가 몇 시간에 걸친 긴 작업을 신뢰성 있게 수행하려면 "이 작업을 2시간 동안 반복하라" 같은 지시를 따를 수 있어야 한다. 시간을 계속 잘못 판단하는 에이전트는 제어하기 어렵다. 다음 단계로 저자들은 에이전트가 정해진 작업 시간을 준수할 수 있는지 테스트할 계획이다. 경과 시간을 알려주는 도구에 접근할 수 있을 때는 에이전트가 거의 매번 정확하게 맞혔다.

원문 보기
원문 보기 (영어)
AI agents have no sense of time and are not aware of it Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 30, 2026 Nano Banana Pro prompted by THE DECODER A new study finds that popular coding assistants can't predict how long a task will take, and they can't reliably tell how long they've already been working. That's a problem for long-running jobs. When an AI assistant works on a task, it often has no idea how much time is passing. That's the takeaway from a study by two independent AI researchers, done as part of the MATS research program. The pair tested two widely used coding assistants, Anthropic's Claude Code and OpenAI's Codex, on their sense of time. Before each coding task, the agents had to estimate how long they'd need. Then they solved the task and, looking back, reported how much time had passed. The test material came from 200 tasks in a collection called ProgramBench, plus the researchers' own suite of 18 benchmarks. In the tests , the agents consistently overestimated how much time they'd need. On ProgramBench, both models mostly guessed around 90 minutes, no matter the difficulty. In the second round, Claude was off by three times on average, Codex by six to ten times. The estimates were worst for short tasks, and only in the multi-hour range did some predictions come close to reality. The same AI behaves completely differently depending on its setup The results shift based on the software setup the models run in. Claude Code keeps working until it thinks the task is done, a median of about 90 minutes. Codex, on the other hand, stops after roughly half an hour, almost regardless of the task. According to the study, the same language model takes 2.5 times more steps in Claude Code than in Codex on average. So runtime depends on the model and heavily on the surrounding software, known as the harness. The agents are just as unreliable at judging the quality of their own work. The older models, Opus 4.8 and GPT-5.5, overrated their results by 20 points on average and handed themselves high marks even on failed tasks. In one case, both figured their work was about 70 percent successful. The actual scores were 7 and 14.5 percent. The researchers say this ability to self-assess matters. For an agent to work reliably on long tasks that run for hours, it has to follow instructions like "iterate on this task for two hours." An agent that constantly misjudges the time is hard to control. Next, the authors want to test whether agents can stick to a set work duration. When the agents got access to a tool that reports elapsed time, they got it right almost every time. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->