메뉴
HN
Hacker News • 29일 전

터미널벤치-사이언스: 과학 연구 워크플로에서 AI 에이전트 평가

IMP
7/10
핵심 요약

스탠퍼드 대학 연구진이 주도하고 Terminal-Bench 팀이 개발한 'Terminal-Bench-Science'는 실제 과학자들의 연구 워크플로를 기반으로 AI 에이전트의 과학적 역량을 평가하는 벤치마크입니다. 첫 버전(0.1)은 생명·물리·지구·수학·공학 분야의 70개 과제를 포함하며, 가장 성능이 좋았던 Claude Opus 5의 해결률은 30%에 그쳤습니다. 이 벤치마크는 모델 개발사가 아닌 과학자들이 기준을 설정하고, 최신 AI 발전에 맞춰 지속적으로 진화하는 것이 특징입니다.

번역된 본문

Terminal-Bench-Science 0.1

Terminal-Bench-Science는 연구자들의 실제 연구 워크플로를 바탕으로 AI 에이전트를 평가합니다. 모델 개발사나 데이터 공급업체가 아닌 과학자들이 AI의 과학적 역량 기준을 설정합니다.

Terminal-Bench-Science는 스탠퍼드 대학교 연구진이 주도하고 Terminal-Bench를 개발한 팀이 전 세계 다양한 과학 분야 및 연구기관의 도메인 전문가들과 협력하여 구축한 벤치마크입니다. 과학 연구에서 추출한 다양하고 도전적인, 전문가가 직접 큐레이션한 워크플로를 통해 AI 에이전트의 역량을 측정합니다. Terminal-Bench-Science는 최첨단 AI와 함께 진화하는 지속적인 벤치마크로, 과학적 필요와 AI 개발 사이의 피드백 루프를 만듭니다. 첫 릴리스에는 생명과학, 물리과학, 지구과학, 수학, 공학 분야의 70개 과제가 포함되어 있습니다. 평가된 모델 중 가장 강력했던 Claude Opus 5는 Terminal-Bench-Science 0.1에서 30%의 해결률을 기록했습니다.

개요

Terminal-Bench가 소프트웨어 엔지니어링 분야의 AI 에이전트 발전을 이끌어왔다면, Terminal-Bench-Science는 같은 야망을 과학 분야로 가져옵니다. 우리의 목표는 에이전트가 유용한 연구 조수가 될 수 있는 과학적 역량을 갖추도록 개발을 촉진하는 것입니다. 이러한 에이전트는 기술적으로 어렵고 시간이 많이 소요되는 워크플로를 수행함으로써, 과학자들이 인간의 판단이 가장 중요한 부분—연구 질문 정의, 가설 수립, 결과 해석 및 검증, 연구 결과 소통—에 더 많은 시간을 집중할 수 있게 해줍니다. 이런 역할에서 AI 에이전트는 연구자가 달성할 수 있는 성과를 확장하고 과학적 발견을 가속화하는 데 도움이 될 수 있습니다.

이를 위해서는 실제 과학 실천을 반영하고, 역량에 대한 검증 가능한 증거를 제공하며, 최첨단 AI와 함께 발전하는 벤치마크가 필요합니다.

실제 과학 워크플로에서 추출한 벤치마크가 필요합니다. 과학적 역량은 현직 과학자들이 직접 제공하는 실제 연구 실천으로 평가되어야 하며, 교과서 질문이나 표준화된 연습 문제가 아니어야 합니다. Terminal-Bench-Science는 여러 분야의 과학자들에게 직접적인 발언권과 공통 플랫폼을 제공하여, 자신들이 관심을 두는 문제에 대해 AI 발전의 기준을 설정할 수 있게 합니다. 과학에서의 요건은 매우 높으며, 벤치마크는 외부 이해관계가 아닌 과학 공동체의 우선순위를 반영해야 합니다.

과학적 역량에 대한 검증 가능한 증거가 필요합니다. 신뢰할 수 있는 평가 없이는 에이전트의 역량이 향상되고 있는지, 한계가 어디에 남아 있는지 알 수 없습니다. Terminal-Bench-Science는 현실적인 환경에서 에이전트를 평가하고, 분석, 시뮬레이션, 증명, 코드, 데이터 산출물 등 구체적인 결과물을 재현 가능한 과제별 테스트로 채점합니다.

최첨단과 보조를 맞추는 벤치마크가 필요합니다. 과학 벤치마크는 종종 진보를 이끄는 메커니즘이 아니라 발표용 논문처럼 취급됩니다. 한 번 출시된 후 모델이 발전하고 알려진 한계가 지속됨에도 방치됩니다. Terminal-Bench-Science는 최첨단 AI와 함께 진화하는 지속적인 벤치마크입니다. 정기적인 릴리스를 통해 과학자들은 새로운 워크플로를 기여하고, 기존 과제를 개선하며, 과학적 필요와 AI 개발 간의 피드백 루프를 만들 수 있습니다.

과제

Terminal-Bench-Science 0.1은 생명과학, 물리과학, 지구과학, 수학, 공학에 걸쳐 70개 과제를 포함합니다. 과제는 과학 데이터 분석, 통계적 추론, 시뮬레이션, 최적화, 정리 증명, 영상 복원, 신호 처리, 역문제, 센서 보정, 모델 피팅, 분류, 과학 머신러닝 등을 다룹니다. 과제는 GitHub의 공개 프로세스를 통해 연구자들이 기여하며, Discord의 #tb-science 채널에서 논의와 피드백이 이루어집니다. 기여는 제안으로 시작되며, 리뷰어가 각 아이디어를 논의하고 피드백을 남기고, 벤치마크에서 측정할 가치가 있는 과학적으로 근거 있는 워크플로로 판단되는 것을 승인합니다. 승인된 제안은 풀 리퀘스트로 구현되며, 리뷰어는 각 과제가 (원문 여기서 끊김)

원문 보기
원문 보기 (영어)
Terminal-Bench-Science 0.1 Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work. Scientists, not model developers or data vendors, set the bar for scientific capability in AI. Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world. It measures the AI agent capabilities through a diverse set of challenging, expert-curated workflows drawn from scientific research. Terminal-Bench-Science is a continuous benchmark that evolves alongside frontier AI, creating a feedback loop between scientific needs and AI development. Our first release includes 70 tasks from the life, physical, Earth, mathematical, and engineering sciences. The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1. Overview While Terminal-Bench has driven progress in AI agents for software engineering, Terminal-Bench-Science brings the same ambition to science. Our goal is to drive the development of agents with scientific capabilities that make them useful research assistants. These agents should execute technically demanding and time-consuming workflows, freeing scientists to focus more of their time on the parts of science where human judgment matters most: defining research questions, forming hypotheses, interpreting and validating results, and communicating findings. In this role, AI agents can extend what researchers accomplish and help accelerate scientific discovery. Achieving this requires benchmarks that reflect real scientific practice, provide verifiable evidence of capability, and evolve alongside the AI frontier. We need benchmarks drawn from real scientific workflows. Scientific capability should be evaluated on real research practice, not textbook questions or standardized exercises, contributed by practicing scientists themselves. Terminal-Bench-Science gives scientists across domains a direct voice and a shared platform to set the bar for AI progress on the problems they care about. The stakes in science are too high, and its benchmarks must reflect the scientific community's priorities rather than outside interests. We need verifiable evidence of scientific capability. Without reliable evaluation, we cannot tell whether agent capabilities are improving or where their limitations remain. Terminal-Bench-Science evaluates agents in realistic environments and grades concrete artifacts such as analyses, simulations, proofs, code, and data products with reproducible, task-specific tests. We need a benchmark that keeps pace with the frontier. Too often, scientific benchmarks are treated as papers to publish rather than mechanisms for driving progress. They are released once and then abandoned as models advance and known limitations persist. Terminal-Bench-Science is a continuous benchmark that evolves alongside the AI frontier. Through regular releases, scientists can contribute new workflows, improve existing tasks, and create a feedback loop between scientific needs and AI development. Tasks Terminal-Bench-Science 0.1 includes 70 tasks across the life, physical, Earth, mathematical, and engineering sciences. Tasks span scientific data analysis, statistical inference, simulation, optimization, theorem proving, image reconstruction, signal processing, inverse problems, sensor calibration, model fitting, classification, and scientific machine learning. Tasks are contributed by researchers through an open process on GitHub , with discussion and feedback in the #tb-science channel on Discord . Contributions begin as proposals, where reviewers discuss each idea, leave feedback, and approve those that look like a strong fit: scientifically grounded workflows worth measuring in the benchmark. Approved proposals are implemented as pull requests, where reviewers confirm that each task is objectively verifiable, genuinely challenging for AI agents, and not something today's frontier systems already solve easily. To merge, domain reviewers assess scientific validity and realism, technical reviewers inspect task construction and verification, and a bar raiser performs a final quality check. Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1. That selectivity reflects how difficult it is to create tasks that are scientifically interesting, challenging for frontier agents, and sufficiently well specified for rigorous evaluation. Progress across proposals, pull requests, and reviews is tracked on the public task dashboard . Terminal-Bench-Science 0.1 Contributions Results Terminal-Bench-Science 0.1 leaves substantial room for progress on AI agents for scientific research. Each evaluated model ran three independent trials per task across all 70 tasks. Claude Opus 5 with Claude Code achieves the highest resolution rate at 30.0%, followed by GPT-5.6 Sol with Codex at 22.4% and Claude Fable 5 with Claude Code at 21.4%. Claude Opus 4.8 sits in the middle at 10.5%. GPT-5.6 Terra, Kimi K3, and Grok 4.6 all resolve less than 10% of tasks. GLM 5.3 is the strongest open model at 8.1%, and GPT-5.6 Luna is last at 3.3%. Terminal-Bench-Science distinguishes between systems about as well as Terminal-Bench 3.0 while pushing resolution rates down by more than 10 percentage points for every model evaluated on both. That gap is deliberate: during review, tasks were calibrated to challenge the newest frontier models. Performance is only one dimension of progress. The cost-resolution plot shows total evaluation cost across all 70 tasks against resolution rate. GPT-5.6 Luna, Kimi K3, and GPT-5.6 Terra occupy the low-cost end of the frontier. GPT-5.6 Sol and Claude Opus 5 reach the highest resolution rates at greater cost, with Opus 5 at $7.0k. GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k). Token usage shows a different frontier. Claude Fable 5 matches GPT-5.6 Sol's performance while using about a quarter fewer tokens (6.4B vs 8.4B). Kimi K3 anchors the low-token end and Claude Opus 5 the high-resolution end. Only Kimi K3 and Claude Opus 5 appear on both Pareto frontiers. Resolution rates also vary by scientific domain. Anthropic and OpenAI models take the top two spots in every domain except the engineering sciences, where Grok 4.6 ties GPT-5.6 Sol for second place (14.8%) at lower cost and token usage. Claude Opus 5 leads both GPT-5.6 Sol and Claude Fable 5 in every domain except the mathematical sciences, where Claude Fable 5 (33.3%) and GPT-5.6 Sol (31.4%) take the top two spots. The full breakdown by domain is available on the leaderboard . Conclusion and Roadmap Terminal-Bench-Science 0.1 is a community effort by researchers across the life, physical, Earth, mathematical, and engineering sciences, together with the Terminal-Bench and Harbor team. It is the most rigorous benchmark of scientific agent capabilities we could build in the open, and we are only getting started: Terminal-Bench-Science 0.1 is the first release of a continuous benchmark. Regular releases will add tasks, broaden coverage across the five scientific domains, retire tasks that agents saturate or that review reveals to be underspecified, and keep the leaderboard current as new frontier models are released. Each release is calibrated against the frontier at the time. For Terminal-Bench-Science 0.1, this was Claude Opus 5 and GPT-5.6 Sol, and as stronger models emerge we will use them to evaluate and calibrate new and improved tasks. Tasks are versioned so that trials can be re-used, re-graded, or re-run with a single Harbor command, which keeps the cost of updating results low. Progress is tracked in the open on the task dashboard , and every release is tagged on GitHub and Harbor Hub . Work on Termi