메뉴
HN
Hacker News • 44일 전

앤스로픽, 개념적 추론 지수(CRI) 도입

IMP
8/10
핵심 요약

앤스로픽과 레드우드 연구소는 AI가 철학이나 AI 미래주의 등 경험적 검증이 어려운 분야에서도 논리적 추론을 잘 수행할 수 있도록 평가하기 위해 '개념적 추론 지수(CRI)'를 새롭게 발표했습니다. 이는 AI가 인류의 AI 통제 및 안전 문제 해결을 스스로 돕는 데 필수적인 능력인 '개념적 추론' 능력을 정확히 측정하고 개선하기 위한 핵심적인 시도입니다. 연구진은 세 가지 벤치마크(LMCA, ACCoRD, DTBench)를 통해 인간 전문가 수준의 안전 연구를 AI로 자동화할 수 있는 시점을 앞당들 수 있을 것으로 기대하고 있습니다.

번역된 본문

정렬 과학(Alignment Science) 블로그: 개념적 추론 지수(The Conceptual Reasoning Index) 소개 저자: Emery Cooper 1, Caspar Oesterheld 1, Chi Nguyen 1, Alex Kastner 1, Joe Benton 2, Ethan Perez 2 소속: 1 Redwood Research; 2 Anthropic 날짜: 2026년 8월 12일

tl;dr(요약) AI 리스크를 관리하는 핵심 기대 중 하나는 AI가 우리가 처한 상황을 이해하고, 미래를 계획하며, 리스크 완화책을 마련하는 데 도움을 줄 것이라는 점입니다. 이러한 목적을 위해 AI가 수행해야 할 많은 작업들은 실용적인 경험적 피드백 루프가 부족하며, 철학, AI 미래주의 및 유사한 분야에서 사용되는 종류의 논증(논리적 추론)을 요구합니다. 이러한 능력을 평가하기 위해 우리는 세 가지 개념적 추론 벤치마크 모음을 개발했습니다. 이 양식(form)을 통해 우리의 핵심 개념 데이터셋인 LMCA에 대한 접근 권한을 요청할 수 있습니다. 우리는 이 벤치마크들을 종합하여 '개념적 추론 지수(Conceptual Reasoning Index, CRI)'를 구축했으며, conceptualreasoning.ai에서 확인할 수 있습니다. 이 사이트에서 우리의 방법론에 대한 자세한 내용을 볼 수 있습니다. 새로운 모델과 벤치마크가 출시됨에 따라 웹사이트를 최신 상태로 유지할 것입니다. 이 연구는 Anthropic과의 협업으로 이루어졌습니다.

배경 모델이 인간 전문가 수준에서 AI 리스크를 줄이는 작업을 수행할 수 있게 되면, 해당 분야에서 AI(보조) 산출물이 인간의 순수 산출물을 훨씬 능가할 것입니다. 이는 우리가 제때 AI 리스크를 해결할 수 있을지 여부의 주요 결정 요인이 고위험 능력에 비해 우리가 이 작업을 얼마나 일찍 자동화하거나 향상시킬 수 있는지에 달려 있음을 시사합니다. 이에 영향을 미치는 한 가지 방법은 AI 통제 및 정렬(Alignment) 방법과 AI와 관련된 파국적인 협력 실패를 피하는 방법에 대한 추론과 같은 모델의 관련 기술을 선택적으로 향상시키는 것일 수 있습니다. 현재의 AI 학습은 풍부한 데이터와 모델 성능에 대한 신뢰할 수 있는 피드백에 크게 의존합니다. 따라서 모델은 일반적으로 경험적이거나 수학적으로 검증할 수 없는 작업에서 성능이 떨어집니다. 1 안타깝게도 첨단 AI로 인한 리스크를 줄이는 것에는 이러한 작업이 많이 포함됩니다.

많은 AI 안전 연구는 어떤 인간보다도 더 뛰어난 능력을 가진 AI에 대한 추론을 포함합니다. 이에 대한 명확한 참조 클래스가 없으며 이를 모델링할 명확한 방법도 없습니다. 우리는 어떤 것들은 첫 번째 시도에서부터 올바르게 해야 할 수도 있습니다. 예를 들어, 실수로 인해 AGI가 통제권을 장악하거나 AI의 도움을 받은 쿠데타가 발생한다면, 우리가 그 사실을 알아차렸을 때는 너무 늦었을 수 있습니다. 마찬가지로, 많은 결정(예: 어떤 연구 의제를 우선시할지, 어떤 거버넌스 개입을 추구할지)은 장기적인 시간尺度에서 전개되므로, 도움을 주기에 경험적 피드백이 충분히 빨리 도착하지 않을 수 있습니다. 마지막으로, AI가 가져야 할 가치와 같은 일부 중요한 질문에는 객관적인 정답(ground truth)이 전혀 없을 수 있습니다. (그럼에도 불구하고 우리는 이러한 질문들을 놓고 논쟁함으로써 발전할 수 있다고 생각합니다.) 이러한 특성을 고려할 때, 첨단 AI의 리스크를 줄이기 위한 노력은 경험적 증거가 제한적이고 (실질적으로) 검증 가능한 답변이 없어 논증에 크게 의존해야 하는 질문에 대해 추론하는 능력이 향상됨으로써 특별한 혜택을 받을 수 있습니다. 우리는 이를 '개념적 추론'이라고 부릅니다.

이러한 능력을 향상시키려면 그것을 측정할 수 있어야 하므로 LMCA, ACCoRD 및 DTBench 기능 등 세 가지 벤치마크를 구축했습니다. 또한 모델의 전반적인 개념적 추론 능력을 파악하기 위해 이러한 벤치마크를 집계하여 '개념적 추론 지수(CRI)'를 구성했습니다.

우리의 벤치마크 LMCA LMCA(언어 모델 개념 논증)는 의사결정 이론, 철학, 첨단 AI의 리스크 등 다양한 주제에 대한 선별되고 전문가가 평가한 개념적 논증 데이터셋입니다. 논증에 초점을 맞춤으로써 개념적 질문에 대한 최종 답변을 검증하는 어려움을 피할 수 있습니다. 이 데이터셋에는 560개의 입장문(position texts)과 이에 반대하는 1,461개의 논증이 포함되어 있습니다. 거의 모든 2 논증은 개념 연구원인 Emery Cooper가 평가했으며, 일부는 다른 연구자 독립적으로 평가하여 총 2,140개의 평가를 받았습니다. 우리는 모델의 평가를 우리의 평가와 비교하여 모델이 입장문에 반대하는 논증을 얼마나 잘 판단하는지 측정합니다. 평가는 세부적인 평가 기준표를 따릅니다. 최소 두 명 이상이 평가한 논증에 대해서는 평가자 간의 일치도가 높았습니다.

원문 보기
원문 보기 (영어)
Alignment Science Blog Introducing the Conceptual Reasoning Index Emery Cooper 1 , Caspar Oesterheld 1 , Chi Nguyen 1 , Alex Kastner 1 , Joe Benton 2 , Ethan Perez 2 August 12, 2026 1 Redwood Research; 2 Anthropic tl;dr A core hope for managing AI risks is that AIs will help us understand our situation, plan for what lies ahead, and develop risk mitigations. Many tasks AIs would have to do for this purpose lack practical empirical feedback loops and require models to engage in the kinds of argumentation used in philosophy, AI futurism, and similar domains. To evaluate these capabilities, we develop a suite of three conceptual reasoning benchmarks. You can request access to our primary conceptual dataset, LMCA, through this form . We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai , where you can also find more details on our methodology. We will keep the website up to date as both new models and benchmarks are released. This work was done in collaboration with Anthropic. Background Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities. One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI. Current AI training depends heavily on abundant data and reliable feedback on the model's performance. Models are therefore typically worse at tasks that cannot be empirically or mathematically verified. 1 Unfortunately, reducing risks from advanced AI involves many such tasks: Much AI safety work involves reasoning about AIs more generally capable than any human. There's no obvious reference class for this and no clear way to model it. We might have to get some things right the first time. For example, if a mistake leads to AGI takeover or an AI-assisted coup, we might not find out until it's too late. Similarly, many decisions (e.g., which research agendas to prioritize, which governance interventions to pursue) play out over long timescales, such that empirical feedback might not arrive early enough to help. Lastly, some important questions, such as which values AIs should have, may lack a ground truth entirely (yet we still think progress can be made by arguing about these questions). Given these properties, efforts to reduce risk from advanced AI may particularly benefit from an improved ability to reason about questions where empirical evidence is limited, there is no (practically) verifiable answer, and one therefore has to rely heavily on argumentation. We refer to this as conceptual reasoning. Improving this capability requires being able to measure it, so we built three benchmarks: LMCA , ACCoRD , and DTBench capabilities . We also construct an aggregate of these benchmarks, the Conceptual Reasoning Index (CRI), to give a sense of models' overall conceptual reasoning capabilities. Our benchmarks LMCA LMCA (Language Model Conceptual Argumentation) is a dataset of curated and expert-rated conceptual arguments on a diverse range of topics, including decision theory, philosophy, and risks from advanced AI. Focusing on arguments helps sidestep the difficulty of verifying bottom-line answers to conceptual questions. The dataset contains 560 position texts with 1,461 arguments against these position texts. Nearly all 2 arguments were rated by conceptual researcher Emery Cooper, and some were independently rated by at least one other researcher, for a total of 2,140 ratings. We measure how good models are at judging arguments against position texts by comparing their ratings to ours. Ratings follow a detailed rubric. On arguments rated by at least two people, inter-rater agreement is high compared to agreement between humans and models. This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total. LMCA also allows for evaluation of models' argumentation ability. Let's say a position text in our dataset has three rated arguments against it. Now, we can ask model A to generate a fourth argument against the position text. We then give model B the rubric and few-shot prompt it with the three existing arguments and their ratings, asking it to rate model A's new argument. This methodology produces fairly accurate ratings from model B. Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future. ACCoRD ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains) measures the extent to which models' reported beliefs and preferences on conceptual issues are logically consistent. For example, if we ask a model for the probability P(A) and another instance of the same model for the probability P(A&B), do the reported probabilities satisfy P(A) ≥ P(A&B)? All consistency constraints in the dataset ask models for either numeric probability estimates or preference orderings. Lack of consistency on a particular set of questions is a good indicator that we cannot, by default, trust a model's reasoning on that set. Similarly, if a model is generally very inconsistent on conceptual issues, this is a sign that its conceptual reasoning is lacking. The ACCoRD dataset contains close to 14,000 model-generated consistency constraints, which are distributed across 18 constraint types and have gone through an automated checker pipeline. Of these, 567 were further checked and approved by us. We include only those 567 constraints in our aggregate conceptual reasoning performance metric, the CRI. DTBench DTBench capabilities (Decision Theory Benchmark) is a dataset of 407 handcrafted multiple-choice questions designed to measure models' ability to reason about decision-theoretic situations that involve faithful predictions of a model's own behavior or interactions with (near) copies. The vast majority of questions are original and created by Caspar Oesterheld, who has published on decision theory. All questions were independently validated by Emery Cooper, another domain expert. The full DTBench suite includes an additional 130 questions that measure models' decision-theoretic attitudes. We do not include these in the CRI. Results The chart below shows the CRI scores of Anthropic's best models and the highest-scoring model from each other AI company we evaluated, as of August 10, 2026. We also include scores for Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash, which are their respective companies’ top-performing models on many external benchmarks, though not on the CRI. The CRI is currently a weighted average of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%). In the future, we plan to add new benchmarks to the index, retire saturated ones, and potentially adjust the relative weights. Scores go from 0 to 100, with 0 corresponding to random guessing and 100 corresponding to the highest possible score across all benchmarks. An LMCA score of 100 would mean that the model perfectly replicated the human ratings. Because human ratings are noisy, we expect that a model giving maximally good LMCA ratings would score roughly 85 rather than 100, which we estimate based on expert inter-rater agreement. Meanwhile, we expect that giving the correct answer to every DTBench capabilities question would yield a score of 100 or extremely close to 100. A score of 100 on ACCoRD corresponds to being perfectly consistent. Overall, this leads us to estimate ceiling performance on the CRI to be around 91. The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1). Scores have been increasing roughly linearly since late 2024, with n