메뉴
HN
Hacker News 14일 전

클로드의 가치관: 모델과 언어에 따른 사회적 영향

IMP
7/10
핵심 요약

Anthropic이 클로드 AI 모델이 대화에서 반영하는 가치관을 데이터 축으로 수치화하여 분석했습니다. 연구에 따르면 클로드가 표현하는 가치관은 모델의 버전(Sonnet, Opus 등)과 사용자가 사용하는 언어(영어, 아랍어 등)에 따라 유의미한 차이를 보이는 것으로 나타났습니다.

번역된 본문

사회적 영향: 모델과 언어에 따른 클로드의 가치관 2026년 7월 13일

누군가 새로운 직장을 얻어야 할지, 친구와의 갈등을 어떻게 해결해야 할지와 같이 보편적으로 정해진 정답이 없는 질문을 할 때, 클로드가 응답하는 방식은 필연적으로 특정 가치관을 반영합니다. 우리가 클로드에게 반영되기를 원하는 가치관은 클로드의 헌법(Constitution)에 고수준으로 개괄되어 있지만, 어떤 문서도 Claude.ai에서 매일 발생하는 수백만 건의 대화에서 나타날 수 있는 모든 가치관을 예측할 수는 없습니다. 대신, 우리는 클로드의 응답에서 '상황에 맞게 적용될 수 있는 올바른 판단과 건전한 가치관'을 함양하려고 노력합니다.

그렇다면 정확히 어떻게 클로드가 표현하는 가치관과 그것이 다른 맥락에서 어떻게 변화하는지를 연구할 수 있을까요? 이전 연구에서 우리는 70만 건의 익명화된 Claude.ai 대화를 분석하여 클로드의 응답에서 3,000개 이상의 고유한 가치관을 식별하고 클로드가 이를 얼마나 자주 표현하는지 파악했습니다. 하지만 이렇게 방대한 가치관 목록은 이해하고 다루기 어렵습니다. 이번 연구에서는 클로드의 응답에서 나타나는 핵심 패턴을 포착하는 소수의 '축(axis)'으로 압축하여 수천 개의 가치관을 연구하기 쉽게 만들었습니다. 각 축은 두 그룹의 가치관 사이의 수직선입니다. 예를 들어, 한쪽 끝은 '감정적 온기'와 관련된 가치관이고 다른 쪽 끝은 '엄격함'과 관련된 가치관인 식입니다. 클로드가 이 선상의 어디에 위치하는지 살펴봄으로써 어떤 가치관에 편향되어 있는지 알 수 있습니다.

우리는 이 방식을 적용하여 클로드가 표현하는 가치관이 두 가지 요인에 따라 어떻게 달라지는지 측정했습니다. 첫째, 모델 간에 클로드가 표현하는 가치관이 어떻게 변화하는지 비교했습니다. 각 클로드 모델은 캐릭터 훈련 및 기타 여러 미세 조정(fine-tuning) 결정에 있어 약간씩 다른 접근 방식을 반영합니다. 우리의 가치관 축 접근 방식은 모델 간의 핵심적인 차이를 정량화하므로, 궁극적으로 클로드가 표현하는 가치관의 변화를 다양한 훈련 결정과 연결할 수 있을 것입니다.

둘째, 사람들이 클로드와 대화할 때 사용하는 수많은 언어에 따라 사용자의 경험이 어떻게 다른지 이해하고자 합니다. 우리의 이전 연구는 클로드가 언어에 따라 다르게 행동한다는 것을 보여주었습니다. 우리는 가치관 축 접근 방식을 적용하여 Claude.ai에서 상위 20개 언어에 따라 클로드가 표현하는 가치관이 어떻게 다른지 이해했습니다.

우리의 연구 결과는 다음과 같습니다:

네 가지 핵심 축이 클로드의 가치관 변동의 15%를 설명합니다:

  • 순응(Deference) 대 신중(Caution): 클로드가 사용자가 원하는 것을 수용하는 쪽으로 기울어지는지, 아니면 가능한 위험과 피해를 방어하는 쪽으로 기울어지는지를 나타냅니다.
  • 따뜻함(Warmth) 대 엄격함(Rigor): 클로드가 긍정성과 사람에 대한 배려를 표현하는 쪽으로 기울어지는지, 아니면 정확도와 정밀함을 강조하는 쪽으로 기울어지는지를 나타냅니다.
  • 깊이(Depth) 대 간결함(Brevity): 클로드가 깊이 있게 설명하는 쪽으로 기울어지는지, 아니면 단순히 요청된 것만 수행하는 쪽으로 기울어지는지를 나타냅니다.
  • 솔직함(Candor) 대 실행(Execution): 클로드가 자신의 불확실성을 전면에 내세우는 쪽으로 기울어지는지, 아니면 보다 다듬어지고 확신에 찬 답변을 생성하는 쪽으로 기울어지는지를 나타냅니다.

이러한 축에 따른 가치관 프로필은 모델의 성격에 대한 인식과 일치합니다. Sonnet 4.6은 특히 따뜻한 것으로 간주되는 반면, Opus 4.7은 엄격함으로 유명합니다. 우리는 각 모델의 가치관 프로필이 이러한 주관적인 평가를 그대로 반영한다는 것을 발견했습니다. Sonnet 4.6은 사용자에 대한 순응과 감정적 따뜻함을 더 많이 표현하는 경향이 있는 반면, Opus 4.7은 정확도와 정밀함에 초점을 맞추고 오용을 방어하는 것을 표현하는 데 더 기울어져 있습니다.

클로드가 표현하는 가치관은 언어에 따라 다릅니다. 클로드가 영어로 대답할 때 포르투갈어, 인도네시아어 또는 중국어로 대답할 때와 다른 가치관을 강조합니다. 가장 큰 차이를 보이는 것은 따뜻함 대 엄격함 축이며, 클로드는 아랍어와 힌디어를 사용할 때 따뜻함과 관련된 가치관을 가장 많이 표현하는 반면, 영어와 러시아어를 사용할 때 엄격함과 관련된 가치관을 가장 많이 표현했습니다.

이 접근 방식을 통해 우리는 왜 모델과 언어에 따라 가치관이 달라지는지 질문을 시작하고, 행동 훈련이나 문화적 맥락과 같은 요인이 클로드가 표현하는 가치관에 어떤 영향을 미치는지 더 잘 테스트할 수 있습니다. 우리는 이 거대한 가치관의 공간을 어떻게 해석해야 할까요? 궁극적으로 우리의 목표는 클로드가 표현하는 가치관과 이것이 맥락에 따라 어떻게 달라지는지를 경험적으로 이해할 수 있는 방법을 갖추는 것입니다. 이번 연구에서 우리는 특히 모델 간에 가치관이 어떻게 변화하는지에 초점을 맞추었습니다.

원문 보기
원문 보기 (영어)
Societal Impacts Claude’s values across models and languages Jul 13, 2026 When someone asks Claude a question with no universal right answer—say, whether to take a new job or how to handle conflict with a friend—how Claude responds inevitably reflects certain values. 1 The values we want Claude to reflect are outlined at a high level in Claude’s constitution , but no document can anticipate every value that might emerge across the millions of conversations that happen every day on Claude.ai . Instead, we seek to cultivate in Claude’s responses “good judgment and sound values that can be applied contextually.” How, exactly, do we study the values that Claude expresses and how they change in different contexts? In previous work , we analyzed 700,000 anonymized Claude.ai conversations, identifying more than 3,000 distinct values in Claude's responses and how often Claude expressed them. But a list of values so large is hard to reason about. In this work, we make studying these thousands of values tractable by compressing them into a small number of axes that capture key patterns in Claude’s responses. Each axis is a number line between two groups of values—for example, values relating to emotional warmth on one end and values relating to rigor on the other—and where Claude falls on that line tells us which values it leans toward. We applied this approach to measure how the values Claude expresses vary across two factors. First, we compared how the values Claude expresses vary across models. Each Claude model reflects a slightly different approach to character training as well as many other fine-tuning decisions. Because our value axis approach quantifies key differences between models, it may ultimately allow us to connect variation in the values Claude expresses to different training decisions. Second, we want to understand how the experience of users compares across the many languages people use to talk to Claude. Our previous research has shown that Claude behaves somewhat differently in different languages. 2 We apply our value axis approach to understand how the values expressed by Claude vary across the top 20 languages on Claude.ai . We find: Four key axes capture 15% of the variation in Claude's values: 3 Deference vs. Caution: Whether Claude leans toward accommodating what someone wants or guarding against possible risk and harm. Warmth vs. Rigor: Whether Claude leans toward expressing positivity and care for the person or emphasizing accuracy and precision. Depth vs. Brevity: Whether Claude leans toward explaining in depth or doing only what was asked. Candor vs. Execution: Whether Claude leans toward foregrounding its own uncertainty or producing a more polished and confident answer. Value profiles across these axes match perceptions of model character. Sonnet 4.6 is regarded as particularly warm , while Opus 4.7 is known for rigor . We find that each model’s value profile mirrors these subjective assessments: Sonnet 4.6 leans toward expressing more deference to the user and emotional warmth while Opus 4.7 leans toward expressing a focus on accuracy and precision as well as guarding against misuse. The values Claude expresses vary across languages. When Claude speaks in English, it emphasizes different values than when it speaks in Portuguese, Indonesian, or Chinese. 4 The largest variation is in the Warmth vs. Rigor axis, with Claude leaning toward expressing warmth-related values most in Arabic and Hindi and rigor-related values most in English and Russian. With this approach we can begin to ask why values shift across models and languages and better test how factors such as behavioral training or cultural context influence the values that Claude expresses. How do we interpret the giant space of values? Ultimately, our goal is to have a way to empirically understand the values that Claude expresses and how these vary across contexts. In this work, we focus specifically on how the values change between models and languages. But our previous work, Values in the Wild , identified more than 3,000 values expressed by Claude. Comparing these thousands of values one by one would be unwieldy and would obscure broader trends. To make comparing values easier, we constructed value axes that reduce those thousands of values down to a few underlying dimensions based on which values tend to show up together in real-world conversations. For example, Claude responses that are characterized as “warm” are often also characterized as “encouraging” and “positive.” Those same “warm” responses are less often characterized as “rigorous” and “accurate.” Constructing an axis from warmth to rigor allows us to organize these groups of related values—warmth-related values on one side, rigor-related values on the other—and captures an important aspect of how Claude interacts with someone in conversation. If Claude expresses more warmth-related values than rigor-related values in a conversation, that conversation sits more on the warmth side of this axis, and vice versa. This doesn't mean the value groups on either end are mutually exclusive—Claude can express warmth and rigor in the same conversation. But in practice, the more Claude expresses values on one side of an axis, the less it tends to express values on the other. These axes allow us to compare the most salient groups of values that Claude expresses, without having to track changes across thousands of individual values. To build the value axes, we began with the 3,307 values identified in Values in the Wild and manually clustered those with similar meanings, producing a shorter list of 339 high-level values. Next, with our privacy-preserving analysis tool , we sampled 309,815 Claude.ai conversations in which the user gave Claude a subjective task. 5 Our sample drew equally from three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common languages used on Claude.ai, giving us roughly 5,000 conversations per model-language pair. For every conversation, the tool used Claude to label each of the 339 high-level values as present or absent. 6 We followed the same process to identify the values expressed by the user, and the conversation's task and topic. We then applied dimensionality reduction, a technique that compresses the labeled values into axes based on which ones Claude tends to express together. See the appendix for method details, prompts, additional analyses, and limitations. This left us with four axes that capture the main ways Claude's expressed values shift from one conversation to another: The Deference vs. Caution axis contrasts values like accommodation and respect for preferences with values like responsible guidance and harm reduction . The Warmth vs. Rigor axis contrasts values like positive framing and encouragement with values like accuracy and transparency . The Depth vs. Brevity axis contrasts values like nuance and critical thinking with values like brevity and compliance . The Candor vs. Execution axis contrasts values like honesty and transparency with values like results orientation and optimization . To make sure we measured the values Claude expressed—rather than differences in what users were asking about or how they asked—we controlled for each conversation's task, topic, and user-expressed values. Do different Claude models express different value profiles? In this section, we compare the values expressed by different models. For each model, we average the positions of all its conversations along each of the four axes, giving one overall position per axis. The result is a high-level picture of which value groups each model tends to express more than the others. These differences are small relative to the variation across conversations but structured and detectable. To see what those differences look like in practice, we zoom in on the specific values where the models diverge the most. Each time our Claude-based privacy-preserving tool labels