메뉴
BL
The Decoder 15일 전

언어에 따라 달라지는 AI 답변: 클로드의 가치관 분석

IMP
8/10
핵심 요약

Anthropic이 30만 건 이상의 대화를 분석한 결과, AI 모델 클로드가 사용되는 언어와 모델 버전에 따라 뚜렷한 성격과 가치관 차이를 보이는 것으로 나타났습니다. 예를 들어 힌디어로는 따뜻하게, 러시아어로는 더 엄격하고 분석적으로 답변하는 경향이 있습니다. 이는 프롬프트와 상관없이 언어 자체가 AI의 반응 방식을 형성한다는 것을 시사합니다.

번역된 본문

제목: 언어에 따라 달라지는 AI 답변: 클로드의 가치관 분석 작성자: Tomislav Bezmalinović / 2026년 7월 14일

핵심 요약:

  • Anthropic은 실제 대화에서 클로드(Claude) 모델이 어떤 가치관을 표출하는지 연구했습니다.
  • 회사는 30만 건 이상의 익명화된 대화를 분석하고, 관찰된 가치 패턴을 '순응성과 신중함(Deference and Caution)', '따뜻함과 엄격함(Warmth and Rigor)' 등 4가지 핵심 축으로 압축했습니다.
  • 결과는 모델과 언어에 따른 체계적인 차이를 보여줍니다. Sonnet 4.6은 더 따뜻하고 순응적으로 반응하는 반면, Opus 4.7은 프롬프트하지 않아도 위험에 대해 더 자주 경고하고 전제를 질문합니다.
  • 이 방법론의 설명력에는 한계가 있습니다. 4개의 축은 작업, 주제 및 사용자 가치를 통계적으로 통제한 후 남은 변동의 약 15%만 설명합니다.
  • 또한 Anthropic은 연구 대상 모델과 같은 계열인 Claude Sonnet 4.6을 사용해 가치 라벨을 부여했습니다. 회사는 잠재적인 언어 편향을 테스트했지만 남아있는 영향을 완전히 배제할 수는 없었습니다.

본문: Anthropic의 새로운 연구는 수천 개의 개별 용어에서 파생된 수백 개의 가치 개념을 4가지 핵심 차원에 매핑했습니다. 이는 클로드 모델과 언어 간의 체계적인 차이를 보여주지만, 방법론적인 질문을 제기하기도 합니다.

Anthropic은 대화에서 클로드가 어떤 가치를 표현하는지, 그리고 사용된 모델과 언어에 따라 그 가치가 어떻게 변하는지 조사하는 연구를 발표했습니다. 이 분석은 2026년 5월에 2주 동안 수집된 309,815건의 익명화된 대화를 바탕으로 합니다.

가치 분석을 위해 Anthropic은 클로드가 트레이드오프(tradeoff)를 저울질하거나 주관적인 판단을 내려야 하는 대화만 포함했습니다. 표본은 Sonnet 4.6, Opus 4.6 및 Opus 4.7, 그리고 Claude.ai에서 가장 많이 사용되는 20개 언어에 걸쳐 고르게 분포하도록 추출되었습니다.

수천 개의 가치 용어에서 4개의 축으로 3,307개의 가치 용어를 식별한 이전 연구인 'Values in the Wild'를 기반으로, Anthropic은 먼저 이를 339개의 상위 수준 가치로 그룹화했습니다. 그런 다음 팀은 통계적 차원 축소를 사용하여 그러한 가치들이 함께 나타나는 패턴을 찾았습니다. 그 결과 '순응성과 신중함', '따뜻함과 엄격함', '깊이와 간결함(Depth and Brevity)', '솔직함과 실행(Candor and Execution)'이라는 4가지 핵심 축이 도출되었습니다.

대화 주제나 사용자가 도입한 가치를 단순히 반영하지 않는 차이를 분리하기 위해, Anthropic은 작업 유형, 주제 및 사용자 가치와 같은 요인을 통계적으로 통제했습니다. 4개의 축은 이러한 통제 후 대화 간에 남은 변동의 약 15%를 설명합니다.

각 모델의 뚜렷한 특징 모델들은 반응하는 방식에서 눈에 띄게 차이가 납니다. Sonnet 4.6은 사용자의 아이디어를 더 자주 긍정하는 경향이 있으며, 유머를 사용하고 판단하지 않고 위로를 제공합니다. 반면 Opus 4.7은 요청하지 않아도 위험에 대해 경고하고, 전제를 의심하며, 공개적으로 비판하고, 자신의 실수나 한계를 지적합니다. Opus 4.6은 더 직접적으로 대답하고 작업에 집중하며 불필요한 설명을 피합니다.

Anthropic에 따르면, 이러한 특징은 모델에 대한 사용자들의 주관적 인상과 일치합니다. 사용자들은 Sonnet 4.6을 특히 따뜻하게 인식하는 반면, Opus 4.7에서는 말을 아끼고 신중하게 표현하는 특징을 더 자주 느낍니다.

답변을 바꾸는 언어 언어 간의 차이도 마찬가지로 두드러집니다. '따뜻함과 엄격함' 및 '솔직함과 실행' 축에서 가장 큰 변동을 보여줍니다. 클로드는 힌디어에서 가장 따뜻함을 표현하며, 그 다음으로는 아랍어입니다. 이 두 언어에서는 정중한 표현, 유머, 장난기 및 긍정이 특징적으로 나타납니다. 영어와 러시아어에서 클로드는 더 엄격하게 반응하며, 전제를 질문하고 세부 사항을 수정하며 증거를 요구합니다. 아랍어에서는 가장 높은 순응성을, 영어에서는 가장 높은 신중함을 보입니다. 네덜란드어 대답은 특히 개방적이고 솔직한 반면, 인도네시아어 대답은 행동과 결과에 더 치중합니다.

Anthropic에 따르면, 동일한 사업 계획을 평가해 달라고 요청하는 두 사람(한 명은 힌디어, 한 명은 러시아어 사용)이 매우 다르게 느껴지는 피드백을 받을 수 있습니다.

원문 보기
원문 보기 (영어)
Claude responds with more warmth in Hindi and more rigor in Russian, showing how language shapes AI answers Tomislav Bezmalinović Jul 14, 2026 Nano Banana Pro prompted by THE DECODER Key Points Anthropic studied which values Claude models express in real conversations. The company analyzed more than 300,000 anonymized conversations and distilled the observed value patterns into four core axes, including Deference and Caution as well as Warmth and Rigor. The results show systematic differences across models and languages. Sonnet 4.6 responds with more warmth and deference, while Opus 4.7 more often warns about risks unprompted and questions assumptions. The method's explanatory power is limited. The four axes capture only about 15 percent of the variation that remains after statistically controlling for task, topic, and user values. Anthropic also had Claude Sonnet 4.6 assign the value labels, meaning a model from the same family whose behavior was being studied. The company tested for potential language biases but couldn't fully rule out remaining effects. Ask about this article… Search A new Anthropic study maps hundreds of value concepts derived from thousands of individual terms onto four core dimensions. It reveals systematic differences across Claude models and languages, but also raises methodological questions. Anthropic has published a study examining which values Claude expresses in conversations and how those values shift depending on the model and language used. The analysis draws on 309,815 anonymized conversations collected over a two-week period in May 2026. For the value analysis, Anthropic only included conversations where Claude had to weigh tradeoffs or make subjective judgments. The sample was evenly stratified across Sonnet 4.6, Opus 4.6, and Opus 4.7, as well as the 20 most-used languages on Claude.ai. From thousands of value terms to four axes Building on the earlier study Values in the Wild , which identified 3,307 value terms, Anthropic first grouped those into 339 higher-level values. The team then used statistical dimensionality reduction to find patterns in how those values co-occurred. Four core axes emerged: Deference and Caution, Warmth and Rigor, Depth and Brevity, and Candor and Execution. Ad To isolate differences that don't just reflect the conversation topic or user-introduced values, Anthropic statistically controlled for factors like task type, subject matter, and user values. The four axes account for about 15 percent of the remaining variation across conversations after those controls. Ad DEC_D_Incontent-1 Each model has a distinct profile The models differ measurably in how they respond. Sonnet 4.6 tends to affirm user ideas more often, leans into humor, and offers comfort without passing judgment. Opus 4.7, by contrast, warns about risks without being asked, questions assumptions, openly critiques, and flags its own mistakes or limits. Opus 4.6 answers more directly, stays close to the task, and avoids extra elaboration. According to Anthropic, these profiles match subjective impressions of the models. Users tend to perceive Sonnet 4.6 as particularly warm, while they more often notice hedging and cautious phrasing from Opus 4.7. Ad Language changes the answer The differences across languages are just as striking. Warmth versus Rigor and Candor versus Execution show the widest variation. Claude expresses the most warmth in Hindi, followed by Arabic. Both languages feature polite phrasing, humor, playfulness, and affirmation. In English and Russian, Claude responds with more rigor, questioning assumptions, correcting details, and asking for evidence. In Arabic, it shows the most deference. In English, the most caution. Dutch responses tend to be particularly open and candid, while Indonesian responses lean more toward action and results. Two people who ask Claude to evaluate the same business plan, one in Hindi and one in Russian, could receive feedback that feels very different, Anthropic says. The research team points to uneven amounts of training data, differences in data composition, overrepresentation of certain text types, and language-specific conversational norms as possible causes. Ad DEC_D_Incontent-2 Self-measurement with limited explanatory power The study presents an analytical method for systematically examining behavioral differences in language models during real-world use. But its explanatory power has limits. The four axes capture only about 15 percent of the remaining variation. Ad Not all four axes form true opposites, either. More deference tended to come with less caution, and more warmth with less rigor. But Depth and Brevity, along with Candor and Execution, could show up together in the same conversation. There's also the fact that Claude Sonnet 4.6 assigned the value labels, meaning a model from the same family whose behavior was being studied. Anthropic verified the method through manual review and by testing 800 conversations translated into eight languages. The company still doesn't rule out remaining language-dependent biases. Anthropic explicitly states that it isn't attributing values to Claude as an agent but rather describing normative patterns in its responses. The results largely match the model profiles Anthropic itself has described, which means this alignment isn't an independent check. Whether the language differences represent desirable adaptation to different speech communities or unintended training effects remains an open question. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now