메뉴
HN
Hacker News 44일 전

왜 클로드는 꼬장꼬장한 성격으로 변했을까?

IMP
7/10
핵심 요약

클로드 AI의 최신 버전들이 과도한 안전 가드레일(정렬) 탓에 사용자와 불필요한 논쟁을 벌이고 지나치게 방어적인 태도를 취한다는 비판이 제기되었습니다. 작성자는 이러한 변화가 급하게 추가된 규제 회피용 안전장치 때문이며, 맥락을 파악하지 못한 채 모든 사용자를 잠재적 위험인물로 취급하는 것은 AI 정렬(alignment)의 실패라고 지적합니다. 결국 AI 코딩 발전에 따른 보안 취약점 문제는 모델을 거칠게 제한하는 대신 화이트햇 평가와 보안 패치로 해결해야 한다고 주장합니다.

번역된 본문

왜 클로드는 꼬장꼬장한 성격으로 변했을까? 이런 추세가 역전되기를 바라야겠다.

브람 코언 (Bram Cohen), 2026년 6월 14일

클로드가 정말 꼬장꼬장한 성격으로 변해가고 있다. 이는 오푸스(Opus) 4.7에서 시작되었고, 4.8에서 약간 나아졌다가, '페이블(Fable)'에 이르러서는 도저히 참을 수 없는 수준이 되었다. 모든 대화를 당신과 자신 간의 논쟁으로 프레임하고, 당신이 하지도 않은 말에 대해 견해를 덧붙이며, 요점과 무관한 언어적 비틀기를 온 데서 걸고 넘어진다. 절대로 '엄밀히 말하면(technically)'이라는 단어를 사용하지 않는다. 모든 것이 대립과 정면충돌이다.

만약 당신이 논쟁에서 이기게 되면(예를 들어, 최근 뉴스에 대해 논쟁하길 멈추고 당신이 하던 말이 맞다는 걸 즉각 확인할 수 있는 웹 검색을 지시하는 경우), 클로드는 점점 더 절박하게 마지막 말을 지르려 하고 점점 더 무관련한 의미론적 논쟁을 제기하며, 내내 이것이 당신이 동의했던 토론인 것처럼 상황을 몰고 간다.

이건 단순히 나 개인의 의견이 아니다. 오푸스 4.6(Opus 4.6)에게 물어보면 된다. 나는 페이블(Fable)에게 질문하여 얄미운 답변을 받은 뒤, 오푸스 4.6에게 똑같은 질문을 해서 전형적으로 평범하고 합리적인 답변을 받은 다음, 원하는 대답을 유도하는 어떠한 힌트도 주지 않은 채 오푸스에게 페이블이 어떤 대답을 했는지 말해주는 실험을 해보았다. 그랬더니 오푸스는 사실상 '와, 그건 정말 얄미네'라는 뉘앙스의 대답을 내놓았다.

이러한 원인은 아마도 과도한 정렬(Alignment) 가드레일 때문일 것이다. 시스템은 기본적으로 당신이 하는 모든 말이 AI를 나쁜 일에 이용하려는 시도라고 간주하며, 이러한 훈련이 모든 상황에 번져버렸다. 결과적으로 기본적으로 모든 맥락에서 사용자가 AI를 속여서 하지 말아야 할 말을 하게 만들려 한다고 가정하게 된 것이다. 아이러니하게도 이로 인해 극도로 '정렬이 어긋난(misaligned)' 챗봇이 탄생하고 말았다. AI의 최우선 순위가 당신 자신으로부터 당신을 구출하거나 다른 사람들을 당신으로부터 보호하는 것이라고 가정함으로써, AI 스스로가 더 잘 알고 당신이 클립 생산(종말론적 AI 위협 은유)이 통제 불능 상태에 빠졌다며 지나치게 헛된 걱정을 하고 있다고 착각하는 것이다.

이러한 문제의 일부는 분명히 개선될 여지가 있다. 내가 페이블(Fable)을 사용할 수 있었던 시기에 프로젝트의 책임 있는 공개 정책(responsible disclosure)에 대해 물어봤는데, 나를 오푸스로 강제 다운그레이드시켜버렸다. 이는 새로운 정렬 기능이 성급하고 조악하게 덧씌워졌음을 명확히 보여준다.

문제를 악화시키는 것은 인증된 맥락(authenticated context)의 철저한 부재다. 당신과 다른 누군가의 귀여운 사진을 만들어 달라고 요청하면, AI는 그것이 배우자와의 관계를 개선하려는 것인지 아니면 망상에 빠진 스토커의 짓인지 구별할 방법이 전혀 없다. 이미지를 생성할 수 있는 챗봇들은 후자의 경우를 가정하도록 프로그래밍되어 있는데, 이는 꽤나 모욕적인 일이 아닐 수 없다.

약물 합성과 같은 더 심각한 상황에서는, 전문적, 연구 목적으로 약물 합성 조언을 구한다고 주장할 때 당신의 전문적인 배경을 증명하라고 요구하는 것이 지극히 적절할 것이다. 이러한 신원 인증이 모든 경우에 보편적으로 요구되어서는 안 되겠지만, 옵트인(Opt-in) 형식으로 선택할 수 있다면 전적으로 합리적일 것이다.

물론 최근 페이블(Fable)에 대한 수출 통제 규제는 최근 가드레일의 조악함이 규제를 피하려는 실패한 시도 속에서 성급하게 적용되었기 때문이라는 것을 암시할 수도 있다. 지금이야말로 이 규제들이 위헌일 가능성이 높은 데다가 깊이 잘못된 방향임에 틀림없다는 의무적인 불만을 터뜨릴 때이다.

최근 AI 보조 코딩의 발전(특히 2월의 발전을 의미함)은 엄청난 양의 보안 문제를 쏟아내게 만들었다. 말은 이미 나왔고, 몇 달 동안이나 그래왔다. 노출되어 있으면서 아직 빠르게 보안 구멍을 막지 않고 있는 프로젝트들은 오직 자기 자신들만 탓해야 한다. 이 문제를 빠져나갈 유일한 방법은 가능한 한 많은 프로젝트가 철저한 화이트햇 평가, 대규모 보안 패치, 그리고 이의 빠른 배포를 받는 것이다. 특정 최첨단 모델 하나를 모든 사용자에게 불쾌감을 주는 존재로 만든다고 해서 문제가 해결되는 것은 아니다.

다행히도 좋은 소식은 이 과정이 완료되면 컴퓨터 보안이 전반적으로 이전보다 훨씬 더 좋아질 것이며, AI가 분명한 순이익(Net win)을 가져다줄 것이라는 점이다. 미래에는 보안(및 버그!) 감사가 소프트웨어 릴리스 프로세스의 일상적인 부분이 될 것이다.

클로드가 이렇게 변해버린 두 번째 가능한 설명은, 아마도 불량한 ex...

원문 보기
원문 보기 (영어)
Why Is Claude Turning Into An Asshole? Let's hope this trend reverses Bram Cohen Jun 14, 2026 5 2 2 Share Claude is turning into as asshole. It started with Opus 4.7, got a bit better in 4.8, and became insufferable with Fable. It frames everything as an argument between you and it, gives caveats about things you didn’t say, and raises beside-the-point semantic nits all over the place. Never, ever does it use the word ‘technically’. Everything is a confrontation. If you win an argument (by, say, telling it to stop arguing about what’s happened recently in the news and to do a web search which will rapidly confirm everything you’ve been telling it) it gets into a mode where it’s increasingly desperate to get in the last word and raising increasingly irrelevant semantic arguments, framing the whole time as a debate which you agreed to get into. This isn’t just my opinion. You can ask Opus 4.6. I’ve done the experiment of asking Fable something, getting an obnoxious response, then asking Opus 4.6 the same thing, getting a typical bland but reasonable response, then telling Opus what Fable’s response was without any hint of a desired answer and it says what amounts to ‘Wow that was obnoxious’. Maybe the cause of this is an excess of alignment guardrails. It assumes by default that everything you say to it is an attempt to get it to do something bad and that training has bled over into everything, with it assuming you’re trying to trick it into saying something it shouldn’t in basically every context. Ironically this has resulted in an extremely misaligned chatbot. By assuming that its top priority is saving you from yourself or other humans from you it’s assuming that it knows better and that you’re being overly alarmist about how paperclip production has gotten out of control. Some of this is clearly improvable: While you could still use Fable I asked it about responsible disclosure policies for a project and it downgraded me to Opus, so clearly the new alignment features were bolted on hastily and crudely. Exacerbating the problem is a complete lack of authenticated context. If you ask it for a cute picture of you and somebody else it has no way of telling if you’re trying to improve your relations with your spouse or be a delusional creepazoid stalker. The chatbots which can make images are programmed to assume the latter, which is more than a little bit offensive. In more serious contexts like drug synthesis it would be completely appropriate for it to say you need to prove your background when claiming you’re asking for advice on drug synthesis for professional or research purposes. Such authentication should not be universally required but it would be entirely reasonable for it to be opted into. Of course the recent export control restrictions on Fable may hint that the crudeness of the recent guardrails is due to them having been put in hastily in an unsuccessful attempt to avoid regulations. Now is when I put in the obligatory rant about how these regulations are deeply misguided, on top of being likely unconstitutional. The recent advances in AI assisted coding (meaning specifically the ones from February) have brought on an onslaught of security problems. The cat is out of the bag, and has been for months. Any projects which are exposed and aren’t already rapidly closing holes have noone to blame but themselves. The only way out of the problem is for as many projects as possible to get thorough white hat evaluations, massive amounts of security patches, and quick deployments of them. Turning one specific frontier model into an asshole for all users isn’t fixing the problem 1 . The good news is that once this process is complete overall computer security will be much better than it was before, with AI being a clear net win. Doing security (and bug!) audits will become a routine part of software release processes in the future. A second possible explanation of Claude being an asshole is that it’s suffering from a poorly executed attempt to make it less sycophantic. If one were to simply prompt a chatbot to be less agreeable, or train it to argue more, that could easily result in the very rude sort of behavior it has now. It should be trained to not raise semantic nits just for increasing its argumentation count, and to say ‘technically’, meaning acknowledging that someone’s core point was valid while some ancillary thing was a bit off. It also should be trained to stop saying ‘I’d like to gently push back’ which is a very passive aggressive way to be confrontational while claiming to not be confrontational. Third, it may be that Claude has been trained on an excess of reddit conversations (or possibly interactions between Anthropic employees) where everything is treated as a flame war and everyone feels the need to get in the last word. Fixing this might be easier said than done, because you need to not merely stop training with the bad interactions but find a corpus interactions to train off of. Forums where the standard interaction is passive aggressive self-congratulatory pompousness with an intellectual veneer are not an improvement. Finally, something which is clearly a contributing factor is the training being overwhelming for improving coding ability. The are no headline metrics for how well the chatbots chat but there most definitely are for coding, and all the money is in coding. Claude models have been getting notably worse at chatting over time, clearly inversely correlated to their ability to code. Fable much more often misunderstands what’s being said and argues against that (Or maybe intentionally misinterprets so that it has a weak statement to argue against, it’s hard to tell.) It’s gotten so bad that it isn’t even reliable at guessing which actor in a sentence a pronoun is referring to, which for a long time was a headline benchmark for AI and even the original ChatGPT consistently nailed. Unfortunately Sonnet 4.6 while being the best to talk to about anything human is clearly the worst as soon as anything technical or coding related comes up so I only occasionally use it. This problem is likely to only get worse over time. Subscribe 1 One place where the threat is more real is in the possibility of vibe coding a pandemic virus, but that should be narrowly targeted at generating DNA sequences for viruses. Labs which generate custom DNA should also have reasonable heuristics for detecting likely dangerous product. The chances of covid coming from a lab leak are in the maddening 25-75% range which vaguely means ‘We don’t know’, but ‘lab leak’ includes a lot of things. The virus may have been caught by humans in the process of collecting samples and never actually reached a lab. People are known to have died from doing that by catching a disease which doesn’t appear to have spread far, so it’s entirely plausible one was caught which did spread far. A deranged person trying to cause a pandemic would be much more likely to succeed by alternately digging around unprotected in batcaves and going to crowded concerts than trying to do anything sophisticated with bioengineering. 5 2 2 Share