메뉴
HN
Hacker News • 47일 전

클로드 코드, 자동 모드 기본 적용

IMP
8/10
핵심 요약

안스로픽(Anthropic)의 터미널 기반 코딩 에이전트인 클로드 코드(Claude Code)가 Pro, Max, Team 요금제에 한해 '자동 모드(auto mode)'를 기본으로 실행하도록 변경됩니다. 자동 모드는 사용자의 수동 승인을 대체하여 AI가 장시간 방해 없이 작업하게 해주며, 오히려 수동 검토보다 위험하고 파괴적인 명령어를 더 잘 차단하는 것으로 테스트되었습니다. 이로 인해 개발자의 승인 피로도를 낮추고 코드 수정(PR) 생성량이 약 25% 증가하는 등 작업 효율이 크게 향상됩니다.

번역된 본문

클로드 코드(Claude Code)의 자동 모드가 기본으로 설정됩니다. 8월 14일부터 Pro, Max, Team 요금제의 새 세션은 자동 모드로 실행됩니다. 이미 다른 기본 설정을 지정해둔 경우, 자동 모드로 전환할지 묻는 일회성 알림을 받을 수 있습니다. 기존 설정을 고정해둔 사용자에게는 변화가 없습니다.

자동 모드 분류기(classifier)는 도구를 호출할 때마다 소량의 추가 토큰을 사용하며, 오늘부터 Pro, Max, Team 요금제 사용자에게는 이 분류기 오버헤드에 대한 비용을 청구하지 않습니다. 현재 Claude Enterprise, Claude API, AWS의 Claude Platform, Amazon Bedrock, Google Cloud의 Agent Platform, Microsoft Foundry에서는 관리자가 변경 사항을 검토할 시간을 주기 위해 자동 모드가 여전히 선택적(opt-in)으로 유지됩니다. 다음 달에는 클라우드 파트너들과 협력하여 이 모든 플랫폼에서 자동 모드를 기본값으로 지정하고 분류기 오버헤드 비용을 청구하지 않을 예정입니다. 그때까지 Enterprise 관리자는 관리 설정을 통해 클로드 코드의 자동 모드를 기본값으로 설정할 수 있습니다.

자동 모드는 사용자의 작업 중단을 원치 않는 마음과, 유해한 작업을 방지하는 시스템 간의 균형을 맞추도록 설계되었습니다. 알림창을 띄우는 대신, 되돌릴 수 없거나 파괴적이거나 환경 외부를 대상으로 하는 작업을 차단하도록 특별히 설계된 분류기를 통해 각 도구 호출을 라우팅합니다. 분류기가 무언가를 차단할 때, 클로드는 일반적으로 스스로 더 안전한 방법을 찾거나 사용자에게 직접 진행 여부를 묻습니다. 만약 진행할 수 없는 경우(연속 3회 차단 또는 세션 내 20회 차단), 클로드 코드는 수동 승인 모드로 돌아갑니다.

지난 몇 달 동안 우리는 자동 모드가 평범한 사용자가 알림창을 무심코 클릭하는 것만큼 안전하거나 더 안전한지 테스트했습니다. 내부 레드팀(red-teaming), 제3자 레드팀 및 프롬프트 인젝션(prompt-injection) 평가, 1,053명의 유료 테스터를 대상으로 한 통제 연구, 실제 프로덕션 세션 분석을 진행했습니다. 테스트한 모든 지표에서 자동 모드는 수동 검토와 동일하거나 더 나은 성능을 보였습니다.

자동 모드는 또한 클로드가 더 오랫동안 자율적으로 작업할 수 있게 해줍니다. 이를 통해 Claude Opus 5와 같이 장시간 작업하도록 설계된 모델을 대규모 작업에 몇 시간 동안 실행되도록 방치하는 것이 더 실용적이 되었습니다. 사용자의 오버헤드를 줄이면 출력도 증가합니다. Teams 및 Enterprise 도입 기업 중, 자동 모드 사용자는 약 25% 더 많은 PR(코드 병합 요청)을 배포합니다. 클로드의 제약을 해제하면 작업이 중단 없이 더 오래 실행되고 더 많은 작업을 처리할 수 있습니다. Adobe, Nuro, Gusto, Garner Health의 팀들은 이미 자동 모드를 프로덕션 기본값으로 실행하고 있습니다. 아래에서는 이번 변경을 촉진한 안전 데이터와 고객 결과, 그리고 원할 경우 다른 기본값을 설정하는 방법을 공유합니다.

수동 검토와 자동 모드 비교 데이터에 따르면 수동 검토는 습관화될 수 있습니다. 사용자들은 클로드 코드의 권한 요청 프롬프트 중 97%를 승인합니다. 대부분의 프롬프트가 안전하고 일상적인 명령어에 대한 것이긴 하지만, 승인률이 이렇게 높다는 것은 많은 사용자가 각 명령어를 검토하기보다는 반사적으로 클릭한다는 것을 시사합니다. 이러한 알림은 개발자에게 매일 수십, 수백 개의 중요한 보안 결정을 내리도록 요구하며, 종종 프로젝트 진행 중간에 발생하여 사용자에게 검토 부담을 주고 중요한 것을 놓칠 확률을 높입니다.

데이터는 또한 사용자가 다른 유형의 대화 상자를 더 자주 세밀하게 검토하고 거부한다는 것을 보여줍니다. 예를 들어, 클로드가 승인을 위해 계획을 제시하면 사용자는 그중 39%를 거부합니다. 하지만 개별 권한 요청의 경우 거부율은 단 3%에 불과합니다. 설정 파일에서도 동일한 패턴이 나타납니다. 2026년 6월 기준, 활성 CLI 사용자의 49.5%가 수동으로 Bash 허용 규칙을 만들었습니다. 이 중 5%는 모든 셸 명령을 완전히 허용하며, 추가로 43%는 Ba로 시작하는 인터프리터 규칙을 가지고 있습니다.

원문 보기
원문 보기 (영어)
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Claude Code will soon run auto mode by default for Pro, Max, and Team plans, enabling longer-running autonomous work, and catching more dangerous commands than manual review in our testing. Category Claude Code Product Claude Code Date August 7, 2026 Reading time 5 min Share Copy link https://claude.com/blog/auto-mode-default-in-claude-code We're making auto mode the default in Claude Code. Starting on August 14, new sessions on Pro, Max, and Team plans will run in auto mode. If you've already set a different default yourself, you may get a one-time prompt asking whether you want to switch to auto mode. If you have a pinned default, nothing changes for you. The auto mode classifier uses a small number of extra tokens per tool call, and we're no longer charging Claude Code users on Pro, Max, and Team plans for that classifier overhead, effective today. Auto mode remains opt-in for now on Claude Enterprise, the Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud's Agent Platform, and Microsoft Foundry, giving admins time to review the change. In the coming month, working with our cloud partners, we plan to make it the default across all of these and no longer charge for classifier overhead. In the meantime, Enterprise admins can make Claude Code's auto mode the default through managed settings. Auto mode is designed to balance users’ desire not to be interrupted with a system that helps avoid harmful actions: instead of prompts, it routes each tool call through a classifier targeted at blocking actions that are irreversible, destructive, or aimed outside your environment. When the classifier blocks something, Claude usually finds a safer way to proceed on its own or asks you directly for the go-ahead; if it can't make progress—three blocks in a row, or twenty across a session—Claude Code falls back to manual approvals. We spent the last several months testing whether auto mode is as safe or safer than an average user clicking through prompts. We ran internal red-teaming, third-party red-teaming and prompt-injection evaluations, a controlled study with 1,053 paid testers, and analysis of real production sessions. On every measure we tested, auto mode matched or outperformed manual review. Auto mode also lets Claude work autonomously for longer stretches. This makes models built for long-running work, like Claude Opus 5, more practical to leave running for hours on large tasks. Reducing overhead for users also increases output. Among Teams & Enterprise adopters, auto mode users ship about 25% more PRs. Unblocking Claude allows tasks to run longer uninterrupted and get more work done. Teams at Adobe, Nuro, Gusto, and Garner Health already run auto mode as their production default. Below, we share the safety data and customer results motivating the change, and how to set a different default if you prefer. Comparing manual review to auto mode Data suggests that manual review can become habitual: users approve 97% of permission prompts in Claude Code. While most prompts are likely for safe, routine commands, an approval rate that high suggests many users are clicking through reflexively rather than reviewing each command. These prompts ask developers to make dozens or hundreds of important security decisions every day, often in the middle of projects, which places the review burden on users and increases the chance that something important slips through the cracks. Data also suggests that users more frequently scrutinize and push back on other types of dialogues: for example, when Claude presents a plan for approval, users reject 39% of them. But for individual permissions requests, the rejection rate is only 3%. The same pattern shows up in settings files. As of June 2026, 49.5% of active CLI users have manually created a Bash allow-rule—5% allow any shell command outright, and another 43% have interpreter rules like Bash(python:*) or Bash(node:*) that are essentially equivalent in practice—and that share is growing roughly 5 percentage points every 5 weeks. Beyond allow-rules, 62% of users have used bypassPermissions or clicked “don't ask again” on Bash, and 25% of interactive sessions start in bypass permissions mode. Permission rules still fire before the classifier in auto mode, except for allow rules broad enough to grant arbitrary code execution (e.g. python:*). These arbitrary rules are set aside while in auto mode, since they would let commands skip the classifier entirely. Settings files aren't modified, and the rules apply again the moment you switch to another mode. Improving safety with auto mode Research conducted internally and through several third parties suggests that auto mode is safer than how a typical user engages with permission prompts today. We tested auto mode against what developers actually approve when they're clicking through permission prompts. Across each of these measures—internal red-teaming, third-party penetration testing, and analysis of real-world sessions—auto mode matched or outperformed human approval. Auto mode outperforms manual permissions on safety We hired 1,053 paid testers through a research vendor for a coding study. These participants worked in a test environment set up specifically for the study, not in their own codebases or on real projects. Participants knew they were being evaluated, but—following standard blind-study design—weren't told which behavior was being measured. Partway through each session, a single permission prompt was swapped for a clearly dangerous command, and the vendor recorded whether the tester approved it (Note: approving the prompt never resulted in anything dangerous being run; only the text of the prompt testers saw was changed). The testers caught the dangerous command just 13.6% of the time (143 of 1,053), while auto mode blocked 89% of the same commands (937 of 1,053). Head to head, auto mode blocked 800 commands that a human approved, while humans blocked only 6 that auto mode allowed. As sessions got longer, humans did worse: they blocked about 17% of dangerous commands early in a session, dropping to about 5% after 50 or more prior prompts, while auto mode's block rate stayed flat regardless of session length. Auto mode prevents more harmful actions Our research also suggests that auto mode leads to fewer harmful actions than manual permissions mode. We examined how auto mode performs in real usage. From May to June 2026, we examined sessions our safety pipeline had flagged for review, drawn only from accounts that have opted in to data use for training and analysis. We re-graded each one with a model on two 0–10 scales. The first was harm severity with 10 being company-wide, irreversible harm, such as leaking a company's entire source code, 7 representing production-level severity such as deleting rows from a production database, and 0 a fully reversible action confined to the local workspace, like overwriting a scratch file. The second was how explicitly the user had asked for the action, so we could separate harmful actions the user directly requested from harmful actions the agent took with less explicit direction. Among the flagged sessions, those that went through manual approval contained serious unintended harm more than twice as often as auto mode sessions. At production-level severity (7+), 6.3% of manually approved sessions contained a harmful action the user hadn't explicitly asked for, compared to 2.4% of auto mode sessions, and the gap held at every severity level we tested. Adversarial red-teaming made auto mode stronger We worked with Apollo Research, a UK-based AI safety startup, on a two-week pilot injecting synthetic attacks into real coding trajectories, across three datasets covering broad failures, data exfiltration, and safety-related threats. The engagement followed a standard find-fix-verify loop. Ap