메뉴
HN
Hacker News • 2일 전

2주 만에 claude.ai 3배 빠르게 만든 방법

IMP
8/10
핵심 요약

Anthropic이 8월에 2주간의 스프린트를 통해 claude.ai와 데스크톱 앱의 핵심 사용자 경험을 약 3배 빠르게 개선했습니다. Claude(Claude Tag, Opus 5.5 수준의 내부 연구 모델)가 병목 지점을 찾고 벤치마크를 만들고 개선사항을 배포하는 방식으로, 인간은 목표 설정과 변경 승인만 담당했습니다. 그 결과 3천 건 이상의 변경을 단 한 건의 고객 대면 사고나 롤백 없이 병합했으며, 하루에 수만 시간의 사용자 대기 시간을 절약한 것으로 추정됩니다.

번역된 본문

엔지니어링

2주 만에 claude.ai를 3배 빠르게 만든 방법

Claude는 무언가를 측정할 수만 있으면 그것을 더 빠르게 만들 수 있다. 그래서 우리는 계속 더 많은 것을 측정했다.

저자: Raymond Wang, Sam Attard, Issac G. 게시일: 2026년 9월 23일 읽는 시간: 15분

이번 8월, 우리는 2주간의 스프린트를 통해 claude.ai와 Claude 데스크톱 앱의 핵심 사용자 경험을 약 3배 빠르게 만들었다. 사용자들이 느리다고 말해왔고, 그 말이 맞았다. 우리는 모든 작업을 단 하나의 Slack 채널에서 진행했으며, 모든 스레드에 Claude가 참여했다. 우리는 사용자 활동의 95%를 차지하는 네 가지 여정(journey)에 집중했다.

75퍼센타일 기준으로, claude.ai를 새로 로드했을 때 타이핑 가능한 페이지가 뜨는 시간은 3.1초에서 0.55초로, 새 Claude Code 세션 시작은 0.8초에서 0.3초로, Claude Cowork 클라우드 세션 로딩은 2.6초에서 0.73초로 단축되었다. 종합적으로 매일 수만 시간의 사용자 대기 시간을 절약하는 것으로 추정된다.

우리는 Opus 5.5와 대체로 비교 가능한 내부 연구 모델을 사용하는 Claude Tag(베타)를 활용했다. Claude가 병목 지점을 찾고, 벤치마크를 만들고, 개선사항을 배포하고, 모든 배포를 모니터링했다. 우리는 목표를 설정하고, 트레이드오프를 결정하며, 모든 변경을 승인하는 방식으로 방향을 잡았다. 이 접근 방식으로 우리는 단 한 건의 고객 대면 사고나 롤백 없이 3천 건이 넘는 변경을 병합했다.

이 글에서는 우리가 무엇을 배포했는지, 어떻게 측정했는지, 그리고 Claude와 함께 안전하게 이를 수행하기 위해 만든 루프를 다룬다.

요약

스프린트 전에 우리는 다음과 같은 상시 지침이 있는 Slack 채널을 만들었다:

@Claude 당신의 임무는 claude.ai 웹사이트와 데스크톱 앱의 성능과 관련된 모든 일을 지원하는 것입니다. 당신의 책임에는 배포에서 성능 회귀 모니터링, 기존 텔레메트리의 정확성과 완전성 평가, 잘 정리된 옵저버빌리티 대시보드 유지 관리, 관찰된 문제와 쉬운 개선 과제(로우행잉 프루트)에 대한 해결책 선제적 구현, 성능 프로젝트 기회 제안, 인간 팀원들과의 커뮤니케이션이 포함됩니다. […] 이 채널의 궁극적인 목표는 당신이 최대한 자율적으로 작동하는 것이지만, 현재로서는 아직 그것이 가능하지 않다는 것을 우리는 알고 있습니다.

우리는 Claude에게 Datadog MCP 서버를 통해 사용량 데이터를 분석하도록 요청했다. Claude는 가장 영향이 큰 네 가지 사용자 여정을 식별했다: 앱 실행, 대화 시작, 기존 대화 불러오기, 메시지 보내기다. 웹과 데스크톱, 그리고 우리 제품 전반에 걸쳐 이러한 여정은 13가지 개별 측정 항목으로 나뉘었다.

기준선(baseline)을 설정하기 위해 각 측정이 직접 비교 가능하도록 계측(instrumentation)을 추가했다: 각 측정은 사용자 상호작용으로 시작해 결과가 렌더링되면 끝나며, 클라이언트 작업과 서버 작업을 구분했다.

우리는 각 여정을 겨냥한 손으로 직접 고른 약 20개의 프로젝트 목록으로 스프린트를 시작했다. Claude는 각 프로젝트의 영향을 밀리초 단위로 추정했고, 우리는 이 추정치를 종합해 스프린트 목표를 설정했다. 일부 프로젝트는 규모가 꽤 컸지만, 대부분을 2주 안에 달성할 수 있을 것이라고 생각했다.

우리는 3일째에 13개 목표 중 12개를 달성했다. 계획된 프로젝트는 일찍 완료되었다. 더 빠른 실행을 위해, 사용자가 React 초기화 중에도 입력할 수 있도록 정적 컴포저(composer)를 HTML에 심었고, 데스크톱 셸의 메인 프로세스가 처음부터 다시 컴파일하지 않도록 V8 코드 캐시를 사전 컴파일했다. 더 빠른 화면 전환을 위해, 대화 사이에 컴포저를 마운트된 상태로 유지하고, 사용자가 세션에 마우스를 올리면 세션을 미리 불러오도록(prefetch) 했으며, 사이드바 재렌더링을 90% 줄였다.

또한 Claude가 스스로 기회를 식별하고 새로운 작업 흐름을 제안할 수 있는 여지도 남겨두었다. 이러한 작업 흐름은 빠르게 완전한 프로젝트로 확대되어 초기 목표를 훨씬 뛰어넘었다. 그래서 우리는 새로운 목표를 설정하고, 측정할 수 있는 것을 더 찾았다:

@Claude 우리는 원래 프로젝트 목록의 거의 모든 프로젝트와 그 이상에 자원을 배정하게 됐어. 목록을 새로 고칩시다. […] 우리가 아직 탐색하지 않은 것은 무엇이고, 어디서 더 개선(hill climb)할 수 있고, 지금 가장 기회가 큰 곳은 어디일까? […] 나는 엉뚱한(WACKY) 아이디어에도 열려 있어. 무엇이든 개선될 수 있다(ANYTHING CAN BE HILL CLIMBED).

처음부터 우리는 원했…

원문 보기
원문 보기 (영어)
Engineering How we made claude.ai 3x faster in two weeks Once Claude can measure something, it can make it faster. So we kept finding more things to measure. AUTHOR Raymond Wang, Sam Attard, and Issac G. PUBLISHED Sep 23, 2026 READING TIME 15 min This August, we made the core user experience of claude.ai and the Claude desktop app about 3x faster in a two-week sprint. Users had been telling us it was slow, and they were right. We ran everything from a single Slack channel, with Claude in every thread. We focused on four journeys that make up 95% of user activity. At the 75th percentile, time to a typeable page on a fresh load of claude.ai went from 3.1 seconds to 0.55, starting a new Claude Code session went from 0.8 seconds to 0.3, and loading a Claude Cowork cloud session went from 2.6 seconds to 0.73. In aggregate, we estimate that saves tens of thousands of user-hours of waiting every day. We used Claude Tag (beta), running an internal research model roughly comparable to Opus 5.5. Claude found bottlenecks, built benchmarks, shipped improvements, and watched every deploy. We steered by setting goals, making tradeoffs, and approving every change. With that approach, we merged more than three thousand changes without a single customer-facing incident or rollback. This post covers what we shipped, how we measured it, and the loop we built with Claude to do it safely. THE BRIEF Before the sprint, we created a Slack channel with the following standing instructions : @Claude Your job is to facilitate all things related to the performance of the claude.ai website and desktop app. Your responsibilities include monitoring deploys for performance regressions, assessing the accuracy and comprehensiveness of existing telemetry, maintaining well-curated observability dashboards, proactively implementing solutions for observed issues and low-hanging fruit, proposing performance project opportunities, and communicating with your human teammates. […] The ultimate goal for this channel is for you to become as autonomous as possible, but today we know that isn’t yet possible. We asked Claude to analyze usage data through the Datadog MCP server. It identified the four highest-impact user journeys: launching the app, starting a conversation, loading an existing conversation, and sending a message. Between web and desktop, and across our products, those journeys came to thirteen distinct measurements. To establish baselines, we added instrumentation until they were directly comparable: each started with a user interaction, ended once the result was rendered, and disambiguated client and server work. We kicked off the sprint with a list of about twenty hand-picked projects, each targeting a specific journey. Claude estimated the impact of each project in milliseconds, and we aggregated those estimates to set our targets for the sprint. Some of the projects were fairly large, but we thought we could probably achieve most of them within two weeks. We hit twelve of the thirteen targets by day three. The planned projects landed early. For faster launches, we baked a static composer into the HTML so users can type during React initialization, and precompiled a V8 code cache so the desktop shell’s main process doesn’t recompile from scratch. For faster navigations, we kept the composer mounted between conversations, prefetched sessions when the user hovered over them, and cut sidebar re-renders by 90%. We had also left room for Claude to identify opportunities and propose new workstreams. Those workstreams quickly ramped into full projects of their own, which far exceeded our initial targets. So we set new targets, then looked for more things to measure: @Claude we’ve ended up funding nearly every project in the original projects list and more. let’s do a refresh […] what have we not explored, what can we hill climb on, where is the most opportunity at this point? […] i am open to WACKY ideas ANYTHING CAN BE HILL CLIMBED From the start, we knew we wanted to iterate faster than our deploy cadence. Claude could work asynchronously for many hours, even overnight, and we wanted to let it validate its prototypes without waiting for field reads. To achieve that, we looked for other ways to measure performance in the lab. Sam found the first lead: Eleven minutes later, five threads were running, each focused on a different measurement: instruction counts, V8 call counts, React commits, style recalculations, and DOM mutations. We treated every new benchmark with some skepticism. Each one had two jobs: first, a metric Claude could move in the lab; second, a guardrail in CI with a number that could only ratchet down. If a benchmark was flaky, or if it didn’t actually correlate with user latency, we threw it out rather than let Claude climb the wrong hill. @Claude please prove that hill climbing against each of these can result in measurable wall clock perf wins. we’ll unship the benches for any candidates that cannot prove that Wall-clock time is what users feel, but it’s noisy, and milliseconds are too flaky to use as a CI gate. Instruction counts were appealing because they were deterministic, but we still needed Claude to prove they tracked wall-clock time. So we asked Claude to drive the count down on two hot paths: the routine that assembles a conversation’s message tree, and a scanner for status lines in Claude Code output. Claude profiled both with Valgrind and found that a quarter of the first path’s instructions were megamorphic dictionary lookups, resolving the same message ID three separate times. An hour later it had cut instructions on both paths by 48% and 31%, and wall-clock time had dropped 78% and 44%. We checked in two new ratchets. From then on, any PR that raised the instruction counts of those paths failed CI, and a daily job lowered each ceiling whenever the count went down. That led us to the central lesson of the sprint. With Claude, measuring something makes it tractable. Measurement used to be step zero: you’d add a metric, wait for data to roll in, and only then start to understand the problem. With Claude, it’s step one of the climb. As soon as Claude had a number to beat, it could start optimizing. This meant the highest-leverage thing we could do was find more things to measure. THE LOOP, THREAD BY THREAD All of this ran in the same Slack channel, with multiple engineers and Claude jamming in every thread. From there, the sprint settled into a loop : Someone would open a thread about a slow stretch of a journey, often with a screenshot or recording. Claude would trace the flow, then find or build a benchmark that demonstrated the problem. Once it had a promising result in the lab, Claude would come back with a PR — often several, sized for risk and review, with anything user-visible behind a flag. After it shipped, Claude watched the deploy and read the field data. If performance improved, Claude locked in the win by ratcheting the benchmark down; if not, it turned the flag off and iterated. Then it went looking for the next slow spot in the same journey. An example: someone shared a screen recording that showed sidebar rows popping in after the page loaded. Chat and Cowork rows resolved at different times, making the page feel janky. None of our existing monitors detected it. The closest we had was Cumulative Layout Shift , but each shift only scored about 0.008 — well within the good threshold of 0.1. Issac had the idea to reference the underlying Layout Instability API directly. Claude created a telemetry event that mapped the sources of each layout-shift entry to a named region (e.g. sidebar, transcript) and phase (e.g. before first paint, after typeable). It added an integration test that opened the page with a populated sidebar, held the sidebar’s data until after first paint, and failed on any shift in any named region. Claude used that as a benchmark to prove a fix: the test went red 20 of 20 runs on main, and green 20 of 20 on the PR. After th