메뉴
HN
Hacker News • 21일 전

AI가 장애 대응을 대신하며 엔지니어는 시스템 감각을 잃는다

IMP
7/10
핵심 요약

AI 기반 장애 대응 도구(AI SRE)가 루틴한 장애를 자동 해결할수록 인간 대응자의 실전 경험은 줄어들고, 자동화로 해결 불가능한 복잡한 고장애 발생 시 대응 능력이 크게 떨어진다는 문제를 제기한다. 저자는 항공 산업의 시뮬레이터 훈련처럼 소프트웨어 업계에도 장애 시뮬레이터가 필요하다고 주장한다.

번역된 본문

블로그로 돌아가기 2012년 제가 LinkedIn에서 SRE로 일할 때, 스스로 치유하고 과거 장애로부터 학습하는 시스템을 설계한 적이 있습니다. 당시 AI 능력은 지금과 비교할 수 없었고 프로토타입에 그쳤지만, 이제는 현실이 되었습니다. 이런 도구는 모든 것을 처리합니다. 알림을 점검하고, 가설을 세우고, 텔레메트리를 조회하고, 최근 배포와의 상관관계를 분석하고, 심지어 수정 사항을 직접 구현하기까지 합니다. 이런 모습을 보면 무척 기쁘지만, 큰 우려가 하나 있습니다. 우리가 우리의 시스템과 점점 멀어지고 있다는 것입니다. 이런 도구가 일상적인 장애를 해결하는 데 뛰어날수록, 인간 대응자가 쌓을 수 있는 실습 기회는 줄어듭니다. 그리고 자동화가 해결할 수 없는 모호하고 심각도가 높은 장애가 발생하면, 대응하는 엔지니어들은 곤경에 처할 것입니다.

자동화는 인간에게 가장 어려운 장애만 남긴다

이런 AI 지원 장애 대응 도구, 흔히 'AI SRE'라고 불리는(개인적으로 별로 좋아하지 않는 용어입니다) 도구들은 여러 면에서 훌륭합니다. 밤에 일상적인 장애를 처리해서 용량 문제 때문에 깨어나지 않아도 될 때는 특히 마법처럼 느껴집니다. 문제는, 일상적인 장애야말로 대응자들이 시스템이 어떻게 동작하고 실패하는지에 대한 직관을 '안전하게' 키우는 방법이라는 점입니다. AI가 해결할 수 없는 어렵고 처음 보는 장애를 만나면, 엔지니어는 예전보다 실습이 부족한 상태로 넘겨받아야 합니다.

인적 요인 연구자 리산 베인브리지(Lisanne Bainbridge)는 이 역설을 1983년의 유명한 논문 '자동화의 아이러니(The Ironies of Automation)'에서 설명했습니다. 자동화는 운영자가 일상적인 업무를 연습할 기회를 줄이면서도, 새롭고 비정상적인 상황에 대한 책임은 여전히 운영자에게 남겨둔다는 것입니다. 따라서 그녀는 운영자가 자동화 이전보다 더 숙련되어야 하고 더 많은 훈련을 받아야 한다고 주장했습니다. 앞으로 대부분 장애의 평균 MTTR은 AI 지원 장애 대응 덕분에 낮아지겠지만, 복잡한 장애의 해결 시간은 오히려 치솟을 것이라고 예측합니다. 장애 대응자들이 시스템과 멀어져 조사에 어려움을 겪을 것이기 때문입니다.

항공 산업은 희귀한 고장에 대비해 조종사를 훈련시킨다

항공 산업에서 영감을 얻을 수 있습니다. 비행기 자동화가 비행의 대부분을 처리하지만, 조종사는 자동화가 처리할 수 없는 상황, 즉 엔진 고장, 신뢰할 수 없는 계기, 이륙 중단, 실속 및 기타 비정상 상황에 대한 책임을 여전히 지고 있습니다. 이런 사건들은 극히 드뭅니다. 예를 들어 최신 터빈 엔진은 엔진 비행 시간 10만 시간당 1회 미만의 비행 중 정지를 겪습니다. 다시 말해, 상업용 조종사는 시뮬레이터 밖에서는 경력 전체를 마쳐도 한 번도 겪지 않을 정도로 드뭅니다. 하지만 고장이 발생하면 조종사는 빠르고 정확하게 대응해야 합니다. 예를 들어 트랜스아시아 항공 235편에서는 이륙 직후 오른쪽 엔진의 프로펠러가 자동으로 페더링되었습니다. 항공기는 왼쪽 엔진만으로도 계속 비행할 수 있도록 설계되어 있었지만, 승무원이 문제를 잘못 판단했습니다. 항공기는 실속에 빠져 첫 번째 경고 후 단 117초 만에 추락했습니다.

항공사 조종사는 정기적으로 시뮬레이터로 돌아가 희귀한 비상 상황을 연습합니다. 미국 FAA 규정에 따르면 기장은 6개월마다 재훈련 또는 숙련도 검사를 완료해야 하며, 여기에는 이륙 중 엔진 고장 같은 시나리오가 포함됩니다. 대부분의 소프트웨어 장애는 생명을 위협하지 않지만, 그렇다고 우리 기술을 갈고닦지 않을 이유는 되지 못합니다.

소프트웨어 업계에는 장애 시뮬레이터가 필요하다

문제를 만든 기술이 그 문제를 해결하는 데 도움을 줄 수도 있습니다. 제가 근무하는 Rootly에서는 Uptime Labs와 협업해 현실적인 장애 시뮬레이션을 통해 이 아이디어를 실현했습니다. 엔지니어는 시뮬레이션된 전자상거래 장애 상황에서 사고 총지휘관(incident commander) 자리에 앉아 관측 가능성(observability) 도구를 사용하면서 Slack에서 LLM 기반 이해관계자들과 조율합니다. 그 결과는 매우 현실적입니다. 무엇이 잘못되었는지 조사하면서 대응을 체계적으로 유지하고 CEO와 고객 지원팀까지 대응해야 합니다. 장애 중에 정말 중요한 기술을 연습할 수 있는 것입니다. 불완전한 정보를 이해하고, 명확하게 소통하고, 협업하고…

원문 보기
원문 보기 (영어)
Back to the blog When I was an SRE at LinkedIn, back in 2012, I designed a system that could heal itself and learn from previous incidents. AI capabilities were nowhere near what we have today, and that remained a prototype, but this is now a reality. These tools do it all: inspect alerts, form hypotheses, query telemetry, correlate recent deployments, and even implement the fix themselves. As much as I love to see it, I have a major concern: we are losing touch with our systems. The better these tools become at resolving routine incidents, the less practice human responders will get. And when an ambiguous, high-severity incident comes in that automation cannot solve, responding engineers will be in trouble. Automation leaves humans with the hardest incidents # These AI-assisted incident response tools, more commonly called “AI SREs” – a term I don’t particularly like – are fantastic in many ways. They feel especially magical when they handle a routine incident at night and you don’t have to wake up for a capacity issue. The problem is that routine incidents are also how responders “safely” develop an intuition for how their systems behave and fail. When AI runs into a hard, never-seen-before incident it cannot solve, engineers will have to take over with less practice than they would have had before. Human-factors researcher Lisanne Bainbridge described this paradox in her famous 1983 paper, The Ironies of Automation. She explained that automation reduces operators’ opportunities to practice routine work while leaving them responsible for new and abnormal situations. She argues that, therefore, operators need to be more skilled and receive even more training than before automation. In the years to come, I predict that the average MTTR for most incidents will go down – thanks to AI-assisted incident response – but that the resolution time will shoot up for complex incidents because incident responders lost touch with their system and are struggling to investigate. Aviation trains pilots for rare failures # We can look at the aviation industry for inspiration. Plane automation handles much of the flying, but pilots remain responsible for situations that automation cannot manage: engine failures, unreliable instruments, rejected takeoffs, stalls, and other abnormal conditions. These events are extremely rare. Modern turbine engines, for example, experience fewer than one in-flight shutdown per 100,000 engine flight hours. In other words, that is rare enough that a commercial pilot may complete an entire career without experiencing one outside a simulator. But when a failure occurs, pilots must react quickly and correctly. For example, on TransAsia Airways Flight 235 , the right engine’s propeller autofeathered shortly after takeoff. And while the aircraft was designed to continue flying on its left engine, the crew misidentified the problem. The aircraft stalled and crashed only 117 seconds after the first warning. Airline pilots regularly return to simulators to rehearse rare emergencies. Under US FAA rules, captains must complete recurrent training or a proficiency check every six months, including scenarios such as an engine failure during takeoff. While most software incidents do not threaten lives, that is no reason not to perfect our craft. Turns out the technology that created the issue can also help close it. The software industry needs incident simulators # At Rootly, where I work, we partnered with Uptime Labs to apply this idea through realistic incident simulations. Engineers take the incident commander’s seat during a simulated e-commerce outage, using observability tools while coordinating with LLM-powered stakeholders in Slack. The result feels real. You have to investigate what’s going wrong while keeping the response organized and dealing with the CEO and customer support. You get to practice the skills that matter during an incident: making sense of incomplete information, communicating clearly, coordinating people, and actually running the response. AI can also help preserve these skills # But what about using AI as a trainer? Responders can ask an agent to explain the steps it took, the signals it examined, and the evidence behind its diagnosis. But explanation and observation are not substitutes for practice. You might pick up a few things from watching Serena Williams play, but you only learn tennis by getting on the court, and incident response is no different. I spent more than half a decade of my career building a software engineering school around progressive education: learning by doing. It was in-person, but we had no teachers; students worked on projects instead of listening to lectures. When Dropbox told me graduates it hired were still too inexperienced at troubleshooting, I created projects that gave students broken infrastructure and required them to diagnose and repair it. For most hands-on skills, I believe hands-on education beats passive instruction by a lot. Incident simulation should become part of on-call readiness # As LLMs do more of our work, engineering teams risk accumulating comprehension debt: a growing gap between how their systems work and how well responders understand them. Engineers should regularly interact with the system they watch over, handle unfamiliar failures, practice working under pressure, and rehearse the coordination and communication required during a SEV0. Tabletop exercises and chaos engineering are nothing new, but practice has become even more important in the LLM era. Researcher Bainbridge recommended giving operators regular hands-on control and using simulation to prevent their skills from decaying. That’s the irony of automation, the more successful it becomes, the less prepared humans may be for the moment it fails. Sylvain Kalache AI Labs lead and DevRel at Rootly. Former LinkedIn SRE and co-founder of Holberton School. Related writing article LLMs Broke the SRE Runbook. Now What? article Is AI-assisted coding an incident magnet? article Vibe Coding Is Here — But Are You Ready for Incident Vibing?