메뉴
BL
Wired AI • 22일 전

오픈AI·앤스로픽 동시 서비스 중단, 원인은 밝혀지지 않아

IMP
6/10
핵심 요약

9월 3일(목) 오전 안스로픽, 오픈AI, xAI의 프론티어 AI 모델이 동시에 서비스 중단을 겪었다. xAI는 스페이스X의 멤피스 컴퓨트 센터 장애가 원인이라고 밝혔지만, 오픈AI는 라우팅 오류라고만 설명했고 안스로픽은 원인 코멘트를 거절했다. 동시 다발 장애는 보통 공유 클라우드 제공업체 문제를 시사하지만, Cloudflare·AWS·Azure 모두 장애를 보고하지 않아 진짜 원인은 베일에 싸여 있다.

번역된 본문

안스로픽, 오픈AI, xAI의 프론티어 모델이 모두 목요일 오전 이례적인 서비스 중단을 겪으며 각 사의 AI 챗봇이 일시적으로 사용 불가 상태가 되었다.

xAI의 모회사인 스페이스X는 목요일 오후, Grok의 문제는 "오늘 아침 멤피스 컴퓨트 센터의 장애" 때문이었다고 밝혔다. 문제들이 시간상 일치했기 때문에 처음에는 서로 연관된 것으로 보였다. 아마도 공통 제3자 서비스 제공업체 때문일 수 있었지만, 오픈AI와 안스로픽 모두 목요일 WIRED에 대한 코멘트에서 외부 원인을 언급하지 않았다. 스페이스스는 WIRED의 코멘트 요청에 응답하지 않았지만, 목요일 공개 코멘트에서 "영향을 받은 컴퓨트 파트너들에게도 사과드리고 싶다"고 밝혔다. 안스로픽과 xAI는 5월에 스페이스X와의 "컴퓨트 파트너십"을 발표한 바 있다.

오픈AI 대변인 캐슬린 차이코스키(Kathleen Chaykowski)는 WIRED에 다음과 같이 말했다: "9월 3일 목요일 태평양 표준시(PT) 오전 7시 43분경 시작된 라우팅 오류로 인해 일부 사용자가 플랫폼 전반에서 ChatGPT와 Codex를 사용할 수 없었다. PT 오전 8시 17분경 해결책이 성공적으로 적용되었으며 계속 모니터링 중이다."

안스로픽은 이 사건에 대한 코멘트를 거절했다. 이 회사는 목요일 PT 오전 6시 23분에 "Claude Mythos 5.1, Claude Fable 5.1, Claude Opus 5 요청에서 오류 증가"를 포함한 "부분적 서비스 중단" 경보를 시작했다. 직후 회사는 "원인을 파악했으며" "수정 사항이 배포되었다"고 밝혔다. PT 오전 9시 16분에 문제가 해결된 것으로 표시되었다. Claude Sonnet 5는 PT 오전 9시 직후 잠시 유사한 문제를 겪은 것으로 보였다.

xAI는 PT 오전 6시 30분에 서비스 상태 페이지에 "장애 조사 중"을 게시하며 모든 플랫폼과 서비스에서 Grok 장애를 보고했다. 해당 페이지는 "Grok에 문제가 발생하고 있습니다. 최대한 빨리 서비스를 복구하기 위해 노력 중입니다"라고 밝혔다. PT 오전 10시 5분에 이 사건은 완료로 표시되었으며, 회사는 "상황을 해결했으며 트래픽이 다시 정상입니다"라고 작성했다.

목요일 오전 구글 Gemini의 가능한 장애에 대한 산발적인 보고도 있었지만, 회사는 이를 확인하지 않았고 서비스 상태 대시보드에 사고를 기록하지도 않았다. 구글은 출판 전 WIRED의 코멘트 요청에 응답하지 않았다.

일반적으로 같은 분야에서 동시에 여러 장애가 발생하면 클라우드 제공업체, 콘텐츠 전송 네트워크(CDN) 또는 기타 제3자 벤더의 문제가 여러 고객에게 영향을 미친 것을 의미한다. 그러나 오픈AI와 안스로픽은 잠재적 공통 원인을 지목하지 않았고, Cloudflare, 아마존 웹 서비스(AWS), 마이크로소프트 Azure를 포함한 주요 인터넷 인프라 기업들은 목요일 장애를 보고하지 않았다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Frontier models from Anthropic , OpenAI , and xAI all experienced rare outages on Thursday morning, creating downtime for their corresponding AI chatbots. SpaceX, xAI’s parent company, said on Thursday afternoon that the issues with Grok resulted from “an outage at our Memphis compute center this morning.” The issues initially appeared to be linked because they coincided—perhaps the result of a shared third-party service provider—but neither OpenAI nor Anthropic cited an external source in comments to WIRED on Thursday. SpaceX , xAI’s parent company, did not respond to WIRED’s request for comment. But the company said as part of its public comments on Thursday: “We’d also like to apologize to our impacted compute partners.” Anthropic and xAI announced a “compute partnership” with SpaceX in May. OpenAI spokesperson Kathleen Chaykowski tells WIRED: “A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms. As of about 8:17 am PT on Thursday, a solution was successfully implemented and is continuing to be monitored.” Anthropic declined to comment on the episode. The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.” Shortly after, the company said it had “identified the cause” and that “a fix has been deployed.” The company marked the issue as resolved by 9:16 am PT. Claude Sonnet 5 seemed to briefly have similar issues shortly after 9 am PT. xAI reported Grok outages across all of its platforms and services beginning at 6:30 am PT when the company posted “investigating outage” on its service status page . “Grok is experiencing issues. We are working on restoring service as quickly as possible,” the page said. At 10:05 am PT the episode was marked complete. “We have resolved the situation, and traffic is healthy again,” the company wrote. There were scattered reports of a possible Google Gemini outage on Thursday morning as well, but the company did not confirm this or record any incidents on its service status dashboard. Google did not respond to WIRED’s request for comment ahead of publication. Typically, multiple outages in the same sector at the same time would point to a cloud provider, content delivery network, or other third-party vendor having issues affecting multiple customers. But OpenAI and Anthropic did not point to a potential shared cause, and major players in the internet infrastructure space—including Cloudflare, Amazon Web Services, and Microsoft Azure—did not report outages on Thursday.