메뉴
BL
TechCrunch AI 41일 전

오픈AI도 주목하는 로봇 학습 데이터, XDOF의 700억 규모 도전

IMP
8/10
핵심 요약

최근 AI 로봇 분야의 가장 큰 병목 현상은 언어 모델을 능가하는 고품질 '물리적 상호작용 데이터'의 부족입니다. 스타트업 XDOF는 원격 조종, 데이터 정제 및 파이프라인 구축을 전담하는 인프라 기업으로서 7천만 달러(약 950억 원)를 투자받았습니다. UC 버클리와 협력하여 역대 최대 규모의 로봇 훈련 데이터셋인 ABC를 공개하며, 선도적 AI 연구소들의 로봇 모델 학습을 가속화하고 있습니다.

번역된 본문

2주 전, OpenAI가 2021년에 폐쇄했던 로봇 공학 프로그램을 재가동하겠다고 발표했습니다. 이는 최대 규모의 AI 연구소들이 물리적 세계에서 작동하는 기계를 가르치기 위해 경쟁하고 있다는 최신 신호입니다. 하지만 유능한 로봇을 구축하려면 AI 산업이 아직 갖추지 못한 것이 필요한데, 바로 언어 모델(LLM)에 사용되는 것과 필적하는 수준의 학습 데이터입니다. 이러한 격차는 새로운 종류의 인프라 비즈니스를 창출하고 있습니다.

공개된 방대한 텍스트를 기반으로 학습된 LLM과 달리, 로봇은 물리적 상호작용을 포착하는 데이터가 필요하지만 이런 종류의 데이터는 거의 존재하지 않습니다. 유튜브 동영상이나 긱 워커(Gig worker)들이 촬영한 영상은 해상도가 낮고(fidelity), 물리적 세계와 조화시키기 어렵습니다.

오늘 스텔스 모드(비공개 상태)를 벗어난 XDOF(발음: '엑스-도프')는 AI의 다음 위대한 병목 현상이 모델이나 칩이 아니라, 로봇에게 물리적 세계와 상호작용하는 방법을 가르치는 데 필요한 '데이터 피드백 루프'가 될 것이라베팅하고 있습니다. 이 스타트업은 최첨단 연구소와 로봇 공학 기업이 자체적으로 구축하기 어려운 데이터 파이프라인, 수집 도구 및 어노테이션(주석) 시스템을 구축하는 것을 목표로 하며, 이를 위해 Thrive Capital, Spark Capital, a16z, Lux, WndrCo로부터 7천만 달러(약 950억 원)의 자금을 조달했습니다.

공동 창립자이자 CEO인 Philippe Wu에 따르면, 약 60명의 직원을 보유한 XDOF는 여러 최첨단 AI 연구소를 포함하여 20개의 고객과 이미 협력하고 있으나 구체적인 이름은 밝힐 수 없다고 합니다.

Wu는 "모든 최고 수준의 연구소들이 로봇 공학을 추구하려 하고 있다"며, "우리는 언어 모델 경쟁에서 뒤처지면서 겪었던 몇 가지 실패를 이미 목격했습니다... 이 기술을 너무 늦게 쫓는 상황에 놓이고 싶지 않으며, 이제 모두가 '물리적 AI(Physical AI)'가 다음 프론티어라는 배에 탔습니다."라고 말했습니다.

Wu는 UC 버클리의 박사과정 학생 시절 직접 이 문제에 부딪혔습니다. 그의 초점은 로봇이 대규모 데이터 세트로부터 기술을 학습하도록 하는 것이었습니다. 한 가지 문제가 있었죠. 그는 TechCrunch와의 인터뷰에서 "우리가 다룰 대규모 데이터가 없었습니다. 닭이 먼저냐 알이 먼저냐 하는 문제가 있었습니다. 로봇을 위한 파운데이션 모델(Foundation model)을 어떻게 훈련시킬지 묻기도 전에 먼저 데이터를 실제로 수집해야만 했습니다."라고 회상했습니다.

Wu와 향후 XDOF 공동 창립자 겸 CTO가 될 Fred Shentu는 인간 조작자가 로봇 팔을 제어하여 학습 데이터를 생성할 수 있는 저렴한 원격 조종(Teleoperation) 시스템인 GELLO라는 프로젝트를 진행했습니다. Wu는 "이 프로젝트는 결국 로봇 공학 분야에서 매우 영향력 있는 논문이 되었습니다. 많은 사람들이 비슷한 요구와 병목 현상을 겪고 있었고, 데이터 수집을 위해 이러한 유형의 장치를 활용하기 시작했습니다."라고 말했습니다.

이 기회를 포착한 Wu, Shentu, 그리고 세 번째 공동 창립자 겸 최고 운영 책임자(COO)인 Nemo Jin은 로봇 모델을 추구하는 기업들을 위한 데이터 생태계를 제공하기 위해 2024년 10월 XDOF를 설립했습니다. 단순한 데이터 제공만으로는 비즈니스의 막다른 길이 될 수 있다는 점을 인지하고, 이 회사는 데이터 정제, 도구 제공, 그리고 어노테이션에도 집중하며 로봇 훈련자를 위한 자가 강화형 피드백 루프를 만들고 있습니다.

출발점으로서 이 회사는 UC 버클리의 AI 연구소와 협력하여 지금까지 구축된 것 중 가장 방대한 고품질 로봇 학습 데이터 컬렉션인 'ABC'를 공개했습니다. 여기에는 로봇 조작 데이터 궤적(Trajectory) 13만 개, 시뮬레이션 300시간, 평가 100시간 분량이 포함되어 있습니다. 이 정도 규모의 사전 학습(Pre-training) 데이터는 학계에 이전에 결코 제공된 적이 없었습니다.

이 데이터 공개를 조직하는 데 도움을 준 UC 버클리의 박사과정 학생인 David McAllister는 TechCrunch에 "우리는 언어, 이미지 생성 및 기타 분야에서 모델과 데이터가 공개되었을 때, 커뮤니티가 반드시 기대하지 않았던 것들을 달성하는 것을 보아왔습니다."라고 전했습니다. 팀은 이미 이 데이터를 사용하여 티셔츠 접기, 상자 평평하게 만들기 또는 에어팟 케이스에 에어팟 넣기와 같은 벤치마크 작업에서 로봇을 학습시켰습니다.

무한한 자유도 (Unlimited degrees of freedom)

이 회사는 3계층의 데이터 피라미드로 일할 계획입니다. 가장 가치 있는 계층은 실제 배포되는 로봇에서 수집된 원격 조종 데이터입니다. 다음은 GELLO처럼 원격 조종되는 로봇이 더 일반적인 데이터를 수집하는 것이며, 마지막은 인간이 일상적인 작업을 수행하면서 수집하는 '자아중심적(Egocentric)' 데이터로, XDOF는 이 모든 과정에서 핵심적인 역할을 수행할 예정입니다.

원문 보기
원문 보기 (영어)
Two weeks ago, OpenAI said it would relaunch the robotics program it shuttered in 2021 — the latest signal that the biggest AI labs are racing to teach machines to operate in the physical world. But building capable robots requires something the AI industry doesn't yet have, which is the training data to match that used for language models. That gap is creating a new kind of infrastructure business. Unlike LLMs that were trained on a vast sea of publicly available text, robots need data that captures physical interaction, and that kind of data barely exists. YouTube videos and footage captured by gig workers are low-fidelity and hard to reconcile with the physical world. XDOF (pronounced "ecks-doff"), emerging from stealth today, is betting that the next great bottleneck in AI isn't models or chips, but the data feedback loop needed to teach robots how to interact with the physical world. The startup aims to build the data pipelines, collection tools, and annotation systems that frontier labs and robotics companies can't easily build themselves — and has raised $70 million from Thrive Capital, Spark Capital, a16z, Lux, and WndrCo to do it. Co-founder and CEO Philippe Wu says XDOF, which has about 60 employees, is already working with 20 customers including several frontier AI labs, but cannot name them. "All of the top labs are trying to pursue robotics," Wu said. "We've already seen some of the downfalls of falling a little bit behind in the language model race … you don't want to be in this type of situation where you pursue this technology too late, and everyone is in this boat where physical AI is the next frontier." Wu ran into this problem himself as a PhD student at UC Berkeley. His focus was on enabling robots to learn skills from large-scale data sets. There was just one problem. "We didn't have large-scale data to work with," he told TechCrunch. "There was this chicken-and-egg problem — we first needed to actually collect data before we could even ask how to train a foundation model for robotics." Wu and his future XDOF co-founder and CTO, Fred Shentu, worked on a project called GELLO, a low-cost teleoperation system that lets a human operator control a robotic arm to generate training data. "It ended up becoming a very influential paper in robotics, because a lot of people had similar needs and bottlenecks, and many started leveraging this type of device for data collection," Wu said. Spotting the opportunity, Wu, Shentu, and third co-founder and Chief Operating Officer Nemo Jin launched XDOF in October 2024 to provide a data ecosystem for companies pursuing robotics models. Mindful that data provision alone can be a dead-end business, the company is also focused on data cleaning, tooling, and annotation — creating a self-reinforcing feedback loop for robot trainers. As a starting point, the company is partnering with UC Berkeley's AI Research lab to release what it believes is the largest collection of high-quality robot training data ever assembled, dubbed ABC. It includes 130,000 trajectories of robot manipulation data, 300 hours of simulation, and 100 hours of evaluations. That kind of scaled-up pre-training data has never been available to academia before. "We've seen in language, image generation, and other fields, that when models and data are released, the community achieves things that you wouldn't necessarily have expected," David McAllister, a Berkeley PhD student who helped organize the release, told TechCrunch. The team has already used the data to train robots on benchmark tasks like folding T-shirts and flattening boxes, or loading AirPods into their cases. Unlimited degrees of freedom The company plans to work across three tiers of a data pyramid. The most valuable tier is teleoperation data collected on the actual robot being deployed; next comes teleoperated robots gathering more general data, as with GELLO; and finally "egocentric" data gathered by humans performing everyday tasks, for which XDOF plans to build its own wearable sensors. "Your camera choice is going to affect the quality of your data — which is going to affect how your hand-tracking algorithm performs," Wu said. "If you don't design the hardware well from the start, the data you collect might have very specific problems that you didn't anticipate." The company plans to hire and train armies of teleoperators and egocentric data operators around the world — a labor-intensive model that raises an obvious question: Why aren't the major labs doing this data production work themselves? "You need a warehouse of hundreds of thousands of square feet with hundreds of robots," Wu said. "You need to maintain these robots, calibrate their physical parameters, and properly train operators." It's a build-out that requires focus, capital, and operational scale that most AI labs would rather outsource — which is precisely the market XDOF is betting on. The name XDOF is a play on the robotics term "degrees of freedom," which describes the number of independent motions a robot can perform. Your arm, from shoulder to wrist, has seven degrees of freedom . Humanoid robotics company Figure.AI's latest robot has 30. The X in the company's name captures its ambition: "Arbitrary degrees of freedom, unlimited degrees of freedom," Wu says. Topics AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio June 18 Los Angeles Get an inside look at what it takes to scale and succeed from leaders at Mach Industries, Founders Fund, and Shinkei Systems. Through candid fireside chats and high-impact networking, you'll walk away with valuable insights and new connections. REGISTER NOW Most Popular SpaceX to acquire Cursor for $60B in stock, days after blockbuster IPO Sean O'Kane The US government's Anthropic models ban was never about an AI jailbreak Zack Whittaker The AI layoff wave is becoming a powder keg Connie Loizos Amazon CEO reportedly raised Anthropic model concerns before government crackdown Anthony Ha The FBI built its own replica small town to simulate real-world cyberattacks Zack Whittaker Meta's months-old AI unit is a soul-crushing gulag, say the engineers stuck inside it Connie Loizos Jeff Bezos's Prometheus raises $12B to build an ‘artificial general engineer' for the physical world Marina Temkin