메뉴
HN
Hacker News 9일 전

통합 모션 파운데이션 모델 'Inertia-1' 공개

IMP
8/10
핵심 요약

Inertia-1는 다양한 기기, 센서 부착 위치, 샘플링 레이트에 구애받지 않고 동작하는 통합 모션 파운데이션 모델입니다. 손목 데이터만으로 사전 학습해도 다른 신체 부위나 새로운 센서로 우수한 일반화 성능을 보이며, 단일 모델로 활동 인식부터 장기적인 건강 예측까지 아우릅니다. 그동안 파편화되었던 모션 인식 모델의 한계를 극복한 연구로, 웨어러블 및 헬스케어 AI 실무자에게 중요한 의미를 지닙니다.

번역된 본문

위치 전환 (Placement transfer): 손목에서 전신으로 손목 데이터만으로 사전 학습되었음에도 불구하고, 학습 과정에서 보지 못한 엉덩이, 발목, 가슴 등 다른 신체 부위로도 매끄럽게 전환 및 적용됩니다.

다중 스트림 (Multi-stream): 하나의 센서가 여러 개로 추가 센서와 부착 위치를 결합하면 정확도가 크게 향상됩니다. 이 스트림들은 중복되지 않고 서로 보완적인 역할을 합니다.

샘플링 레이트 강인성 (Rate-robust): 1Hz에서 20Hz까지 안정적 사전 학습된 표현(Pretrained representation)은 1Hz와 같은 저주파 샘플링 레이트에서도 성능을 유지합니다. 더 높은 주파수는 미세한 건강 신호를 포착하는 데 도움을 줍니다.

범용 태스크 (Task-universal): 하나의 모델로 모든 작업 수행 동일한 핵심 백본(Backbone) 네트워크가 활동 인식, 보행 분석, 장기적 질병 예측을 모두 처리합니다.

건강 인사이트 (Health insight): 움직임을 건강으로 해석 일상적인 수동적 움직임에도 장기적인 건강 지표가 담겨 있으며, 이를 통해 일상의 움직임과 임상 결과를 연결합니다.

기기 독립성 (Device-agnostic): 기기 및 센서를 가리지 않음 새로운 기기와 센서 방식에 강건하므로, 기존 웨어러블 기기에 즉각적으로 통합 및 활용할 수 있습니다.


Inertia-1: 통합 모션 파운데이션 모델을 향한 개방형 탐구

핵심 아이디어: 하나의 범용 모션 모델을 향하여 움직임은 보편적이지만, 이를 위해 만들어진 모델들은 그동안 보편적이지 않았습니다. Inertia-1는 이러한 파편화된 모델 환경을 하나로 통합합니다.

  1. 파편화된 분야 기존 데이터셋들은 샘플링 레이트, 윈도우 길이, 센서 방식, 신체 부착 위치, 심지어 신호 형식과 같은 기본적인 부분에서도 서로 달랐고, 각 작업마다 별도로 맞춤형 모델이 필요했습니다. 특정 환경에서 얻은 연구 결과를 다른 환경에 적용하기란 거의 불가능했습니다.

  2. 하나의 통합된 탐구 Inertia-1는 고립된 개별 사례가 아닌, 단일하고 통제된 환경 내에서 데이터, 센싱, 목표, 규모 등 모션 모델의 전체 라이프사이클을 연구합니다.

  3. 범용 표현(General representation) 그 결과로 얻은 이점은 바로 부착 위치, 기기, 작업에 걸쳐 적응할 수 있는 단일 표현입니다. 학습된 환경을 훨씬 뛰어넘어 동작하는 동일한 핵심 백본 구조를 제공합니다.

(모델이 지원하는 신체 부위: 머리, 가슴, 등, 팔, 손목, 손, 엉덩이, 허벅지, 무릎, 정강이, 발목) (모델이 지원하는 센서: 가속도계(Accelerometer), 자이로스코프(Gyroscope), 자력계(Magnetometer), 3축 ENMO) (모델이 지원하는 환경: 0.2Hz ~ 20Hz 샘플링 레이트, 10초 ~ 2시간 윈도우 길이, 시간 영역(Time domain) 및 주파수 영역(Frequency domain)) (수행 가능한 작업: 활동 인식, 보행 감지, 종단적 건강 분석)


우리의 발견 (What we found) Inertia-1는 기존 벤치마크를 넘어, 모션 모델이 실제 현실 세계에서 제대로 작동할 수 있는지를 결정하는 핵심 요소들을 발견했습니다.

손목으로 학습하고, 신체 어느 곳에든 사용하라. 손목 데이터로 한 번 사전 학습한 후에는 모델을 어디든 적용할 수 있습니다. 학습 중에 한 번도 보지 못한 새로운 신체 부위는 물론, 자이로스코프나 자력계와 같은 새로운 센서 유형에서도 성능을 유지합니다. 신체 부위가 바뀔 때마다 재학습할 필요가 없습니다.

더 많은 스트림을 결합할수록 더 많은 신호를 얻는다. 추가 부착 위치, 자이로스코프, 자력계 등 더 많은 스트림을 결합하면 학습된 표현이 더 정확하고 명확해지며, 각 활동이 더 뚜렷한 클러스터로 분리됩니다. 이 스트림들은 서로 보완적이므로 각각 다른 스트림이 놓치는 정보를 포착해 냅니다.


알아둘 만한 추가 사항 센싱 디자인은 가장 중요한 1순위 선택입니다. 어떻게 모션을 포착하느냐가 모델의 한계를 결정합니다. 본 연구에서 얻은 실용적인 몇 가지 지침은 다음과 같습니다.

  • 샘플링 레이트: 사전 학습된 모델은 활동 인식에서 1Hz의 낮은 주파수에서도 강력한 성능을 유지합니다. 하지만 더 미세한 건강 신호를 분석하려면 더 높은 샘플링 레이트가 도움이 됩니다.
  • 윈도우 길이: 대부분의 작업에서 30~60초 윈도우가 가장 적합합니다. 문맥을 파악하기에 충분히 길면서도 날카로운 예측을 유지하기에 충분히 짧은 최적의 지점(Sweet spot)입니다.
  • 3개 축 모두 유지: 완전한 3축 입력은 벡터 크기(Vector-magnitude)로 요약된 데이터보다 항상 더 나은 성능을 냅니다. 추가적인 축에는 보존할 가치가 있는 신호가 담겨 있습니다.
  • 시간 영역(Time domain) 유지: 시간 영역 모델링이 주파수 영역 재구성보다 보행 및 건강 단서를 더 잘 보존합니다.

작동 원리 원시 신호부터 실제 세계(Real-world)까지 하나의 파이프라인으로 작동합니다.

원문 보기
원문 보기 (영어)
Placement transfer From one wrist to the whole body Pretrained on the wrist alone, it transfers to placements it never saw — hip, ankle, chest, and more. Multi-stream One sensor becomes many Fuse extra sensors and placements and accuracy climbs — the streams are complementary, not redundant. Rate-robust Steady from 1 Hz to 20 Hz Pretrained representations stay strong even at low sampling rates, with finer rates helping subtle health signals. Task-universal One model, every task The same backbone powers activity recognition, gait analysis, and long-horizon disease prediction. Health insight Movement, read as health Passive motion carries long-horizon health markers, linking everyday movement to clinical outcomes. Device-agnostic Works across devices & sensors Robust to new devices and sensor modalities, so it plugs into whatever a wearable already has. Inertia-1 An Open Exploration to a Unified Motion Foundation Model Contact Us The big idea Towards one general motion model Motion is universal — but the models built for it weren't. Inertia-1 brings the whole landscape under one roof. 01 A fragmented field Datasets disagree on the basics — sampling rate, window length, sensor modality, body placement, even signal format — and every task gets its own bespoke model. Findings rarely carry from one setup to the next. 02 One unified exploration Inertia-1 studies the full lifecycle of motion models — data, sensing, objectives, and scale — inside a single, controlled space instead of isolated one-offs. 03 A general representation The payoff: one representation that adapts across placements, devices, and tasks — the same backbone, working far beyond the setting it was trained on. Head Chest Back Arm Wrist Hand Hip Thigh Knee Shin Ankle Accelerometer Gyroscope Magnetometer Triaxial ENMO 0.2 Hz 1 Hz 5 Hz 20 Hz 10 s window 30 s window 60 s window 2 hr window Frequency domain Time domain Activity recognition Gait detection Longitudinal health Head Chest Back Arm Wrist Hand Hip Thigh Knee Shin Ankle Accelerometer Gyroscope Magnetometer Triaxial ENMO 0.2 Hz 1 Hz 5 Hz 20 Hz 10 s window 30 s window 60 s window 2 hr window Frequency domain Time domain Activity recognition Gait detection Longitudinal health body) ============ --> What we found Beyond benchmarks, Inertia-1 surfaces the choices that decide whether a motion model actually works in the real world. Learn it on the wrist. Use it anywhere on the body. Discover more Go back Pretrain once on the wrist, then point the model anywhere. It holds up on body placements — and even sensor types like gyroscope and magnetometer — that it never saw during training. No retraining for each new spot on the body. Wrist accelerometer Other sensors · gyro, mag Other placements Fused representation higher accuracy · cleaner motion clusters Add more streams. Get more signal. Discover more Go back Stack on more streams — extra placements, gyroscope, magnetometer — and the learned representation gets both more accurate and cleaner, with activities separating into tighter clusters. The streams are complementary: each one catches something the others miss. Also worth knowing Sensing design is a first-order choice How you capture motion shapes what a model can do with it. A few practical rules of thumb from the study. Sampling rate Pretrained models stay strong even at a low 1 Hz for activity recognition; finer-grained health signals benefit from higher sampling rates. Window length 30–60 second windows hit the sweet spot across most tasks — long enough to capture context, short enough to stay sharp. Keep all three axes Full triaxial input consistently beats collapsed vector-magnitude summaries — the extra axes carry signal worth keeping. Stay in the time domain Time-domain modeling preserves gait and health cues better than frequency-domain reconstruction. How it works One pipeline, from raw signal to real-world insight The general representation comes together in three clean steps. 01 Pretrain at scale Learn from planetary-scale accelerometry — over 18 million hours across global cohorts — with self-supervision, no labels required. 02 Transfer across settings Adapt the same representation to new placements, devices, and sampling rates with light tuning — or none at all. 03 Deploy across tasks Power activity, mobility, and health applications from one backbone — from fitness tracking to clinical screening. Capabilities From movement to meaning The same representation spans the full spectrum of motion understanding. Activity & behavior Recognize everyday activities and behavioral patterns with state-of-the-art accuracy across diverse populations. Gait & mobility Detect subtle gait changes — like freezing of gait — that signal mobility decline and neurological conditions. Health & disease Surface long-horizon health markers from passive motion, linking everyday movement to clinical outcomes. One general model for human motion Inertia-1 is a first step toward a unified motion foundation model — and an open invitation to collaborators with motion data, new tasks, or a shared interest in where the field is headed. Read the paper Contact Us