메뉴
BL
The Decoder 16일 전

구글 SensorFM, 웨어러블 데이터를 건강 인텔리전스로

IMP
9/10
핵심 요약

구글 리서치는 500만 명 이상의 방대한 웨어러블 기기 데이터를 사전 학습한 파운데이션 모델 'SensorFM'을 공개했습니다. 이 단일 모델은 수면, 심혈관 및 대사 건강 등 총 35개의 건강/행동 예측 작업 중 34개에서 기존 특화 모델들을 능가하는 성능을 입증했습니다. 이는 파편화된 기존 웨어러블 건강 기능을 통합하고, 막대한 라벨링 비용 없이도 개인화된 AI 건강 비서 구현의 핵심 기반이 될 수 있어 매우 중요합니다.

번역된 본문

구글의 SensorFM, 지저분한 웨어러블 센서 데이터를 범용 건강 인텔리전스 계층으로 변환하다 Tomislav Bezmalinović 2026년 7월 13일 / THE DECODER

핵심 요약

  • 구글 리서치는 500만 명 이상의 라벨링되지 않은 웨어러블 데이터 1조 분 이상을 학습해 생리 및 행동 패턴의 일반적인 표현을 학습하는 파운데이션 모델, SensorFM을 선보였습니다.
  • 이 모델은 핏빗(Fitbit)과 픽셀 워치(Pixel Watch) 기기의 센서 데이터를 처리하며, 특별히 가공된 웨어러블 특징(feature)을 사용한 기존 지도 학습 비교 모델들을 35개의 건강 및 행동 작업 중 34개에서 성능을 능가했습니다.
  • 실험 환경에서 SensorFM의 예측 결과가 추가된 건강 요약본은 연구된 5개 영역 모두에서 기준 모델 버전보다 훨씬 더 높은 평가를 받았습니다.

본문

구글 리서치는 500만 명으로부터 수집된 웨어러블 센서 데이터로부터 인체 생리 및 행동의 일반적인 표현을 학습하는 파운데이션 모델인 SensorFM을 공개했습니다. 이 모델은 35가지의 다양한 건강 및 행동 작업에 적용될 수 있습니다.

오늘날 웨어러블 기기의 대부분의 건강 기능은 단일 목적을 위해 구축됩니다. 한 모델은 수면 단계를 감지하고, 다른 하나는 심혈관 위험을 추정하며, 또 다른 모델은 스트레스나 대사 지표를 분석합니다. 구글은 이러한 파편화된 접근 방식을 공유 AI 기반으로 대체하고자 합니다. 이를 통해 여러 건강 문제에 걸쳐 연속적이고 종종 비어있는(누락된) 센서 데이터를 이해하고, 비용이 많이 드는 라벨링 학습 데이터의 필요성을 줄이며, 궁극적으로 개인화된 맥락을 AI 건강 비서에 제공할 수 있기를 기대합니다.

구글 리서치는 이제 블로그 게시물과 동반 논문을 통해 SensorFM을 소개했습니다. 이 파운데이션 모델은 대량의 라벨링되지 않은 웨어러블 데이터로부터 생리적, 행동적 패턴의 일반적이고 재사용 가능한 표현을 학습합니다.

연구진은 500만 명의 핏빗 및 픽셀 워치 사용자로부터 얻은 1조 분 이상의 다중 모달 센서 데이터를 사전 학습에 사용했습니다. 이 데이터는 100개국 이상에서 수집되었으며 20가지 이상의 다양한 핏빗 및 픽셀 워치 모델을 통해 모아졌습니다. 저자들에 따르면, 이는 이런 종류의 모델을 학습시키는 데 사용된 역대 최대 규모이자 가장 다양한 웨어러블 데이터셋입니다.

더 많은 데이터와 더 큰 모델이 SensorFM을 향상시킨다

SensorFM은 광학 심박수 모니터링(광전 용적 맥파, PPG), 가속도, 피부 전도도, 피부 온도, 기압 고도 등 5가지 유형의 센서 데이터에서 추출한 34가지 특징(feature)을 처리합니다. 이 특징에는 심박수, 심박수 변이도, 혈중 산소 포화도, 수면 단계 및 모션 데이터 등이 포함됩니다.

이 모델은 의도적으로 가려진 데이터 부분을 복원하는 방식인 자기 지도 학습(Self-supervised learning)을 통해 훈련됩니다. '적응형 및 상속 마스킹(Adaptive and Inherited Masking, AIM)'이라 불리는 이 기술은 실제로 누락된 값과 학습 중 인위적으로 숨겨진 값을 모두 표시하여, SensorFM이 두 가지 유형의 데이터 공백을 모두 처리하는 방법을 학습하게 합니다.

연구진은 모델 크기와 데이터 볼륨이 함께 증가할 때 성능이 체계적으로 향상된다고 보고했습니다. 그들이 테스트한 4가지 모델 변형은 약 10만 개에서 1억 개의 매개변수(parameters)까지 다양하며, 학습 데이터셋은 5,000명에서 500만 명까지 포괄합니다. 가장 큰 학습 데이터셋에서 가장 큰 모델의 재구성 오차는 가장 작은 모델보다 31% 낮았습니다. 가장 큰 구성은 또한 대부분의 다운스트림 예측 작업에서 가장 우수한 성능을 보였습니다.

SensorFM, 35개 작업 중 34개에서 비교 모델을 능가하다

그런 다음 연구진은 총 13,985명의 참가자가 포함된 3개의 개별 연구 데이터를 사용하여 SensorFM을 테스트했습니다. 이 모델은 사전 학습 중에 이 데이터를 본 적이 없습니다. 그들은 심혈관 및 대사 건강, 정신 건강, 수면, 인구 통계 및 라이프스타일을 다루는 35가지 예측 작업에 대해 SensorFM을 평가했습니다.

논문에 따르면, SensorFM이 학습한 표현(representation)을 바탕으로 구축된 단순한 작업 특화 헤드 모델조차도 사람이 수동으로 만든 웨어러블 특징을 사용한 지도 학습 기준선(baseline) 모델들을 35개 작업 중 34개에서 능가했습니다. 확장된 사전 학습은 또한 SensorFM을 훨씬 더 라벨링 효율적으로 만들었습니다.

원문 보기
원문 보기 (영어)
Google’s SensorFM turns messy wearable sensor data into a general-purpose health intelligence layer Tomislav Bezmalinović Jul 13, 2026 Nano Banana Pro prompted by THE DECODER Key Points Google Research has introduced SensorFM, a foundation model designed to learn a general representation of physiological and behavioral patterns from more than one trillion minutes of unlabeled wearable data from five million people. The model processes sensor data from Fitbit and Pixel Watch devices and outperformed monitored comparison models, which used specially prepared wearable features, on 34 out of 35 health and behavioral tasks. In the experimental setup, health summaries that included additional SensorFM predictions were rated significantly higher than the baseline version across all five areas studied. Ask about this article… Search Google Research has unveiled SensorFM, a foundation model that learns a general representation of human physiology and behavior from wearable sensor data collected from five million people. The model can be applied to 35 different health and behavioral tasks. Most health features on wearables today are built for a single purpose. One model detects sleep stages, another estimates cardiovascular risk, and yet another analyzes stress or metabolic markers. Google wants to replace these siloed approaches with a shared AI foundation that can make sense of continuous, often gappy sensor data across many health questions, cut the need for expensive labeled training data, and eventually feed personalized context into AI health assistants. Google Research has now introduced SensorFM in a blog post and an accompanying paper. The foundation model learns a general, reusable representation of physiological and behavioral patterns from large volumes of unlabeled wearable data. The researchers used more than a trillion minutes of multimodal sensor data from five million Fitbit and Pixel Watch users for pretraining. The data came from over 100 countries and was collected with more than 20 different Fitbit and Pixel Watch models. According to the authors, this is the largest and most diverse wearable dataset ever used to train a model of this kind. Ad More data and bigger models make SensorFM better SensorFM processes 34 features drawn from five types of sensor data: optical heart rate monitoring (photoplethysmography, or PPG), acceleration, skin conductance, skin temperature, and barometric altitude. The features include heart rate, heart rate variability, blood oxygen saturation, sleep stages, and motion data, among others. The model is trained in a self-supervised way by reconstructing deliberately masked data segments. The technique, called "Adaptive and Inherited Masking" (AIM), flags both genuinely missing values and values that were artificially hidden during training, so SensorFM learns to handle both types of data gaps. Ad DEC_D_Incontent-1 The researchers report that performance improves systematically when model size and data volume grow together. The four model variants they tested range from about 100,000 to 100 million parameters, and the training datasets span from 5,000 to five million people. On the largest training dataset, the biggest model's reconstruction error was 31 percent lower than the smallest model's. The largest configuration also performed best on most downstream prediction tasks. SensorFM beats comparison models on 34 out of 35 tasks The researchers then tested SensorFM on data from three separate studies with a total of 13,985 participants. The model had never seen this data during pretraining. They evaluated SensorFM on 35 prediction tasks covering cardiovascular and metabolic health, mental health, sleep, demographics, and lifestyle. Ad Even simple task-specific head models built on top of SensorFM's learned representations outperformed supervised baselines with hand-crafted wearable features on 34 of 35 tasks, according to the paper. Scaled pretraining also made SensorFM more label-efficient compared to the supervised baselines. The model could adapt to new tasks with relatively few labeled examples, and as it grew larger, it relied less on extra demographic information. The authors believe scaled pretraining could be especially useful for hard-to-measure traits that vary widely between individuals, such as depression and anxiety symptoms. To adapt SensorFM's learned representations to new tasks, the researchers set up a "classroom" of competing and collaborating LLM agents. These agents repeatedly generated, tested, and refined code for downstream prediction models, running more than 30,000 experiments in the process. The models they found outperformed simple linear head models based on the same SensorFM representations on 28 of 35 prediction tasks. Ad DEC_D_Incontent-2 SensorFM makes a health agent's answers better The researchers also integrated SensorFM into a personal health agent and compared three variants. All three received demographic information and daily summaries computed from wearable data, covering things like activity, sleep, blood oxygen, and skin temperature. One variant also received SensorFM predictions for various health markers, a second received the actual known values for those same markers, and the third got none of this extra information and served as the baseline. Ad Four clinicians evaluated 93 health summaries for 31 real participant profiles, spending more than 40 hours and producing 1,860 individual ratings. The result: summaries augmented with SensorFM predictions scored significantly higher than the baseline across all five dimensions the team measured, which were context, personalization, justifiability, relevance, and safety. There was no statistically significant difference overall between summaries that used SensorFM predictions and those that used actual known health data. That said, this doesn't mean SensorFM can replace clinical measurements or diagnoses. SensorFM remains a research model for now The researchers point to several limitations. SensorFM was trained and tested only on data from Fitbit and Pixel Watch devices. Whether the results transfer to other wearables is an open question. The model also doesn't work with high-resolution raw signals but with data aggregated at the minute level, which means very short or fine-grained patterns can get lost. Many of the health markers the team studied are based on self-reports, medication records, or questionnaires rather than clinically confirmed findings. The study population also doesn't fully represent the general population. And the health agent was only evaluated in a static setup with single responses, not in longer conversations with follow-up questions. SensorFM is purely a research model for now. Google already offers the Gemini-based Google Health Coach , which provides personalized tips on fitness, sleep, recovery, and other health topics. SensorFM could eventually serve as a technical foundation for features like these, but Google hasn't announced any concrete plans to integrate it into Fitbit, Pixel Watch, or the AI coach. More details on SensorFM are available in the Google Research blog post and the open-access paper on arXiv . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Google Research