메뉴
BL
MarkTechPost • 11일 전

사카나 AI, 1000층 신경망 학습 가능한 역전파 대안 'PC-ALM' 공개

IMP
7/10
핵심 요약

Sakana AI 연구진이 레이어 단위 지역 학습 방식인 PC-ALM(Augmented Lagrangian Predictive Coding)을 발표했습니다. 각 층 제약에 라그랑주 승수를 붙여 예측 코딩의 지역적 업데이트를 유지하면서도 선형 신경망에서 정확한 역전파 그래디언트를 복원합니다. 추론 예산 T = 2L에서 폭·깊이 8~128 구간 전반에서 역전파와 동등한 성능을 보이며, MNIST에서 1000층 잔차 MLP를 역전파 대비 약 2포인트 이내 오차로 학습할 수 있습니다.

번역된 본문

Sakana AI 연구진 제프리 실리(Jeffrey Seely)와 줄리언 굴드(Julian Gould)가 역전파(Backpropagation)의 지역 학습(local-learning) 대안인 증강 라그랑주 예측 코딩(Augmented Lagrangian Predictive Coding, PC-ALM)을 소개했습니다. 각 층 제약에 라그랑주 승수(Lagrange multiplier)를 부착함으로써 PC-ALM은 예측 코딩의 레이어 단위 지역 업데이트를 유지하면서도 선형 신경망에서 정확한 역전파 그래디언트를 복원해냅니다. 추론 예산 T = 2L에서 폭과 깊이 8부터 128까지 전 구간에서 역전파와 동등한 성능을 보이며, 기준 셀에서 그래디언트 코사인 유사도를 역전파 대비 0.604에서 0.909로 끌어올렸고, MNIST에서 1000층 잔차 MLP(residual MLP)를 역전파와 약 2포인트 이내 차이로 학습했습니다. MIT 라이선스의 JAX 코드가 공개되어 있습니다.

원문 보기
원문 보기 (영어)
Sakana AI researchers Jeffrey Seely and Julian Gould introduce Augmented Lagrangian Predictive Coding (PC-ALM), a local-learning alternative to backpropagation. By attaching a Lagrange multiplier to each layer constraint, PC-ALM keeps predictive coding's layer-local updates while recovering exact backprop gradients in linear networks. It matches BP across widths and depths from 8 to 128 at an inference budget of T = 2L, lifts gradient cosine to BP from 0.604 to 0.909 in the reference cell, and trains 1000-layer residual MLPs within about 2 points of backprop on MNIST. MIT-licensed JAX code is available. The post Sakana AI Researchers Introduce PC-ALM, a Layer-Local Alternative to Backpropagation That Trains 1000-Layer Networks appeared first on MarkTechPost.