메뉴
HN
Hacker News 5일 전

FLUX 3 x mimic: 차세대 비디오-행동 모델의 등장

IMP
9/10
핵심 요약

Black Forest Labs의 새로운 멀티모달 파운데이션 모델인 FLUX 3가 로봇 제어 및 영상 생성을 하나로 통합하는 '비디오-행동 모델'로 진화했습니다. 이 모델은 영상 생성 과정에서 습득한 물리적 세계의 이해를 바탕으로 로봇의 행동(Action)을 예측하며, mimic 로보틱스와의 협업을 통해 아우디(Audi) 실무 환경에 배포되었습니다. 단일 모델로 영상과 로봇 제어를 모두 처리함으로써 피지컬 AI(Physical AI)가 자연스럽게 확장되는 중요한 기술적 이정표입니다.

번역된 본문

블로그 연구 모델로 돌아가서 FLUX 3 x mimic: 차세대 비디오-행동 모델 우리의 새로운 멀티모달 파운데이션 모델인 FLUX 3의 초기 버전이 현재 로봇에서 실행되고 있습니다. 우리는 mimic robotics에 FLUX.3에 대한 조기 액세스 권한을 제공했습니다. 로봇 학습 및 배포에 있어 그들의 강점과 모델의 세계 지식, 그리고 BFL의 파운데이션 모델 전문성이 결합하여 탄생한 것이 바로 차세대 비디오-행동 모델인 FLUX-mimic입니다. FLUX FLUX 1과 FLUX 2는 이미지를 생성했습니다. FLUX 3는 멀티모달리티로 확장되어 시청각 콘텐츠를 공동으로 생성하며, 동시에 FLUX-mimic의 기반을 제공합니다. mimic와의 협력을 통해 개발된 이 비디오-행동 모델은 아우디(Audi)에서 테스트 및 배포된 로봇을 구동합니다. 언뜻 보기에, 설득력 있는 시각적 콘텐츠를 제작하는 것과 로봇을 제어하는 것은 공통점이 거의 없어 보입니다. 하나는 픽셀을 생성해야 하고, 다른 하나는 물리적 세계가 만지고 조작할 때 어떻게 반응하는지 이해해야 합니다. 만약 하나의 모델이 두 가지를 모두 수행한다면, 그것은 결코 단순한 콘텐츠 생성 모델이 아닙니다. 그것은 세계가 어떻게 작동하는지에 대한 모델이며, 콘텐츠 생성은 그 모델로 할 수 있는 일 중 하나일 뿐입니다. 이것이 바로 FLUX 3입니다. 비디오가 가장 어려운 부분입니다 FLUX 3는 단일 모델로서, 처음부터 이미지, 비디오 및 오디오를 통합하여 학습했습니다. 이 학습의 가장 까다로운 부분(전체 컴퓨팅 비용의 95% 이상을 차지)은 비디오 예측입니다. 현실적인 비디오를 생성하려면 모델은 접촉, 움직임, 무게, 인과관계를 배울 수밖에 없습니다. 이 중 하나라도 틀리면 비디오가 어색하게 보입니다. 세계를 정확하게 렌더링하는 법을 배운다는 것은 세계가 어떻게 작동하는지 배운다는 것을 의미합니다. 상대적으로 오디오는 쉬운 모달리티입니다. 비디오보다 저차원이고 훨씬 세밀하지 않아서, 오디오가 포함된 720p 비디오의 토큰 중 0.5% 미만을 차지합니다. 모델이 비디오 이해를 학습하는 힘든 과정을 일단 마치고 나면, 비디오와 오디오 간의 인과 관계를 학습하여 입술 움직임에 동기화된 음성과 이를 유발하는 물리적 사건에 동기화된 오디오 효과를 예측하게 됩니다. 행동(Action)도 동일한 구조를 따릅니다. 로봇 상태의 저차원 표현은 시각적 관찰과 밀접하게 결합되어 있습니다. 행동, 오디오, 비디오 프레임은 모두 단일 기반 물리적 현실에 대한 부분적 표현입니다. 모델이 비디오와 오디오 이면의 물리적 과정을 학습한 후, 행동 예측은 전혀 새로운 출발이 아닙니다. 그것은 모델이 이미 학습한 현실에 대한 또 하나의 시점일 뿐입니다. 단일 백본(Backbone) 이러한 접근 방식이 옳다면, FLUX 3에게 행동 예측을 가르치는 것이 영구적인 비용을 초래해서는 안 됩니다. 모델이 완전한 성능으로 돌아오기 전에 행동 공간의 구조를 학습하고 내부적인 세계 표현을 이에 맞추기 위해 짧은 혼란 기간이 있을 것으로 예상했습니다. 그리고 그것은 우리가 관찰한 바와 정확히 일치합니다. 대규모 학습 과정에서 우리는 학습 과정에 행동 예측을 추가하고 비디오 생성 품질에 미치는 영향을 관찰했습니다. 모델이 새로운 행동 모달리티를 통합하기 시작하면서 초기에는 텍스트-비디오 및 이미지-비디오에 대한 인간 평가 점수가 최대 10% 하락했습니다. 하지만 3,500 스텝 이후, 모델은 이제 행동까지 예측하면서도 이전의 비디오 생성 작업에 대한 원래의 완전한 품질을 되찾았습니다. 각 계열은 행동 예측이 추가되기 이전의 고유한 품질을 기준으로 정규화되었습니다. 수치가 높을수록 좋습니다. 모델은 입력 및 출력에 행동을 통합해야 했지만, 그렇게 한다고 해서 모델의 용량이 영구적으로 소모되지는 않았습니다. 그저 이 새로운 모달리티가 기존의 세계 모델과 어떻게 관련되는지를 학습해야 했을 뿐입니다. 이것이 파악되고 나자 기존 기능에 대한 성능 저하 페널티는 사라졌습니다. 비디오 생성과 행동 예측은 별도의 파운데이션을 필요로 하지 않습니다. 동일한 백본이 둘 다를 처리합니다. 이 점은 피지컬 AI(Physical AI)가 방향의 전환이 아닌, Black Forest Labs의 로드맵에서 자연스러운 확장이 되게 만듭니다. 콘텐츠 생성은 우리의 멀티모달 FLUX 3 백본이 이미지, 비디오 및 오디오로 수행하는 일입니다. 피지컬 AI는 이 백본이 행동(Action)으로 수행하는 일입니다. 시각적 지능(VIsual intelligence)을 핵심으로 하는 하나의 파운데이션 모델이 모든 것을 가능하게 합니다.

원문 보기
원문 보기 (영어)
Back to blog Research Models FLUX 3 x mimic: The Next Generation of Video-Action Models An early version of FLUX 3, our new multimodal foundation model , is now running on robots. We gave mimic robotics early access to FLUX.3. Their strength in robot learning and deployment, combined with the model's world knowledge and BFL's foundation model expertise, produced FLUX-mimic: the next generation of video-action models. FLUX FLUX 1 and FLUX 2 generate images. FLUX 3 expands into multimodality and generates audio-visual content jointly - and, at the same time, provides the foundation of FLUX-mimic: A video-action model, developed in collaboration with mimic, running robots that have been tested and deployed at Audi. At first glance, producing convincing visual content and controlling robots seem to have little in common. One requires generating pixels, the other an understanding of how the physical world responds when you touch and manipulate it. If one model does both, it was never really only a content creation model. It is a model of how the world behaves, and content creation is one thing one can do with it. That is what FLUX 3 is. Video is the hard part FLUX 3 is one model, jointly trained across images, video and audio from the beginning. The most demanding part of that training - accounting for over 95% of the total compute costs - is video prediction. To generate realistic videos, a model has no choice but to learn contact, motion, weight, cause and effect; get any of them wrong and it looks wrong. Learning to render the world accurately means learning how the world behaves. Relatively speaking, audio is the easy modality. Low dimensional and far less detailed than video, it makes up less than 0.5% of the tokens in a 720p video with audio. Once a model has done the hard work of learning video understanding, it will learn the causal relationship between video and audio to predict speech synchronized to lip movement and audio effects synchronized to the physical events causing them. Actions follow the same shape: a low dimensional representation of a robot's state, tightly coupled to visual observations. Actions, audio and video frames are all partial representations of a single underlying physical reality. After the model has learned about the physical processes behind video and audio, action prediction is not a new departure - it is one more view of the reality it already models. A single backbone If that framing is correct, teaching FLUX 3 to predict actions should not incur lasting costs: we expect a brief phase of disturbance as the model has to learn the structure of the action space and align its internal representation of the world to it, before returning to full performance. That is exactly what we observe. In a large-scale training run, we added action prediction to the curriculum and observed the effect on video generation quality. Human ratings on text-to-video and image-to-video initially fell by up to 10% as the model started to incorporate the new action modality. After 3500 steps, the model had regained its full previous quality on video generation tasks while now also predicting actions. Each series is normalized to its own quality before action prediction was added. Higher is better. The model had to integrate actions into its inputs and outputs - but doing so didn't cost it capacity permanently. It merely had to learn how this new modality relates to its existing model of the world. Once this was figured out, the performance penalty on its existing capabilities was gone. Video generation and action prediction don't need separate foundations. The same backbone carries both. This makes Physical AI a natural extension of our roadmap at Black Forest Labs rather than a change in direction. Content creation is what our multimodal FLUX 3 backbone does with image, video and audio. Physical AI is what it does with actions. One foundation model, with visual intelligence at its core, enabling two families of applications. We didn't build a separate foundation model. We focused on the hard thing: building a model that understands the world. Acting in it is what that understanding makes possible. From lab to reality: FLUX-mimic What happens when we point the FLUX 3 backbone at real automation tasks on real production lines? That's the question mimic and BFL created FLUX-mimic to answer. mimic builds their own robots and brings expertise in robot learning, dexterous manipulation and production deployment; BFL builds visual foundation models and brings multimodal training and modeling expertise. Together, we built a next-generation model for general-purpose manipulation - adapted to industry requirements and integrated into mimic's full-stack deployment system. FLUX-mimic is a video-action model built on the FLUX 3 backbone. Decoding the learned world model Our thesis is that FLUX 3 has to learn an internal representation of the world to be able to generate videos. FLUX-mimic follows through on this thesis and decodes actions from the learned world representation of the FLUX backbone. This approach, pioneered in mimic-video, trains a lightweight action decoder on top of intermediate features extracted from the video prediction path of FLUX. Architecture overview how FLUX-mimic is built on top of FLUX 3 The success of this approach depends on two related but different aspects: the quality of the world model learned by FLUX and the quality of the feature representation of this world model. The quality of the world model is directly related to the generation quality: if a model does not understand how the world behaves, it cannot simulate it. However, even the best world model does not help an action decoder if it is inaccessible: if the feature space keeps the causal relationships between modalities entangled nonlinearly, understanding those relationships from the feature representation remains as difficult as understanding them from the raw inputs - representation quality matters. Generation quality and representation quality have long been studied and approached in isolation from each other. Generative approaches result in high-quality world models that enable simulations and they exhibit scaling laws for predictable returns on compute investments. However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding. As generative models themselves rely on their own features, this divergence in their representation quality seems counter-intuitive. Improved representations within generative models should improve the quality of their world model and make them more usable for downstream tasks. In our work, Self-Flow , we demonstrated how to unify generation and representation learning in a single framework and observed exactly this reciprocal improvement: the world model improved - as measured by generation quality across video, image and audio - and its representation quality improved - as measured by success rate for robot control tasks in simulation. Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better). Scaling the world model Scaling laws remain true with Self-Flow, and FLUX 3 is the application of that: the scaled-up version of Self-Flow. It is trained on tens of millions of hours of general video content to learn world dynamics as broadly as possible from day one, and on hundreds of thousands of hours of video content focused on human and robot manipulation tasks to be ready as a backbone for visual intelligence. This scaling is what translates the success of Self-Flow from the lab to reality. mimic deployed FLUX-mimic in real factory use cases spanning the daily reality