메뉴
BL
The Decoder • 18일 전

알리바바 Qwen-Drive 1.0, 주행·설명 하나의 AI로…단, 설명과 조작이 안 맞을 수도

IMP
7/10
핵심 요약

알리바바가 자율주행용 멀티모달 모델 Qwen-Drive 1.0을 공개했다. Qwen3.5-4B를 기반으로 3D 공간 지각 모듈과 경로 계획 전문가(Planning Expert) 모듈을 추가해, 하나의 모델로 주행 시스템과 콕핏 어시스턴트를 모두 수행한다. 시뮬레이션에서 차선 이탈률을 24%에서 12%로 줄였으나, 모델의 설명이 실제 주행 결정과 항상 일치하지는 않는 한계가 있다.

번역된 본문

Qwen-Drive 1.0은 왜 제동하는지 설명해주지만, 그 설명이 실제 조작과 일치하리라 기대하긴 어렵다.

핵심 포인트

  • 알리바바의 Qwen-Drive 1.0은 하나의 AI 모델로 공간 지각, 교통 관련 질의응답, 경로 계획을 처리한다.
  • 3D 매핑과 경로 계획을 위한 두 개의 모듈이 기존 언어 모델을 확장하여, 기존 지식을 잃지 않으면서 주행 시스템과 콕핏 어시스턴트로 모두 작동할 수 있게 한다.
  • 재학습을 통해 시뮬레이션에서 차가 도로를 이탈하는 비율이 24%에서 12%로 감소했다.
  • 그러나 모델의 설명이 실제 내리는 주행 결정과 항상 일치하지는 않는다.

Qwen-Drive 1.0은 세 가지 작업을 하나의 AI 모델로 처리한다: 주변 환경의 공간 지각, 교통 상황에 대한 질의응답, 그리고 경로 계획이다.

연구진은 텍스트-이미지 모델이 그림을 묘사할 수 있다고 해서 3차원 공간을 자동으로 이해하는 것은 아니라는 점을 확인했다. 기존 주행 모델들은 일반 텍스트-이미지 모델을 주행 데이터, 주로 교통 상황에 대한 질의응답 쌍으로 파인튜닝한다. 논문에 따르면 이 접근법에는 두 가지 약점이 있다. 교통 Q&A 위주로 학습된 모델은 여전히 거리, 위치, 개방 공간을 신뢰할 수 있게 감지하지 못한다. 또한 주행 데이터에 너무 특화되면 연구진이 '치명적 망각(catastrophic forgetting)'이라 부르는 현상으로 원래 학습의 폭넓은 일반 지식을 잃게 되는데, 바로 그런 지식이 드물고 예상치 못한 교통 상황에서 가장 중요하다.

별도 모듈로 모델이 실제로 공간을 이해하는지 검증

Qwen-Drive 1.0은 2월에 출시된 Qwen3.5-4B를 기반으로 두 개의 추가 구성요소를 더했다. 첫 번째는 3D 공간에서 물체를 감지하고, 어떤 영역이 점유되어 있는지 파악하며, 도로 구조를 추적하여 주변의 조감도(bird's-eye-view) 지도를 생성한다. 연구진에 따르면 이 모듈은 모델이 이미지에서 실제로 얼마나 많은 공간 정보를 추출하는지 드러내는 측정 도구로도 기능한다.

두 번째 구성요소인 '플래닝 전문가(Planning Expert)'는 모델 내부 데이터를 사용해 차량의 향후 움직임을 계획한다. 연구진이 추가된 구성요소만 학습시키고 비전-언어 모델은 그대로 둘 경우 공간 정확도가 낮게 유지되어, 이미지를 상세히 묘사할 수 있는 모델이라도 3차원 공간을 자동으로 파악하지는 못한다는 것이 확인됐다. 팀이 비전-언어 모델 자체를 공간 과제에 대해 학습시켰을 때에만 성능이 크게 향상되었으며, 이는 교통 장면에 대한 공간적 이해 능력이 의도적으로 구축되어야 함을 의미한다.

학습은 지각 모듈로 시작해 지각과 질의응답을 결합하고, 이어 경로 계획으로 진행된다. 마지막 단계에서는 강화학습으로 모델의 행동을 다듬는다. 비전-언어 구성요소를 위해 팀은 공개된 24개의 교통 장면 데이터셋을 결합했다. 이 데이터셋들은 구조가 서로 달랐고 오류를 포함하기도 했다. AI 모델이 질문과 답변을 표준화하고 원본 데이터와 정렬했다. 팀은 또한 어떤 물체가 제동을 유발하는지 등 차가 특정 주행 결정을 내려야 하는 이유를 설명하는 자체 예시를 구축했다.

콕핏과 주행 시스템을 위한 하나의 모델

최신 차량에서는 인포테인먼트 시스템과 주행 시스템이 두 개의 분리된 컨트롤러에서 작동하는 대신 하나의 컴퓨팅 유닛으로 통합되고 있다. 저자들은 일반적 능력을 순수한 주행 성능과 맞바꾼 모델은 여기서 도움이 되지 않는다고 주장한다. 콕핏이 대화나 개방형 질문 같은 작업을 위해 여전히 별도의 모델과 추가 컴퓨팅이 필요하기 때문이다.

논문에 따르면 Qwen-Drive 1.0은 교통 장면 관련 질문에서 변형하지 않은 기본 모델 Qwen3.5-4B보다 훨씬 높은 점수를 기록했다. 가장 큰 격차는 차가 왜 제동하거나 회전해야 하는지 등 인과관계를 설명해야 할 때 나타난다. 일반 지식도 유지된다.

원문 보기
원문 보기 (영어)
Qwen-Drive 1.0 tells you why it brakes, just don't expect the explanation to match the maneuver Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 7, 2026 Qwen Key Points Alibaba's Qwen-Drive 1.0 handles spatial perception, traffic questions, and route planning in one AI model. Two modules for 3D mapping and route planning extend the base language model, letting the AI run as both a driving system and a cockpit assistant without losing existing knowledge. Retraining cut the rate at which the car veered off the road from 24 percent to 12 percent in simulations. But the model's explanations don't always match the driving decisions it makes. Ask about this article… Search Qwen-Drive 1.0 handles three tasks in one AI model: spatial perception of the environment, answering questions about traffic, and route planning. The researchers confirm that a text-image model doesn't automatically understand three-dimensional space just because it can describe pictures. Existing driving models take a general text-image model and fine-tune it on driving data, mostly question-and-answer pairs about traffic situations. According to the paper, this approach has two weaknesses. A model trained mainly on traffic Q&A still can't reliably detect distances, positions, and open spaces. And if it becomes too specialized on driving data, it loses the broad general knowledge from its original training through what the researchers call "catastrophic forgetting," which is exactly the kind of knowledge that matters most in rare, unexpected traffic situations. The model Alibaba's research division built is supposed to address both problems. A separate module checks whether the model actually understands space Qwen-Drive-1.0 builds on Qwen3.5-4B, released in February, and adds two extra components. The first generates a bird's-eye-view map of the surroundings by spotting objects in 3D space, figuring out which areas are occupied, and tracing the road layout. The researchers say it doubles as a measuring tool that reveals how much spatial information the model is actually pulling from the images. Ad The second component, the Planning Expert, uses internal model data to plan the car's future movement. When the researchers trained only the added component and left the vision-language model untouched, spatial accuracy stayed low, confirming that a model capable of describing images in detail doesn't automatically grasp three-dimensional space. Only when the team also trained the vision-language model itself on spatial tasks did performance improve significantly, meaning the ability to spatially understand traffic scenes has to be built in deliberately. Ad Training starts with the perception module, then combines perception and question answering, followed by route planning. The final step refines the model's behavior through reinforcement learning. For the vision-language component, the team combined 24 publicly available datasets of traffic scenes. These datasets had different structures and sometimes contained errors. An AI model standardized the questions and answers and aligned them with the original data. The team also built its own examples explaining why the car should make a specific driving decision, like which object triggers braking. One model for both the cockpit and the driving system In modern vehicles, the infotainment system and the driving system are converging on a single computing unit instead of running on two separate controllers. The authors argue that a model trading general capabilities for pure driving performance doesn't help here, because the cockpit would still need its own model and extra compute for tasks like dialog or open-ended questions. Ad Qwen-Drive 1.0 scores well above the unmodified base model Qwen3.5-4B on questions about traffic scenes, according to the paper. The biggest gap shows up when the model has to explain cause and effect, like why the car should brake or turn. General knowledge holds up too: The model shows almost no drop on tests outside of driving and even scores slightly higher on some spatial tasks. From simulation to the road For driving planning, the team tested the model at several difficulty levels, from simple predictions up to a simulator where errors compound over time. In the simulator, the version retrained with rewards cut the rate at which the car veered off the road from 24 to 12 percent. It also drove more cautiously and covered less distance overall. Ad The model's explanations don't always pinpoint the actual cause of a situation, though. A red light in the distance and a child stepping into the road call for very different reaction times, and the model can conflate the two. The planned maneuver also doesn't always match the reasoning the model gave beforehand. Some results rest on test procedures the authors designed or rebuilt themselves, so individual metrics say little about how the system would handle messy real-world driving. Ad The Qwen team measured this gap in spatial understanding itself using its HopChain benchmark . Here, vision-language models misclassified objects and confused spatial relationships even while scoring well on image-text benchmarks. Catastrophic forgetting is a familiar problem, too. When Google Deepmind built PaLM-E in 2023 , a model for language, images, and robot control, the smaller variants lost a big chunk of their language ability after robot training. The largest version, at 562 billion parameters, lost almost nothing. Pairing a language model with a driving function opens up a new attack surface. Researchers at UC Santa Cruz placed a labeled sign in the camera's field of view and tricked the DriveLM driving system into swerving toward crossing pedestrians , even though it had detected them correctly. Research is meanwhile shifting toward World Action Models , which also predict how the environment changes in response to the agent's own actions. The Qwen team is releasing the model to the research community for free on Hugging Face , ModelScope, and GitHub . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Arxiv