메뉴
BL
The Decoder 53일 전

알리바바 '큐웬3.7-플러스', 멀티모달 자율형 에이전트로 진화

IMP
8/10
핵심 요약

알리바바가 시각적 인지와 에이전트 기능을 결합한 멀티모달 모델 '큐웬3.7-플러스(Qwen3.7-Plus)'를 출시했습니다. 이 모델은 그래픽 사용자 인터페이스(GUI)를 자율적으로 조작하고 1만 줄 이상의 코드를 독립적으로 작성하는 등 뛰어난 작업 자동화 능력을 보여줍니다. 복잡한 논리적 추론 벤치마크에서는 경쟁사 최상위 모델들에 미치지 못하지만, 가격 경쟁력을 갖춘 독점 모델로서 알리바바 클라우드를 통해 서비스됩니다.

번역된 본문

알리바바의 큐웨(Qwen) 팀이 텍스트 전용 모델인 Qwen3.7을 기반으로 구축된 멀티모달 모델, Qwen3.7-Plus를 발표했습니다. 이 모델은 시각적 인식 능력과 코딩 및 도구 사용과 같은 기존 에이전트 기능을 결합했습니다.

'멀티모달 인터랙티브 하이브리드 에이전트'로 불리는 이 모델은 실제 세계의 장면을 인식하고, 화면 콘텐츠를 읽으며, 그래픽 인터페이스를 조작하고, 시각적 템플릿에서 코드를 작성하고, 모바일 앱을 처음부터 끝까지 탐색하도록 설계되었습니다. UI 클릭 및 명령줄(Command-line) 명령은 동일한 에이전트 루프 내에서 실행됩니다.

11시간의 자율 앱 개발 큐웨 팀은 Qwen3.7-Plus를 사용하여 하이브리드 에이전트 시스템이 영어 어휘 학습 앱을 구축하도록 했습니다. 큐웨에 따르면, 이 에이전트는 11시간 이상 실행되었으며 1,000회 이상의 에이전트 호출을 통해 10,000줄 이상의 코드를 생성했습니다. 이 과정에는 요구사항 문서화, 자동 코드 생성, 설치, 테스트 케이스 생성, GUI 기반 테스트, 병렬 테스트 시나리오 및 독립적인 버전 관리가 포함되었습니다.

두 번째 데모는 데스크톱 앱을 대상으로 합니다. 보고에 따르면 에이전트는 macOS 기본 주식 앱을 자율적으로 조작하여 UI 구조를 분석하고 이로부터 SwiftUI 코드를 생성하여 앱을 재구성했습니다. 그런 다음 실시간 주식 데이터를 위해 외부 API를 연결하고, 앱을 컴파일하며, 가격 조회 및 검색 필터를 포함하여 10개의 기능 테스트를 독자적으로 실행했습니다.

세 번째 사용 사례는 사이드바 확장 프로그램인 'Qwen for Chrome'을 통한 브라우저 에이전트를 보여줍니다. 사용자의 허가를 받으면 모델이 에이전트 모드로 전환되어 클라우드 콘솔에서 작업을 수행합니다. 예를 들어 이미지, 스토리지 및 보안 그룹을 구성하여 사용 가능한 가장 저렴한 가상 서버 인스턴스를 구매하는 작업이 가능합니다. 큐웨는 후속 작업에서 에이전트가 스케일링 및 유지 관리도 처리한다고 밝혔습니다.

GUI 작업은 뛰어나지만 하드 추론 테스트는 부족해 큐웨가 공개한 벤치마크는 이 모델이 그래픽 인터페이스 조작에 탁월하다는 명확한 그림을 보여줍니다. AndroidWorld 및 ScreenSpot Pro에서 Qwen3.7-Plus는 GPT-5.4 (xhigh), Opus 4.6 Max 및 Gemini 3.1 Pro를 크게 앞서고 있습니다. 또한 에이전트 지향적인 터미널 작업 및 장기적 작업 계획에서도 선도적인 성능을 보입니다.

반면, 전통적인 멀티모달 추론에서는 결과가 엇갈립니다. Qwen3.7-Plus는 일부 시각적 추론 테스트에서 1위를 차지하지만, MedXpertQA-MM과 같은 더 까다로운 과학 작업에서는 Gemini 3.1 Pro 및 GPT-5.4에 미치지 못합니다. 텍스트 측면에서 팀은 성능이 최상위 모델과 동등하다고 설명하지만, 전반적으로 이들을 앞지르지는 못합니다.

크로스 프레임워크 호환성이 차별점 Qwen3.7-Plus는 Anthropic API 프로토콜을 지원하며 Claude Code, OpenClaw 및 알리바바 자체의 Qwen Code와 직접 작동합니다. 또한 API는 이전 대화 턴의 추론 콘텐츠를 유지하는 'preserve_thinking'라는 기능을 제공합니다. 큐웨 팀은 에이전트 작업에 대해 이 설정을 명시적으로 권장합니다.

이미지 처리 외에도 모델은 비디오 이해 및 자율주행 장면 분석을 다루어 임베디드 시스템 및 자율주행을 위한 기반으로 자리매김하고 있습니다.

Qwen3.7-Plus는 알리바바 클라우드 모델 스튜디오를 통해 사용할 수 있으며 텍스트 기반 형제 모델인 Qwen3.7-Max와 마찬가지로 오픈 가중치가 없는 독점 모델입니다. 알리바바는 Plus 등급의 가격을 Max보다 훨씬 낮게 책정했습니다. Qwen3.7-Max의 백만 입력 토큰당 $2.00, 백만 출력 토큰당 $6.00와 비교하여 Qwen3.7-Plus는 백만 입력 토큰당 $0.40, 백만 출력 토큰당 $2.40의 비용이 듭니다.

원문 보기
원문 보기 (영어)
Qwen3.7-Plus is Alibaba's bid to turn multimodal AI into a full-blown autonomous agent Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jun 6, 2026 Qwen Key Points Alibaba has released Qwen3.7-Plus, a new AI model that combines visual understanding with agent capabilities, enabling it to autonomously operate graphical user interfaces and apps. In testing, the system demonstrated its ability to recreate desktop applications, perform cloud tasks, and independently program a complete app with 10,000 lines of code. While Qwen3.7-Plus outperforms competitors in operating user interfaces, it falls short in pure logic benchmarks. The model is available as a proprietary, comparatively inexpensive option through Alibaba Cloud. Ask about this article… Search Alibaba's Qwen team has released Qwen3.7-Plus, a multimodal model built on top of the text-only Qwen3.7. It combines visual perception with classic agent capabilities like coding and tool use. Billed as a "multimodal interactive hybrid agent," the model is designed to recognize real-world scenes, read screen content, operate graphical interfaces, write code from visual templates, and navigate mobile apps end to end. UI clicks and command-line instructions run within the same agent loop. Eleven hours of autonomous app development Using Qwen3.7-Plus, the team had a hybrid agent system build an English vocabulary learning app. According to Qwen, the agent ran for over eleven hours, producing more than 10,000 lines of code across more than 1,000 agent calls. The process covered requirements documentation, automated code generation, installation, test case creation, GUI-based testing, parallel test scenarios, and independent version management. Ad A second demo targets desktop apps: the agent reportedly recreated the native macOS Stocks app by operating it autonomously, parsing the UI structure, and generating SwiftUI code from it. It then connected an external API for real-time stock data, compiled the app, and ran ten functional tests on its own, including price lookups and search filters. Ad DEC_D_Incontent-1 A third use case shows a browser agent via "Qwen for Chrome," a sidebar extension. With user permission, the model switches into agent mode and carries out tasks in a cloud console, like purchasing the cheapest available virtual server instance, including configuring the image, storage, and security groups. In a follow-up task, the agent also handles scaling and maintenance, Qwen says. GUI tasks shine, hard reasoning tests don't The benchmarks Qwen published paint a clear picture: the model excels at operating graphical interfaces. On AndroidWorld and ScreenSpot Pro, Qwen3.7-Plus sits well ahead of GPT-5.4 (xhigh) , Opus 4.6 Max , and Gemini 3.1 Pro . It also leads on agent-oriented terminal work and long-horizon task planning. Ad On classic multimodal reasoning, results are mixed. Qwen3.7-Plus tops some visual reasoning tests but falls short of Gemini 3.1 Pro and GPT-5.4 on tougher scientific tasks like MedXpertQA-MM. On the text side, the team describes performance as on par with max-tier models, without beating them across the board. Cross-framework compatibility sets it apart Qwen3.7-Plus supports the Anthropic API protocol and works directly with Claude Code , OpenClaw , and Alibaba's own Qwen Code . The API also offers a feature called preserve_thinking that retains reasoning content from earlier conversation turns. The Qwen team explicitly recommends this setting for agentic tasks. Ad DEC_D_Incontent-2 Beyond image processing, the model also covers video understanding and driving scene analysis, positioning it as a foundation for embedded systems and autonomous driving. Ad Qwen3.7-Plus is available through Alibaba Cloud Model Studio and, like its text-based sibling Qwen3.7-Max , is a proprietary offering with no open weights. Alibaba prices the Plus tier well below Max: Qwen3.7-Plus costs $0.40 per million input tokens and $2.40 per million output tokens, compared to $2.50 and $7.50 for Qwen3.7-Max. That makes Plus roughly six times cheaper on input and three times cheaper on output and well below the list prices of Western frontier models. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Qwen