메뉴
BL
The Decoder 7일 전

알리바이 Qwen-Image-3.0 출시, 10픽셀 텍스트 완벽 렌더링

IMP
8/10
핵심 요약

알리바이 Qwen 팀은 정보 밀집형 인포그래픽과 신문 레이아웃을 한 번에 생성하는 실용적인 이미지 생성 모델 'Qwen-Image-3.0'을 공개했습니다. 이 모델은 4,500 토큰의 프롬프트를 처리하여 10픽셀 수준의 작은 텍스트와 수학 공식(LaTeX)을 오차 없이 구현하며 12개 언어를 지원합니다. 현재는 초대받은 사용자만 API를 통해 접근할 수 있으며, 향후 Qwen Chat 등에 통합될 예정입니다.

번역된 본문

알리바이(Alibaba)의 Qwen-Image-3.0은 단 한 번의 실행만으로 완전한 형태의 인포그래픽 그리드와 10픽셀 크기의 읽기 쉬운 텍스트를 렌더링합니다.

핵심 요약

  • 알리바이의 Qwen 팀은 신문 레이아웃, 복잡한 인포그래픽 및 정보 밀집도가 높은 시각적 콘텐츠 등 실제 실무에 적용할 수 있도록 설계된 이미지 생성 모델 'Qwen-Image-3.0'을 공개했습니다.
  • 이 모델은 최대 4,500 토큰의 입력을 처리하며, 단 한 번의 실행으로 10픽셀 크기의 작은 텍스트, 수학 공식, 12개 국어의 언어를 읽을 수 있을 정도로 완벽하게 렌더링합니다.
  • Qwen-Image-3.0은 현재 초대형 API(API) 접근 방식으로만 제공되며, 향후 Qwen Chat(Qwen Chat) 등 자사 앱에 통합될 예정입니다.
  • 기존 첫 번째 버전의 Qwen-Image와 달리, 이번 모델 가중치는 오픈 소스 라이선스로 공개될 가능성이 낮습니다.

Qwen-Image-3.0은 한 번의 실행으로 여러 패널로 구성된 인포그래픽을 생성하고, 10픽셀 크기의 텍스트까지 선명하게 읽히도록 만들며, 12개 국어를 지원합니다. AI가 생성한 학술 논문이나 신문 페이지가 단순 정적 이미지로서 얼마나 유용할지는 여전히 논쟁거리입니다.

알리바이의 Qwen 팀은 이미지 생성 모델의 세 번째 버전인 'Qwen-Image-3.0'을 발표했습니다. 팀에 따르면, 첫 번째 버전은 "정밀함"에 초점을 맞췄고, 두 번째 버전은 "정밀함, 다양성, 완성도, 아름다움 및 진정성"을 목표로 했습니다. 이번 버전은 단어 그대로 "실제(Real)"로 요약됩니다. 이 모델은 단순히 매력적인 이미지를 넘어 신문 레이아웃, 스토리보드, 시험지와 같은 실질적인 실무 작업을 처리하도록 고안되었습니다.

더 길어진 프롬프트로 단 한 번의 실행에 복잡한 레이아웃 구현 Qwen-Image-3.0은 최대 4,500 토큰의 프롬프트를 지원합니다. 개발팀에 따르면, 이를 통해 모델이 여러 장의 이미지를 조합하는 대신 단 한 번의 실행에도 정보 밀도 높은 레이아웃을 충분히 생성할 수 있는 여유를 갖게 됩니다.

공개된 데모 중 하나는 9개의 개별 인포그래픽을 3x3 격자 형태로 배치한 것으로, 각 패널에는 고유한 텍스트, 수식 및 삽화가 포함되어 있습니다. 패널의 내용은 터널 주변의 안전 차간 거리, 수직선, 감정과 이성에 대한 유교적 가르침, 회전하는 실린더에서 발사체가 분리되는 속도 등을 다룹니다. 또 다른 패널에서는 간흡충의 생활사, 우측 흉통, 위수 72의 군집에 대한 실로우 정리(Sylow theorems), 은행의 내부 통제, 동식물 세포 내의 DNA 등을 설명합니다.

팀은 이 모델이 중첩된 인터페이스(화면)를 어떻게 처리하는지도 보여줍니다. 한 예시는 Qwen Chat(Qwen Chat) 화면이 띄워진 VSCode 창으로 시작됩니다. 그 화면 내부에는 위챗(WeChat) 대화창이 있으며, 그 안에는 핸드드립 커피 만드는 법을 설명하는 포스터가 포함되어 있습니다.

10픽셀 텍스트와 LaTeX 수식으로 렌더링 정확도 강화 Qwen에 따르면 이 모델은 10픽셀 크기의 작은 텍스트도 읽을 수 있을 만큼 선명하게 구현합니다. 데모 예시에는 텍스트가 빼곡히 들어간 고래상어 인포그래픽과 가상의 대수기하학 논문 한 페이지 전체가 포함되어 있습니다. 해당 논문 이미지에는 아래 첨자, 위 첨자, 중괄호, 분수, 합계 및 곱셈 기호가 포함된 여러 줄의 LaTeX(LaTeX) 수식이 담겨 있습니다. 다른 데모에서는 시뮬레이션된 신문 페이지와 교사의 노트처럼 보이는 빨간색 필체의 손글씨 코멘트를 선보입니다.

또한 Qwen-Image-3.0은 초상화나 사물의 모공, 피부 질감, 개별 머리카락과 같은 사진 수준의 세밀한 디테일 표현도 추구합니다. 또 다른 편집 데모에서는 독수리가 싸우는 전통 수묵화의 훼손된 부분을 복원합니다. 이 과정에서 원래의 붓터치와 먹물 음영에 완벽하게 맞춰 손상된 부분을 채워 넣습니다.

다국어 지원 및 UI 목업으로 활용 범위 확장 Qwen은 이번 모델의 세 번째 핵심 목표를 "깊은 지식"으로 설명합니다. 이 모델은 일본어, 한국어, 스페인어를 포함한 12개 언어를 기본적으로 지원합니다. 공개된 예시는 웹사이트, 게임 및 라이브 스트리밍의 인터페이스(UI)를 재현하는 모습도 보여줍니다.

한 편집 데모에서는 곤충 사진을 분류학, 신체 특징 라벨, 확대 상세 보기 및 축척 바가 포함된 완전한 생물 동정판(identification plate)으로 변환해 보여줍니다.

Qwen에 따르면 이 모델은 인터넷의 실시간 데이터를 끌어와 사용할 수도 있으며, 이를 이용해 항저우의 날씨 예보와 같은 콘텐츠를 생성할 수 있습니다. 또 다른 예시는 중국 전통 수묵화를 배치하는 모습 등을 보여줍니다.

원문 보기
원문 보기 (영어)
Alibaba's Qwen-Image-3.0 renders full infographic grids and readable ten-pixel text in a single pass Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jul 21, 2026 Alibaba Key Points Alibaba's Qwen team has released Qwen-Image-3.0, an image generator built for practical applications like newspaper layouts, complex infographics, and other information-dense visual content. The model processes inputs of up to 4,500 tokens and renders text as small as ten pixels, mathematical formulas, and twelve languages in a legible way in a single pass. Qwen-Image-3.0 is currently available only through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon. Unlike the original Qwen-Image, it is unlikely that the model weights will be released under an open license. Ask about this article… Search Qwen-Image-3.0 can render multi-panel infographics in one pass, produce legible text as small as ten pixels, and write in twelve languages. Whether AI-generated academic papers and newspaper pages are useful as static images remains an open question. Alibaba's Qwen team has released Qwen-Image-3.0, the third version of its image generator. According to the team, the first version focused on "precision," while the second targeted "precision, variety, completeness, beauty, and authenticity." This time, Qwen sums up its goal with one word, "Real." The model is meant to handle practical work such as newspaper layouts, storyboards, and exam sheets, not just produce attractive images. Longer prompts let the model build complex layouts in one pass Qwen-Image-3.0 accepts prompts of up to 4,500 tokens. According to the team, that gives the model enough room to create dense layouts in one pass rather than assemble them from several images. Ad One demo packs nine separate infographics into a 3 x 3 grid, each with its own text, formulas, and illustrations. The panels cover safe following distances near tunnels, perpendicular lines, a Confucian lesson about emotion and reason, and the detachment speed of a projectile from a rotating cylinder. Other panels explain the liver fluke life cycle, right-sided chest pain, Sylow theorems for groups of order 72, internal controls at banks, and DNA in animal and plant cells. Ad DEC_D_Incontent-1 The team also shows how the model handles nested interfaces. One example starts with a VSCode window containing a Qwen Chat screen. Inside that screen is a WeChat conversation, which includes a poster explaining how to make pour-over coffee. Ten-pixel text and LaTeX formulas push rendering fidelity Qwen says the model can produce legible text as small as ten pixels. Its examples include a whale shark infographic packed with text and a full page from a fictional algebraic geometry paper. The paper contains multi-line LaTeX equations with subscripts, superscripts, braces, fractions, sums, and products. Other demos show a simulated newspaper page and red handwritten comments that resemble notes from a teacher. Ad Qwen-Image-3.0 also aims for photographic detail in portraits and objects, including visible pores, skin texture, and individual strands of hair. In another editing demo, the model repairs a damaged traditional ink painting of fighting eagles. It fills in the missing areas while matching the original brushwork and ink shading. Language support and UI mockups broaden the model's range Qwen describes the third area of focus as "deep knowledge." The model supports twelve languages natively, including Japanese, Korean, and Spanish. The published examples also show it recreating interfaces from websites, games, and livestreams. In one editing demo, the model turns an insect photo into a full identification plate with taxonomy, labels for physical features, enlarged detail views, and a scale bar. Ad DEC_D_Incontent-2 The model can also pull in live internet data, according to Qwen, and uses it to generate things like a weather forecast for Hangzhou. Another example places Chinese ink painter Qi Baishi and Vincent van Gogh together in a simulated livestream studio. Ad Alibaba released the direct predecessor, Qwen-Image-2.0 , just this past May. The technical report focused on training and inference efficiency gains, including a faster variant that needed only four instead of 40 steps per image. In tests on Alibaba's own arena platform, Qwen-Image-2.0 landed just behind OpenAI's GPT-Image-2 and Google's Nano Banana Pro . For now, Qwen-Image-3.0 appears to be available only through invite-only API access. The model should show up in first-party apps like Qwen Chat soon. It's unlikely that the model weights will ship under an open license, as they did for the original Qwen-Image . Impressive tech, but the use cases don't always add up The practical value of some demos remains unclear. Researchers typically write and typeset papers in LaTeX rather than render them as images, so AI-generated pages with formulas may be better suited to mockups and visual drafts than final papers. A similar question applies to newspaper pages and complex infographics. Modern image models can edit individual text elements, but searchable and editable formats still offer more flexibility for production work. Even so, the examples show how far text rendering has advanced and may point to useful applications beyond the demos shown here. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Qwen