메뉴
HN
Hacker News 49일 전

LLM 기반 텍스트-CAD 생성·편집 통합 프레임워크

IMP
7/10
핵심 요약

기존에는 텍스트로 CAD 모델을 생성하는 작업과 이를 수정(편집)하는 작업이 분리되어 있어 실용성이 떨어졌습니다. 이 논문은 대형 언어 모델(LLM)을 활용해 모델의 생성과 편집을 하나로 통합한 'PR-CAD' 프레임워크를 제안합니다. 강화학습 기반의 에이전트를 통해 사용자의 의도를 파악하고 정밀하게 설계를 수정하여, 기존 대비 월등한 수준의 제어력과 정확도를 달성하며 CAD 모델링의 효율을 크게 높였습니다.

번역된 본문

컴퓨터 과학 > 계산 및 언어 arXiv:2604.19773 (cs) [2026년 3월 27일 제출]

제목: PR-CAD: 대형 언어 모델을 위한 점진적 정제를 통한 통합적이고 제어 가능하며 충실한 텍스트-CAD 생성 저자: Jiyuan An, Jiachen Zhao, Fan Chen, Liner Yang, Zhenghao Liu, Hongyan Wang, Weihua An, Meishan Zhang, Erhong Yang

초록: 전통적으로 CAD(컴퓨터 지원 설계) 모델의 구축은 노동 집약적인 수동 작업과 전문적인 지식에 의존해 왔습니다. 최근 대형 언어 모델(LLM)의 발전은 텍스트-투-CAD(Text-to-CAD) 생성 연구에 영감을 주었습니다. 그러나 기존의 접근 방식들은 생성(Create)과 편집(Edit) 작업을 서로 분리된 별개의 작업으로 취급하여 실용성을 제한했습니다.

본 논문에서는 제어 가능하고 충실한 텍스트-투-CAD 모델링을 위해 생성과 편집을 통합하는 점진적 정제(Progressive Refinement) 프레임워크인 PR-CAD를 제안합니다. 이를 뒷받침하기 위해 여러 CAD 표현 방식은 물론 정성적 및 정량적 설명을 아우르는, CAD의 전체 라이프사이클을 포괄하는 고품질의 상호작용 데이터셋을 큐레이션했습니다. 이 데이터셋은 편집 작업의 유형을 체계적으로 정의하고 인간의 실제 작업과 매우 유사한 상호작용 데이터를 생성합니다.

LLM에 최적화된 CAD 표현 방식을 기반으로, 의도 이해(Intent Understanding), 매개변수 추정(Parameter Estimation), 정밀 편집 위치 파악(Precise Edit Localization)을 단일 에이전트에 통합한 강화학습 기반 추론 프레임워크를 제안합니다. 이를 통해 설계 생성과 정제(수정) 작업을 모두 수행할 수 있는 '올인원(All-in-one)' 솔루션을 제공합니다.

광범위한 실험 결과, 생성과 편집 작업 간, 그리고 정성적 및 정량적 모달리티 간에 강력한 상호 보완 및 강화 효과가 있음을 입증했습니다. 공개 벤치마크에서 PR-CAD는 생성 및 정제 시나리오 모두에서 최고 수준의 제어 가능성(Controllability)과 충실성(Faithfulness)을 달성했으며, 사용자 친화적임과 동시에 CAD 모델링 효율성을 크게 향상시키는 것으로 입증되었습니다.

주제: 계산 및 언어 (cs.CL), 인공지능 (cs.AI) 인용: arXiv:2604.19773 [cs.CL] (또는 이 버전의 경우 arXiv:2604.19773v1 [cs.CL]) https://doi.org/10.48550/arXiv.2604.19773 DataCite를 통해 발급된 arXiv DOI

제출 이력: 출처: Jiyuan An [이메일 보기] [v1] 2026년 3월 27일 금요일 12:13:20 UTC (10,655 KB)

원문 보기
원문 보기 (영어)
--> Computer Science > Computation and Language arXiv:2604.19773 (cs) [Submitted on 27 Mar 2026] Title: PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models Authors: Jiyuan An , Jiachen Zhao , Fan Chen , Liner Yang , Zhenghao Liu , Hongyan Wang , Weihua An , Meishan Zhang , Erhong Yang View a PDF of the paper titled PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models, by Jiyuan An and 8 other authors View PDF HTML (experimental) Abstract: The construction of CAD models has traditionally relied on labor-intensive manual operations and specialized expertise. Recent advances in large language models (LLMs) have inspired research into text-to-CAD generation. However, existing approaches typically treat generation and editing as disjoint tasks, limiting their practicality. We propose PR-CAD, a progressive refinement framework that unifies generation and editing for controllable and faithful text-to-CAD modeling. To support this, we curate a high-fidelity interaction dataset spanning the full CAD lifecycle, encompassing multiple CAD representations as well as both qualitative and quantitative descriptions. The dataset systematically defines the types of edit operations and generates highly human-like interaction data. Building on a CAD representation tailored for LLMs, we propose a reinforcement learning-enhanced reasoning framework that integrates intent understanding, parameter estimation, and precise edit localization into a single agent. This enables an &#34;all-in-one&#34; solution for both design creation and refinement. Extensive experiments demonstrate strong mutual reinforcement between generation and editing tasks, and across qualitative and quantitative modalities. On public benchmarks, PR-CAD achieves state-of-the-art controllability and faithfulness in both generation and refinement scenarios, while also proving user-friendly and significantly improving CAD modeling efficiency. Subjects: Computation and Language (cs.CL) ; Artificial Intelligence (cs.AI) Cite as: arXiv:2604.19773 [cs.CL] (or arXiv:2604.19773v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2604.19773 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Jiyuan An [ view email ] [v1] Fri, 27 Mar 2026 12:13:20 UTC (10,655 KB) Full-text links: Access Paper: View a PDF of the paper titled PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models, by Jiyuan An and 8 other authors View PDF HTML (experimental) TeX Source view license Current browse context: cs.CL < prev | next > new | recent | 2026-04 Change to browse by: cs cs.AI References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation &times; loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) scite.ai Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle Gotit.pub ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle TXYZ.AI ( What is TXYZ.AI? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs . Which authors of this paper are endorsers? | Disable MathJax ( What is MathJax? )