메뉴
HN
Hacker News • 53일 전

대규모 언어 모델이 표 형식 예측에 실패하는 이유

IMP
8/10
핵심 요약

최신 대규모 언어 모델(LLM)이 가장 흔한 기계학습 작업 중 하나인 표 형식 데이터의 예측 분석에서 성능이 저조한 원인을 규명한 논문입니다. 연구 결과, 노이즈나 CSV 포맷 등의 문제가 아니라 '데이터 차원수(dimensionality)'가 증가할수록 LLM의 예측 능력이 급격히 상실되는 결정적인 한계가 발견되었습니다. 이는 최신 LLM이 수십 년 된 고전적 기계학습 모델에 비해 표 데이터 처리에 취약한 근본적인 이유를 설명해 주며, 전문 실무자들에게 매우 중요한 시사점을 던집니다.

번역된 본문

전산학 > 기계 학습 arXiv:2608.02412 (cs) [2026년 8월 3일 제출] 제목: 대규모 언어 모델이 표 형식 예측에 실패하는 이유 저자: Marta Garnelo, Wojciech M. Czarnecki 해당 논문 PDF 보기 (제목: 대규모 언어 모델이 표 형식 예측에 실패하는 이유, 저자: Marta Garnelo 및 기타 1인) HTML 보기 (실험적 기능)

초록: 대규모 언어 모델(LLM)은 매우 다양한 작업에서 기본적인 도구로 자리 잡았지만, 가장 보편적인 기계 학습 작업 중 하나인 표 형식 데이터에 대한 예측 분석 분야에서는 두드러지게 성공을 거두지 못했습니다. 이러한 격차는 급성장하는 표 형식 파운데이션 모델 분야의 출발 배경이 되었지만, 범용 LLM이 왜 실패하는지에 대한 질문은 여전히 미해결 상태로 남아 있었습니다. 우리는 최신 LLM을 가장 순수한 추론 환경(외부 도구나 에이전트 구조, 파인튜닝 없이 전체 학습 및 테스트 데이터가 포함된 프롬프트를 한 번에 처리하는 방식)에서 연구하고, 실패에 대한 5가지 가설을 체계적으로 평가했습니다. 5가지 가설은 다음과 같습니다: (a) 노이즈가 있거나 선형으로 분리할 수 없는 데이터를 처리하지 못하는 능력; (b) 선형화된 CSV 형식이 열 구조를 모호하게 만듦; (c) 숫자 값의 토큰화 문제; (d) 쿼리당 분류되는 테스트 포인트의 수; (e) 입력 데이터의 차원수. 통제된 실험 결과, 가설 (a)부터 (d)까지는 기각되었습니다. 반면 데이터의 차원수는 결정적인 요인으로 밝혀졌습니다. 31개의 벤치마크 데이터셋에 무작위 선형 투영을 적용해 본 결과, LLM은 테스트를 진행한 9가지 방법 중 유일하게 차원수가 커질수록 정확도가 감소했으며, 모든 고전적인 기준 모델은 정확도가 동일하게 유지되거나 오히려 향상되었습니다. 252개 설정의 고전적 모델과의 행동 비교 결과, 2차원에서 LLM은 국소적이고 거리 기반의 방법처럼 예측하며(최대 91.6%의 격자 일치율)를 보였습니다. 하지만 고차원에서는 튜닝된 차원 종속적 노이즈를 추가하더라도 LLM의 예측을 재현해내는 고전적 모델은 단 하나도 없었습니다. 우리가 LLM의 내부 메커니즘을 정확히 규명했다고 주장하는 것은 아닙니다. 다만, LLM의 기능이 어떤 노이즈가 적용된 고전적 학습 모델과도 다른 방식으로 차원수에 따라 붕괴한다는 것을 결과가 보여준다고 겸허하게 말씀드리고 싶습니다. 이는 왜 다른 곳에서는 매우 뛰어난 능력을 보이는 LLM이 표 데이터에서는 계속해서 50년 전의 고전적 모델에 패배하는지를 설명해 주며, 예측의 정확한 메커니즘은 향후 연구를 위한 열린 질문으로 남겨둡니다.

주제: 기계 학습 (cs.LG) 인용: arXiv:2608.02412 [cs.LG] (또는 현재 버전의 경우 arXiv:2608.02412v1 [cs.LG]) https://doi.org/10.48550/arXiv.2608.02412 자세히 알아보기 (DataCite를 통해 arXiv에서 발급한 DOI, 등록 대기 중) 제출 이력 보낸 사람: Marta Garnelo [이메일 보기] [v1] 2026년 8월 3일 월요일 15:52:51 UTC (10,846 KB) 전체 텍스트 링크: 논문 액세스 (제목: 대규모 언어 모델이 표 형식 예측에 실패하는 이유, 저자: Marta Garnelo 및 기타 1인), PDF 보기, HTML 보기 (실험적 기능), TeX 소스, 라이선스 보기 현재 탐색 컨텍스트: cs.LG < 이전 | 다음 > 신규 | 최근 | 2026-08 다음으로 탐색 변경: cs 참고 문헌 및 인용: NASA ADS, Google Scholar, Semantic Scholar 내보내기: BibTeX 인용 로딩 중... BibTeX 형식 인용 및 x, loading... 제공된 데이터: 북마크, 서지 도구, 서지 및 인용 도구, 서지 탐색기 토글 (탐색기란?), Connected Papers 토글 (Connected Papers란?), Litmaps 토글 (Litmaps란?), scite.ai 토글 (scite 스마트 인용이란 무엇인가?) 코드, 데이터, 미디어: 이 기사와 관련된 코드, 데이터 및 미디어, alphaXiv 토글 (alphaXiv란?), 코드 링크 토글 (논문용 CatalyzeX 코드 파인더, CatalyzeX란?), DagsHub 토글 (DagsHub란?), GotitPub 토글 (Gotit.pub이란?), Huggingface 토글 (Hugging Face란?), ScienceCast 토글 (ScienceCast란?) 데모: 데모, Replicate 토글 (Replicate란?), Spaces 토글 (Hugging Face Spaces란?), Spaces 토글 (TXYZ.AI란?) 관련 논문, 추천인 및 검색 도구: Influence Flower 링크, Influence Flower (Influence Flower란 무엇인가?)

원문 보기
원문 보기 (영어)
--> Computer Science > Machine Learning arXiv:2608.02412 (cs) [Submitted on 3 Aug 2026] Title: Why Large Language Models Fail at Tabular Prediction Authors: Marta Garnelo , Wojciech M. Czarnecki View a PDF of the paper titled Why Large Language Models Fail at Tabular Prediction, by Marta Garnelo and 1 other authors View PDF HTML (experimental) Abstract: Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question. Subjects: Machine Learning (cs.LG) Cite as: arXiv:2608.02412 [cs.LG] (or arXiv:2608.02412v1 [cs.LG] for this version) https://doi.org/10.48550/arXiv.2608.02412 Focus to learn more arXiv-issued DOI via DataCite (pending registration) Submission history From: Marta Garnelo [ view email ] [v1] Mon, 3 Aug 2026 15:52:51 UTC (10,846 KB) Full-text links: Access Paper: View a PDF of the paper titled Why Large Language Models Fail at Tabular Prediction, by Marta Garnelo and 1 other authors View PDF HTML (experimental) TeX Source view license Current browse context: cs.LG < prev | next > new | recent | 2026-08 Change to browse by: cs References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation &times; loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) scite.ai Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle Gotit.pub ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle TXYZ.AI ( What is TXYZ.AI? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) IArxiv recommender toggle IArxiv Recommender ( What is IArxiv? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs . Which authors of this paper are endorsers? | Disable MathJax ( What is MathJax? )