메뉴
BL
Wired AI • 4일 전

AI 챗봇 '아폴로', 훼손된 그리스 파피루스의 비밀을 풀다

IMP
6/10
핵심 요약

오스트리아 과학아카데미가 프랑스 AI 기업 미스트랄(Mistral) 등과 함께 고대 그리스어 전용 대규모 언어 모델 '아폴로(Apollo)'를 공개했습니다. 약 6억 단어의 사본·파피루스·비문 데이터로 학습된 이 모델은 훼손된 고문서의 빈칸을 통계적으로 가장 유력한 단어로 채워 복원 작업을 가속화하며, 향후 라틴어나 이집트어 등 다른 고대 언어에도 적용될 수 있습니다.

번역된 본문

전 세계 학술 도서관에는 수십만 점의 고대 그리스어 파피루스 조각이 소장되어 있다. 많은 조각이 손상되어 그 의미를 영원히 알 수 없지만, 학자들은 빠진 단어나 구절을 체계적으로 채워 넣어 나머지를 복원할 수 있다. 이 힘든 작업을 가속화하기 위해 연구자들이 인공지능(AI)에 주목했다. 수요일, 오스트리아 과학아카데미는 프랑스 AI 연구소 미스트랄(Mistral)과 기술 서비스 기업 세일 리플라이(Sail Reply)와 함께 개발한 "세계 최초의 고급 고대 그리스어 대규모 언어 모델"을 공개한다. 이 모델 '아폴로(Apollo)'는 사본, 파피루스, 비문에서 추출한 약 6억 개의 역사적 그리스어 단어로 학습되었다. 이 모델은 챗봇 인터페이스를 통해 학계에 무료로 제공된다.

그 목표는 학자들이 자신의 세부 전공 분야와 관련된 파피루스 조각과 유망한 새 연구 방향을 더 빠르게 찾도록 돕는 것이다. 문서가 찢기고 해져 있을 때 아폴로는 통계적으로 가장 유력한 단어나 구절로 빈칸을 채우도록 설계되어, 역사적 사건과 관행에 관한 숨겨진 세부 정보를 밝혀낼 수 있다. 세일 리플라이의 파트너 디미트리스 블리타스(Dimitris Vlitas)는 이런 방식으로 지식을 해방시키는 것은 "1년 전에는 상상할 수도 없었다"고 WIRED에 말했다.

지금까지 훼손된 파피루스를 복원하려면 숙련된 학자가 먼저 단어 구분을 식별하고(고대 그리스어에는 띄어쓰기가 없다)—문서의 시기를 정확히 추정하고, 적절한 사회정치적 맥락을 고려하며, 참고 자료를 검토하여 빈칸에 들어갈 적합한 단어를 선택해야 했다. 유니버시티 칼리지 런던(UCL)의 고전학 및 역사언어학 교수 스티븐 콜빈(Stephen Colvin)은 "그리스 역사에서 그 정도로 능숙한 사람은 세계에 매우 드물다"고 말한다. 하지만 이러한 전문 지식이 모두 아폴로에 내장되어 있다. 오스트리아 과학아카데미의 역사학자이자 파피루스학자인 안나 돌가노프(Anna Dolganov)는 "호메로스를 보면 호메로스식 그리스어로 보완하고, 도리스 방언 비문을 보면 도리스 방언을 사용한다"고 설명한다.

고된 복원 작업에 매달려 온 학자들은 아폴로가 작업을 가속화하여, 문서에 무엇이라고 쓰여 있는지 파악하는 대신 역사 문서의 의미와 함의에 집중할 수 있게 해주기를 기대한다. 세계 최대 고대 파피루스 컬렉션을 보유한 옥스퍼드 대학의 고전 언어문학 교수 아르망 당구르(Armand D'Angour)는 "매우 흥미진진하다"며 "'이 빈칸에 들어갈 수 있는 세 가지 단어가 여기 있다'고 알려주는 기계가 있다면 작업이 상당히 빨라질 것"이라고 말했다.

아폴로가 고대 세계에 대한 전반적인 이해를 바꾸지는 못할 것이다. 많은 파피루스가 복원되지 못한 것은 바로 그 내용이 평범하기 때문이다—개인 서신, 혼인 계약서, 공무 문서 등이다. 콜빈은 "일반인이라면 갑자기 소포클레스의 새로운 희곡 몇 편이 나올 것이라고 생각할 수 있지만, 그런 일은 일어나지 않을 것"이라고 말한다. 그러나 이 모델은 고대 생활에 관한 새로운 세부 사항을 밝혀내고 기존 학계의 가정을 뒷받침하는 데 기여할 수 있다. 당구르는 "무언가 새로 생성될 때마다 고대 세계에 대한 작은 지식이 하나씩 추가된다"고 말했다.

블리타스에 따르면 아폴로가 성공한다면 동일한 기법을 라틴어나 이집트어 같은 다른 고대 언어에, 또는 대량의 자료를 정리·색인화하면 도움이 되는 다른 학문 분야에도 쉽게 적용할 수 있다. AI는 일부 분야에서 이미 눈에 띄는 성과를 거두었다. OpenAI는 최근 자사 AI 모델이 200년 된 수학 문제를 해결했다고 밝혔고, 구글 딥마인드(Google DeepMind)는 AI를 활용해 유전자 변이가 분자생물학에 미치는 영향을 매핑한 방대한 데이터셋을 공개했다. 한 가지 우려는 확률에 기반한 언어 모델로 고대 문서의 빈칸을 채울 경우 오류로 역사 기록이 오염될 위험이 있다는 점이다. 이를 방지하기 위해 아폴로는 몇 가지 후보를 제안하도록 설계되었다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Academic libraries across the globe are stuffed with hundreds of thousands of Ancient Greek papyrus fragments. Though many are so damaged that their meaning is probably lost, scholars have the ability to restore the rest by methodically filling in missing words or phrases. To accelerate that laborious task, researchers have turned to artificial intelligence . On Wednesday, the Austrian Academy of Science will release “the world’s first advanced large language model for Ancient Greek,” developed in partnership with French AI lab Mistral and technology services firm Sail Reply. The model, Apollo, is trained on roughly 600 million historical Greek words drawn from manuscripts, papyri, and inscriptions. The model will be freely available to academics through a chatbot interface. The ambition is to help scholars to more rapidly identify papyrus fragments relevant to their specific sub-disciplines, as well as promising new avenues of research. Where documents are tattered and torn, Apollo is built to fill in the blanks with the most statistically likely words or passages, potentially revealing hidden details about historical events and practices. Dimitris Vlitas, partner at Sail Reply, tells WIRED that unlocking knowledge in this way “was unthinkable a year ago.” Until now, restoring a tattered piece of papyrus has required a skilled academic to first identify the word divisions—there are no gaps in Ancient Greek writing—then accurately date the document, weigh the appropriate socio-political contexts, and consult reference materials to help choose suitable words to fill in the gaps. “There are very few people in the world who are that good at Greek history,” says Stephen Colvin, a professor of classics and historical linguistics at University College London. But all of that specialized knowledge is baked into Apollo. “When it sees Homer, it supplements Homeric Greek. When it sees an inscription in Doric dialect, it uses Doric dialect,” says Anna Dolganov, a historian and papyrologist at the Austrian Academy of Science. Academics who find themselves bogged down in painstaking reconstruction work expect Apollo to accelerate things, allowing them to focus on the implications of historical documents, rather than figuring out what they say. “I think it’s very exciting,” says Armand D'Angour, a professor of classical languages and literature at the University of Oxford, home to the world’s largest ancient papyrus collection. “If I had a machine telling me, ‘Here are the three possible words that could fit into that gap,’ it would speed up matters considerably.” Apollo is unlikely to change the broad-strokes understanding of the ancient world; many papyri are yet to be restored precisely because they are mundane—personal letters, marital contracts, civil service papers. “If you were a layperson, you might think suddenly we’ll get a few new plays by Sophocles, but that’s not going to happen,” Colvin says. However, the model could help to uncover new details about life in antiquity and substantiate existing scholarly assumptions. “Every time something is produced, it adds a tiny element of knowledge about the ancient world,” D’Angour says. If Apollo is a success, says Vlitas, the same technique could be readily applied to other ancient languages—Latin or Egyptian, say—or any other academic discipline that would benefit from the distillation and indexing of a large corpus of material. AI has had notable success in some areas; OpenAI recently said its AI models solved a 200-year-old math problem, while Google DeepMind released a vast dataset that maps how genetic mutations affect molecular biology, which it compiled using AI. One concern might be that relying on a language model—which deals in probabilities—to fill in gaps in ancient documents risks polluting the historical record with errors. But to head off that issue, Apollo is built to propose a selection of word options for a scholar to select between. “The crucial point is that human competence needs to remain,” says Dolganov. “If we become totally reliant on AI transcriptions and interpretations of historical material, that’s when the problems start.”