메뉴
HN
Hacker News 42일 전

아랍어의 디지털 표기 문제

IMP
6/10
핵심 요약

라틴 문자 기반 사회에서 발전한 활판 인쇄술과 컴퓨터 기술의 한계로 인해, 아랍어의 복잡한 필기체 연결과 우측-좌측 쓰기 방식이 디지털 환경에서 제대로 구현되지 못하는 문제를 다룹니다. 이는 단순히 기술적 결함을 넘어, 비영어권 언어가 글로벌 시스템에 맞춰져야 했던 구조적 한계를 보여주므로 다국어 처리 및 글로벌 AI 개발자에게 중요한 시사점을 던집니다.

번역된 본문

컴퓨터에서 아랍 문자를 표현하는 것은 처음부터 문제투성이였습니다. 이 시리즈에서는 이에 대한 몇 가지 해결책을 살펴볼 것입니다. 우선, 디지털 세계에서 아랍 문자가 왜 이렇게 문제가 되는지 이해할 필요가 있습니다. 아랍어는 1500년 전에 개발되었으며, 암각 문자와 펜과 잉크를 통해 처음 표현되었습니다. 펜과 잉크는 사실상 표현의 표준이 되었고 아랍 문자는 이 매체에서 발전했습니다. 아랍 문자의 대부분의 기술적인 측면은 꾸란을 가능한 한 정확하게 기록할 필요성에 대응하여 개발되었습니다. 예를 들어, 길게 발음하는 것, 동화 현상, 발음을 멈춰야 하는 지점과 계속 읽어야 하는 지점 등이 있습니다.

인쇄 시대의 아랍어 인쇄 시대로 넘어가 봅시다. 석판 인쇄는 아랍 문자와 잘 어울렸습니다. 왜냐하면 이 기술은 전체 페이지를 하나의 도장처럼 준비한 다음, 원하는 만큼 종이에 찍어내는 원리에 기반했기 때문입니다. 하지만 활판 인쇄(Movable type)라는 기술을 사용하면서 인쇄가 훨씬 더 대중화되었습니다. 이 기술은 페이지를 하나의 도장이 아니라 행으로 배열된 수백 또는 수천 개의 작은 도장들로 구성했으며, 각 도장은 하나의 문자나 발음 기호를 나타냈습니다. 각 도장을 옮길 수 있고 모든 페이지마다 재사용할 수 있어 고유한 도장을 만들 필요가 없었기 때문에 '옮길 수 있는'이라고 불렸습니다. 라틴 문자는 이러한 종류의 인쇄에 아주 쉽게 적용되었습니다. 아랍어 활판 인쇄도 비슷하게 개별 문자 단위로 개발되었습니다. 하지만 아랍어는 개별 문자 단위가 아니라 문자 묶음(letter block) 단위로 작성됩니다. 예를 들어, 'المعروف'라는 단어는 5개의 문자가 아니라 4개의 문자 묶음으로 작성됩니다. 첫 번째는 'ا', 두 번째는 'لمعر', 세 번째는 'و', 네 번째는 'ف'입니다. 특히 두 번째 문자 묶음은 특수한 합자(ligatures)를 알고 있는데, 'mīm' 위에 'lām'이 있고, 'mīm'은 오른쪽에 작은 체크 표시로 나타납니다. 1890년경 Brill 출판사의 초기 인쇄물에서 확대한 발췌 부분을 보면 각 문자에 자체적인 활자(도장)가 있음을 알 수 있습니다. 특히 브릴은 'mīm'과 'ḥa'에 대한 일반적인 합자를 실제로 하나의 활자로 수용했습니다. 그럼에도 불구하고 여기서 취해진 접근 방식은 매우 유연성이 떨어집니다. 더욱이, 넓게 벌어진 공간은 고르지 못한 독서 경험을 제공하며, 때로는 단어가 어디서 끝나는지 구별하기 어렵게 만듭니다. 후자의 문제는 해결되었습니다. 우리는 최근 인쇄된 작품들에서 문자 사이의 공간을 보지 못합니다. 하지만 유연성 부족 문제는 지속되었고, 20세기에 대부분의 출판사가 합자를 덜 사용하면서 논쟁의 여지 없이 악화되었습니다.

컴퓨터에서의 아랍어 요약하자면, 이 일련의 게시물 전체에서 논의하겠지만, 활판 인쇄의 이러한 결함은 아랍어의 디지털 표현으로 그대로 이어졌습니다. 이 문제가 지속되는 이유 중 하나는 인쇄기와 컴퓨터 같은 기술들이 라틴 문자 기반 사회에서 개발되었고, 따라서 그 사용 사례를 당연한 것으로 받아들였기 때문입니다. 매우 다양하고 복잡한 문자를 가진 세계의 나머지 지역은 이러한 규칙에 굴복해야 했습니다. 아랍어가 세계 6위의 언어임에도 불구하고 이러한 문제를 근본적으로 해결할 경제적 인센티브가 거의 없었던 것으로 보입니다. 사실, 디지털 환경은 이 문제를 더욱 악화시켰습니다. 텍스트를 문자로 나누는 활판 인쇄 철학을 채택함으로써, 컴퓨터는 아랍어를 연결된 문자로 표현하는 데 있어 상당한 문제를 겪었습니다. 다음 예를 보십시오: 누군가 불쌍한 영혼이 "what doesn't kill you makes you stronger(당신을 죽이지 못하는 것은 당신을 더 강하게 만든다)"를 아랍어로 번역하고, 컴퓨터로 타이핑한 다음 인쇄하여 타투이스트에게 주고 타투를 새기게 했습니다. 문자들이 'ما لا يقتلك يجعلك أقوى'처럼 연결되어 보여야 할 곳에 모든 문자가 개별적으로 표현되어 있습니다. 게다가 끔찍할 정도로 단순한 글꼴이기도 합니다. (올바르게 된 아랍어 타투입니다) 또한, 컴퓨터는 일부 문화권이 왼쪽에서 오른쪽이 아닌 오른쪽에서 왼쪽으로 쓴다는 사실을 파악하는 데 끔찍할 정도로 어려움을 겪었습니다. 따라서 타이핑할 때 문자가 정확히 반대 순서로 나올 수 있습니다. 여기 시험 예시가 있습니다.

원문 보기
원문 보기 (영어)
Representing Arabic script on a computer has been problematic from the very beginning. In this series, we shall explore some of the solutions offered. First, we need to understand what is so problematic about Arabic script in a digital world. Arabic was developed a millennium and a half ago, and found its first expression through rock inscription and pen and ink. Pen and ink became the de facto mode of expression and the Arabic script developed in this medium. Most of the more technical aspects of Arabic script were developed as a response to a need to record the Koran as precisely as possible, e.g. elongations, assimilations, points at which to stop and at which to keep reciting. Printing Arabic Fast forward to the printing era. Lithography went along well with Arabic script because its technology was based on the principle on preparing an entire page as though a stamp, and then stamping it on as many pieces of paper as one wished. But printing became much more popular using the technology called movable type . This technology constructed a page not as one stamp, but as hundreds or thousands of tiny stamps, arranged in rows, each stamp representing one letter or reading sign. It is called movable because you can move around each stamp, and reuse the stamps for every page instead of having one unique stamp for each page. Latin script lends itself quite easily to this type of printing. Movable type for Arabic was developed similarly, on a per letter basis. Arabic, however, is not written on a per letter basis, but on a per letter block basis. For example, the word المعروف is written not with five letters but with four letter blocks, the first is ا, the second is لمعر, the third is و, and the fourth is ف. Especially the second letter block knows special ligatures, with the lām on top of the mīm , and the mīm presented as a tick to the right. Take this zoomed in excerpt from an early print from Brill Publishers, circa 1890: We see that each letter has its own stamp. Notably, Brill actually accommodated the normal ligature for mīm and ḥa , as one stamp. Nevertheless, the approach taken here makes it very inflexible. Moreover, the open spaces provide for an uneven reading experience, in which it is sometimes hard to distinguish where a word ends. The latter problem has been solved; we do not see the spaces in between the letters in more recent printed works. The inflexibility persisted and has arguably only become more aggravated in the 20th century with most publishers using less ligatures. Arabic on computers In short, as we will discuss throughout this series of posts, these flaws of movable type transferred to the digital representation of Arabic. One way to look at the persistence of this problem is that these technologies, printing press and computers, were developed in Latin script based societies and thus took that use case as a matter of fact. The rest of the world, with its many varying complicated scripts, had to bend to these rules. It seems that even though Arabic is the sixth language of the world, there was little economic incentive to fundamentally solve these issues. In fact, the digital environment aggravated the problem even further. Taking the movable type philosophy of dividing text into letters, computers have had a significant problem in representing Arabic as connected letters. Take the following example: Some poor soul decided to translate "what doesn't kill you makes you stronger" into Arabic, type it out on a computer, print it and give it to a tattoo artist to have it as a tattoo. Where letters should have been connected to look like ما لا يقتلك يجعلك أقوى, all letters are represented individually. In a horribly simple typeface as well, I might add. ( this is Arabic tattoos done right) Further, computers have had a terribly hard time figuring out that some cultures write from right to left instead of left to right. Thus, when typing, the letters could turn out in the exact opposite order. Here is an example from Baltimore-Washington Airport: Arabic fail at BWI security lane (wrong direction and letters not connected) pic.twitter.com/OU3BboEvTW — Pinboard (@Pinboard) April 11, 2017 Or this: Absolute gibberish. Very shoddy from @BarbicanCentre for this Arabic poster. Disjointed, unreadable, left to right. pic.twitter.com/Ldel7fyUxt — Joseph Willits (@josephwillits) June 1, 2015 Lastly, computers have seemingly trouble with encoding Arabic. Encoding means that the visible representation on the screen has a string of 0s and 1s that a computer can actually store, that this string is stable so that a visible representation can be taken and repurposed somewhere else and indeed be the same. Examples will make this clear. In the image above, we can see that if we select the word hādhihi , copy it, and then paste it into the search function, we get a completely different string of characters. In the image above, we typed كشف in the search bar, and even though the third word of the document is exactly كشف, the PDF viewer cannot find anything of the like. In the image above, most things go right. We search for سن and get everything from the instances of the word sinn to sunna , to ḥasan , to isnād . In the search results, however, the PDF viewer wants to highlight the search string by making it bold. This breaks the connection with the previous and following letters, making for an awkward view. Working with Arabic invariably makes you come in contact with these problems at some point. Unicode Solutions to such problems have been proposed, but not widely or accurately implemented. Some problems were proposed to be solved with what is called unicode . This has become the standard for digital text, much like what movable type did for printing. Unicode is is simply a table which all companies agree to use, in which every character of every language is assigned a unique number. This means that the digital text file is a series of these numbers, which every computer can render into readable characters as they wish. For example, different fonts give different shapes to the same letter. So, the font can be changed without any changes occurring to the digital text. In other words, abstract letter and actual representation are separated. Unicode calls this difference a difference between characters and glyphs . The most notable example is that the characters for Chinese, Japanese, and Korean (CJK) have been unified. Only by showing the text in a font designed for Chinese will the text look like a Chinese text. A similar approach could have been implemented for Arabic too, for example by distinguishing rasm from diacritics and vocalization. This would not only solve problems from the printing era, but go over and beyond it by providing an ultra flexible approach to the Arabic script, one that does more justice to the writing practice and hence its philosophy. Instead, Unicode has encoded Arabic letters separatedly. Rather than considering a ta marbūṭa as a ha with two dots, it sees it as a separate letter. Searching for كثيره and كثيرة will in most cases yield different results, even though that it arguably should not do that. A similar issue occurs with vocalization and other super or sub signs. One can type أ, which is alif+hamza, with unicode number U+0623, but one can also type أ, which is an alif and a high hamza, represented by unicode numbers U+0627 and U+0654. Clearly, they are the same. They are both alif-hamza . But because they are encoded in different ways, they are usually not picked up as identical by computers. (In most cases, texts are written with the alif+hamza character.) Not even the simple idea of CJK has been implemented. What I mean is that a kāf should be a kāf , encoded as one and the same letter, but which may find different graphical expressions in different variants of writing, such as Arabic and Persian. But no, we find an entry for an "Arabic kāf " https://codepoints.net/U+0643, and an entry for a "Persian