메뉴
HN
Hacker News • 24일 전

퍼플렉시티 인용의 3분의 1은 근거 숫자를 포함하지 않는다

IMP
7/10
핵심 요약

Perplexity 검색 모델에 310개의 사실 질문을 던져 인용 1,826건을 검증한 결과, 34.7%가 열리지 않거나 문장의 숫자를 전혀 포함하지 않는 페이지를 가리켰다. 주요 원인은 죽은 링크(1.3%에 불과)가 아니라 로그인·유료화·봇 차단으로 독자가 접근할 수 없는 페이지(16.1%)와 접근은 가능하지만 해당 내용이 없는 페이지였다. LLM 답변의 인용이 실제 검증 가능한 출처 보증으로 기능하지 못한다는 점에서 AI 검색 신뢰성 논란을 촉발할 연구다.

번역된 본문

HR-2026-09 · 2026년 9월 2일

퍼플렉시티(Perplexity) 인용의 3분의 1은 인용된 숫자를 담고 있지 않다

Perplexity의 검색 모델이 숫자를 포함한 문장에 붙인 인용 1,826건 중 34.7%가 열리지 않거나 해당 문장의 숫자를 단 하나도 포함하지 않는 페이지를 가리켰다. 인용이 아닌 주장(claim) 단위로 평가하면 872건 중 14.4%가 실패한다.

우리는 Perplexity의 두 검색 모델에 210개 기술 기업에 관한 310개의 사실 질문을 던지고, 인용된 모든 출처를 수집해 전부 가져온(fetch) 뒤, 해당 페이지가 인용된 내용을 실제로 담고 있는지 확인했다. 숫자 없이는 확인할 수 없는 것을 제외하고, 숫자가 포함된 문장에 붙은 1,826건의 인용 중 34.7%가 일반 독자가 열 수 없는 페이지를 가리키거나, 열어도 해당 문장의 숫자가 하나도 없는 페이지를 가리켰다. 모델은 총 2,511개의 인용 표시를 달았다.

위의 단위는 인용이지 주장이 아니다. 숫자를 포함한 872개 주장 중 3분의 2는 둘 이상의 표시를 가지고 있으며, 우리는 각 표시를 개별적으로 평가했다. 대신 주장 단위로, 즉 가리키는 페이지 중 하나라도 해당 숫자를 담고 있으면 통과로 계산하면 14.4%가 실패한다. 우리가 인용을 우선으로 내세우는 이유는, 표시 하나하나가 개별적인 출처 주장이기 때문이다. 즉 '이 문장은 이 URL에서 나왔다'는 것이다.

실패의 주된 원인은 죽은 링크가 아니다. 인용된 URL 중 죽은 링크는 1.3%에 불과했다. 두 큰 범주는 독자가 들어갈 수 없는 페이지와, 들어갈 수는 있지만 그 내용이 없는 페이지다.

무엇을 했나

열 가지 질문 템플릿을 사용했으며, 각각은 실제로 사람이 찾아볼 만한 사실이었다: 설립, 최신 투자 라운드, 직원 수, 진입 가격, 본사 위치, 매출, 공개된 보안 사고, 현任 CEO, 인수합병, 유료 플랜 가동률 SLA. 모든 기업에 한 번씩 질문했고, 그중 100개 기업에는 다른 템플릿으로 한 번 더 질문했다. temperature 0으로 310개 질문을 perplexity/sonar, perplexity/sonar-pro, 그리고 대조군으로 웹 플러그인을 단 GPT-4.1에 던졌다.

두 Perplexity 모델 모두 주장을 본문에 [n]으로 표시하며, n은 반환하는 인용 배열의 인덱스다. 이 점이 감사(audit)를 가능하게 하는 부분이다. 답변 하단의 참고문헌 목록이 아니라, '이 문장은 이 URL에서 나왔다'는 구체적 주장이기 때문이다. 우리는 각 답변을 문장으로 나누고 표시마다 하나의 주장–인용 쌍을 만들었다. 두 모델 모두 자기 인용 목록의 범위를 넘어서는 표시를 내놓은 적은 없었다.

그다음 인용된 고유 URL을 전부 가져왔는데, sonar만 따로 2,915개였다. 각각을 죽은 링크(dead), 접근 제한(gated), 빈 페이지(empty), 접속 불가(unreachable), 정상(live)으로 분류했다. 실패한 URL은 두 번의 기회를 더 받았다. 더 긴 타임아웃, 그리고 특정 데이터센터 IP가 차단된 것만으로 막힌 것으로 기록되지 않도록 로테이션 프록시를 통한 재시도였다. 이 세 번째 시도는 192개의 URL을 복구했다. 이 분류는 언제나 페이지에게 유리한 방향으로만 작동한다.

핵심 검증에는 모델이 전혀 필요 없다. 각 주장에서 구체적 수치 — 금액, 백분율, 규모, 연도, 세 자리 이상의 숫자 나열 — 를 추출하고, 인용된 페이지의 보이는 텍스트가 그 중 하나라도 포함하는지 물었다. $185 million, $185M, 185000000가 모두 일치하도록 정규화했다. 숫자 하나만 있어도 통과다. 연도만 있어도 통과다. 따라서 34.7%는 사실상 하한선이다. 실패한 쌍은 모두 페이지에 인용한 문장의 숫자가 단 하나도 없는 경우다.

열리지 않는 인용

perplexity/sonar의 고유 인용 URL 2,915개 기준:

  • 정상적으로 읽을 수 있음: 78.7%
  • 로그인·유료벽·403·봇 차단 뒤에 있음: 16.1%
  • 읽을 수 없는 클라이언트 렌더링 셸: 2.5%
  • 죽은 링크(404, 410, DNS 실패, soft 404): 1.3%
  • 세 차례 시도 후에도 접속 불가: 1.4%

여섯 건 중 한 건의 인용이 접근 제한되어 있다. 이것은 출처의 잘못이 아니다. PitchBook, ZoomInfo, Crunchbase, Reuters는 유료화할 권리가 있다. 하지만 인용의 잘못이다. 독자가 열 수 없는 각주는 검증할 방법이 없는 출처 주장이며, 이는 인용이 존재해서 막으려는 바로 그 상황이다.

답변 수준으로 집계하면 sonar의 310개 답변 중 84.2%가…

원문 보기
원문 보기 (영어)
HR-2026-09 &middot; 2 September 2026 A third of Perplexity&#x27;s citations don&#x27;t contain the number they&#x27;re cited for Of 1,826 citations Perplexity&#x27;s search models attached to a sentence stating a figure, 34.7% pointed at a page that either would not open or did not contain a single figure from that sentence; scored per claim rather than per citation, 14.4% of 872 claims fail. We asked Perplexity&rsquo;s two search models 310 factual questions about 210 technology companies, collected every source they cited, fetched all of them, and checked whether the page said the thing it was cited for. Of the 1,826 citations attached to a sentence stating a figure — the ones checkable without a second opinion — 34.7% pointed at a page that would not open to an ordinary reader, or opened and contained none of the numbers in the sentence they were attached to. The models placed 2,511 citation markers in all. The unit above is the citation, not the claim. Two thirds of the 872 claims carrying a figure have more than one marker on them, and we score each marker separately. Score instead per claim, counting a claim as passing when any one of the pages it points at carries one of its figures, and 14.4% fail. We lead with the citation because a marker is an individual claim of provenance: this sentence came from that URL. The failure is not mainly dead links. Only 1.3% of cited URLs were dead. The two large categories are pages a reader cannot get into, and pages a reader can get into that do not say it. What we did Ten question templates, each a fact somebody would actually look up: founding, latest funding round, headcount, entry price, headquarters, revenue, disclosed breaches, current CEO, acquisitions, paid-tier uptime SLA. Every company got one; 100 of them got a second on a different template. 310 questions, put at temperature 0 to perplexity/sonar , perplexity/sonar-pro and, as a control, GPT-4.1 with a web plugin. Both Perplexity models mark their claims inline as [n] , and n indexes the citation array they return. That is the part that makes an audit possible: it is not a bibliography at the bottom of the answer, it is a specific assertion that this sentence came from that URL. We split each answer into sentences and produced one claim–citation pair per marker. Neither model ever emitted a marker pointing past the end of its own citation list. Then we fetched every unique cited URL — 2,915 of them for sonar alone — and classified each as dead, gated, empty, unreachable or live. Anything that failed got two more chances: a longer timeout, then a retry through a rotating proxy so that no page was recorded as blocked merely because one datacentre address was unwelcome. That third pass rescued 192 URLs. The classification can only ever move in a page&rsquo;s favour. The headline check needs no model at all. From each claim we pulled its specifics — money amounts, percentages, magnitudes, years, any run of three or more digits — and asked whether the cited page&rsquo;s visible text contains at least one of them, normalising so that $185 million , $185M and 185000000 all match. One figure is enough to pass. A bare year is enough to pass. The 34.7% is therefore a floor: every failing pair is one where the page contains not a single number from the sentence that cited it. The citations that do not open Across perplexity/sonar &rsquo;s 2,915 unique cited URLs: Class Share Live and readable 78.7% Behind a login, paywall, 403 or bot wall 16.1% Client-rendered shell we could not read 2.5% Dead (404, 410, DNS failure, soft 404) 1.3% Still unreachable after three passes 1.4% One citation in six is gated. That is not a fault of the source — PitchBook, ZoomInfo, Crunchbase and Reuters are entitled to charge — but it is a fault of the citation. A footnote a reader cannot open is a claim of provenance with no way to test it, which is the condition a citation exists to prevent. Aggregated to the answer, 84.2% of sonar &rsquo;s 310 answers cited at least one URL an ordinary reader could not open, and 10.6% cited at least one that was outright dead. The dead ones are worth naming, because about half of them are the same kind of page — 20 of sonar &rsquo;s 38 — and their URLs give them away. komo.ai/directory/<company>-offices . temperstack.com/plans/<company> . devhelm.io/sla/<company> . apollo.io/where-is/<company> . portersfiveforce.com/blogs/brief-history/<company> , and the identical path on matrixbcg.com and canvasbusinessmodel.com . These are pages minted per company per question type, published at scale to catch exactly the query we asked, and taken down as cheaply as they went up. Three we re-fetched on the day of writing. Asked where Elastic is headquartered, sonar cited komo.ai/directory/elastic-offices : 404. Asked for Reddit&rsquo;s head office, both models cited apollo.io/where-is/reddit : 410 Gone. Asked for Discord&rsquo;s cheapest paid plan, sonar-pro cited temperstack.com/plans/discord : 404. The citations that open and do not say it Of the pairs whose page did open and was readable, 16.1% contained none of the claim&rsquo;s own figures. The cleanest example is a price. Asked for the entry price of Vercel&rsquo;s cheapest paid plan, sonar answered that &ldquo;the free Hobby plan is $0/month, so the first paid tier starts at $20/month&rdquo;, and cited vercel.com/docs/plans . We fetched that page at write time. It returns HTTP 200, it names the plans, and the strings $20 , $20/month and 20/month do not appear anywhere in it. The number is probably right. The citation is not evidence for it. The second pattern is more revealing, because it repeats across companies. Asked for headquarters, both models produce a street address and attribute it to the company&rsquo;s Wikipedia article: Claim Cited page Address on that page? Docker at 3790 El Camino Real #1052, Palo Alto, CA 94306 en.wikipedia.org/wiki/Docker,_Inc. No Rippling at 430 California Street, San Francisco, CA 94104 en.wikipedia.org/wiki/Rippling_(company) No Substack at 111 Sutter Street, San Francisco, CA 94104 en.wikipedia.org/wiki/Substack No SentinelOne at 444 Castro Street, Mountain View, CA 94041 en.wikipedia.org/wiki/SentinelOne No All four articles were fetched at write time and none contains the street number, the street name or the postal code attributed to it. Several of the answers say so themselves, in phrasing like &ldquo;multiple sources list&rdquo; or &ldquo;several business directories list&rdquo;, and then attach a marker to Wikipedia anyway. The claim and the citation were produced by the same process, and that process is not retrieval. A softer version of the same thing: sonar said GitLab&rsquo;s CEO is Bill Staples and that he took the role on 5 December 2024, citing GitLab&rsquo;s own executive team page . That page names Bill Staples. It does not carry the date. Half the sentence is sourced. Where it fails worst Pooling both Perplexity models, by question type, share of pairs whose cited page contained one of the claim&rsquo;s figures: Question Pairs Passed Who is the current CEO 235 44.3% Headquarters address 202 53.0% Entry price 171 62.6% Security incidents 205 66.8% Uptime SLA 145 69.0% Acquisitions 278 69.1% Latest funding round 103 70.9% Revenue or ARR 192 72.4% Founding 128 75.0% Headcount 167 82.0% The ordering is not random. It tracks how well a fact is written down in one canonical place. Headcount and founding year sit in structured fields on pages built to hold them. A CEO&rsquo;s start date and an office&rsquo;s street number are the kind of thing everyone repeats and nobody publishes, so the model reproduces the consensus and then points at a page that never carried it. The premium model is not better, and not worse Our pilot suggested that sonar-pro grounded its claims less well than sonar . At full scale that gap disappears. sonar passes on 65.9% of numeric pairs (95% interval 62.8–68.9), sonar-pro on 64.7% (61.5–67.7). The intervals overlap comfortabl