기계학습(Machine Learning) 기술로 딥페이크 음성을 합성하여 사람의 목소리를 완벽하게 모방하는 AI 음성 사기가 전례 없는 속도로 확산하고 있습니다. FBI에 따르면 2025년 한 해 동안 AI 기반 사기로 인한 미국 내 피해액이 최소 8억 9,300만 달러에 달하며, 특히 60세 이상 고령층이 주요 표적이 되었습니다. 정교한 공격 기술과 대중의 인식 부재 간의 격차가 커져 금융 기관과 사법 당국이 대응에 어려움을 겪고 있습니다.
번역된 본문
3초의 절도: AI 음성 사기가 모든 방어망을 뚫는 이유
2026년 7월 15일. 샤론 브라이트웰(Sharon Brightwell)은 전화선 너머로 자신의 딸이 우는 소리를 듣는 순간, 그녀가 세울 수 있었던 모든 방어망이 무너졌다. 모든 본능은 그 목소리가 딸 에이프릴(April)의 것이라고 말했다. 똑같은 음색, 고통받는 젊은 여성의 똑같은 불안정한 호흡이었다. 그 목소리는 운전 중 문자 메시지를 보내다 임산부를 치었으며, 휴대전화는 경찰에 압수되었다고 말했다. 이어 자칭 에이프릴의 변호사라고 밝힌 남자가 전화를 받아 현금으로 1만 5천 달러의 보석금이 필요하다고 설명했다. 그는 딸의 신용도가 훼손될 수 있으니 은행에 돈의 용도를 말하지 말 것을 경고했다. 1시간도 채 지나지 않아 플로리다주 도버(Dover)에 거주하는 이 은퇴자는 돈을 인출하여 법원 관계자라고 믿었던 택배 기사에게 건넸다. 그날 오전 내내 일을 했으며 교통사고가 난 적이 전혀 없는 진짜 에이프릴에게 연락이 닿았을 때서야, 그녀는 비로소 딸이 그 전화를 걸지 않았다는 사실을 깨달았다. 인간이 건 것이 아니었다. 그 울음소리는 아주 짧은 오디오 조각으로부터 합성된 것이었으며, 그녀가 구출한다고 생각했던 딸은 타인의 기계 속에 숫자 패턴으로만 존재했을 뿐이다.
2025년 여름 미국 지역 뉴스를 통해 보도된 브라이트웰의 피해 사례는 이제 미국에서 가장 흔한 범죄 중 하나가 되었다. 그러면서도 기술적으로는 가장 진보한 범죄이기도 하다. 기계 학습(Machine Learning)의 최전선 기술이 요구되는 사기가, 부엌에 있는 평범한 할머니를 상대로 아무런 비용 없이 대규모로 자행될 수 있다는 두 가지 사실의 충돌은, 법 집행 기관, 은행, 통신회사 및 규제 당국이 2년 동안 통제하지 못한 이 문제의 핵심적인 특징이다. 문제는 더 이상 이 기술이 작동하느냐가 아니다. 끔찍할 정도로 완벽하게 작동한다. 문제는 공격의 정교함과 대상의 인식 사이의 격차가 몇 달이 아니라 몇 년 단위로 측정될 때, 어떤 의미 있는 보호 조치가 필요한지에 있다.
26년된 장부의 새로운 항목
2026년 4월, FBI 인터넷 범죄 신고 센터(IC3)는 지난해의 온라인 범죄에 대한 연례 보고서를 발표했고, 보고서 26년 역사상 처음으로 인공지능(AI) 기반 사기를 별도의 범주로 구분하여 기록했다. 수치는 충격적이었다. FBI는 AI와 관련된 신고를 2만 2천 건 이상 접수했으며, 집계된 피해액은 8억 9,300만 달러를 넘는 것으로 나타났다. 이 중 3억 5,200만 달러의 피해액은 60세 이상 피해자에게 귀속된 것으로, 고령층이 AI 기반 금융 범죄에서 가장 많이 표적이 되는 인구 집단이 되었음을 보여주었다. AI 관련 수치는 훨씬 더 큰 총액의 일부에 불과했다. 미국 전역의 사이버 범죄 피해액은 1년 만에 26% 증가한 209억 달러에 달했으며, 60세 이상 미국인이 그중 77억 달러를 차지해 전년 대비 약 60%가 급증했다. FBI는 이 수치조차도 실제 문제를 과소 평가한 것이라고 솔직하게 인정했다. 보고서의 AI 귀속 수치는 피해자가 인지하고 신고한 것만을 반영하며, 복제된 음성 통화 피해자 대부분은 기계가 개입했다는 사실조차 알지 못한다. 그들은 샤론 브라이트웰이 처음에 믿었던 것처럼 자신의 자녀와 직접 통화했다고 생각한다. 따라서 8억 9,300만 달러라는 금액은 한계치가 아닌 최소한의 기준선으로 보는 것이 타당하다. 본질적으로 피해자들에게 보이지 않도록 설계된 범주의 겉으로 드러난 부분일 뿐이다. FBI가 이 범주를 신설해야만 했다는 사실 자체가 하나의 신호다. 범죄 통계는 보수적인 도구이며, 수사 기관은 일시적인 유행 때문에 26년 된 보고 체계를 새로 만들지 않는다. 장부의 새로운 기록은 불과 3년 전만 해도 소비자용으로 거의 존재하지 않았던 도구가 이제 주요 절도 수단이 되었다는 인정이다. 국제적으로 상황은 더욱 심각하고 악화되고 있다. 2026년 3월, 인터폴(INTERPOL)은 제2판 글로벌 금융 사기 위협 평가 보고서를 발표하여 전 세계 금융 사기로 인한 피해액을 4,420억 달러로 추정했다.
The Three-Second Theft: Why AI Voice Fraud Outruns Every Defence July 15, 2026 Sharon Brightwell heard her daughter crying down the line, and that was the end of any defence she might have mounted. The voice belonged to April, or so every instinct insisted: the same timbre, the same broken rhythm of a young woman in distress. The voice said she had been texting while driving, that she had hit a pregnant woman, that her phone had been seized by police. A man then took over the call, identifying himself as April's attorney, and explained that bail would cost fifteen thousand dollars in cash. He warned Brightwell not to tell the bank what the money was for, because it might damage her daughter's credit. Within the hour, the retiree from Dover, Florida had withdrawn the money and handed it to a courier she believed was connected to the courts. Only when she reached the real April, who had spent the morning at work and never been near a car accident, did she understand that her daughter had not made the call. No human had. The crying had been synthesised from a fragment of audio, and the daughter she thought she was rescuing existed only as a pattern of numbers in someone else's machine. Brightwell's loss, reported across American local news in the summer of 2025, is now one of the most ordinary crimes in the United States. It is also one of the most technically advanced. The collision of those two facts — that a fraud requiring the absolute frontier of machine learning can be perpetrated against an ordinary grandmother in her kitchen, at scale, for the price of nothing — is the defining feature of a problem that law enforcement, banks, telecoms companies and regulators have spent two years failing to contain. The question is no longer whether the technology works. It works appallingly well. The question is what meaningful protection requires when the gap between the sophistication of the attack and the awareness of the target is measured not in months but in years. A New Line in a Twenty-Six-Year Ledger In April 2026, the FBI's Internet Crime Complaint Center published its annual report on the previous year's online crime, and for the first time in the report's twenty-six-year history it broke out artificial-intelligence-enabled fraud as a distinct category. The numbers were stark. The bureau logged more than 22,000 complaints with an AI nexus and adjusted losses exceeding 893 million dollars. Of that sum, the report attributed 352 million dollars in losses to victims aged sixty and over, making older adults the single most heavily targeted demographic in AI-enabled financial crime. The AI figure sat inside a far larger total: cybercrime losses across the United States rose 26 per cent in a single year to 20.9 billion dollars, with Americans aged sixty and older accounting for 7.7 billion of that — a roughly 60 per cent jump on the previous year. The FBI was candid that even these figures understate the problem. AI attribution in the report reflects only what victims recognised and reported, and most victims of a cloned-voice call never learn that a machine was involved at all. They believe, as Sharon Brightwell initially believed, that they spoke to their own child. The 893 million dollars is therefore best read as a floor, not a ceiling — the visible portion of a category that is, by its nature, designed to remain invisible to the people it harms. That the FBI felt compelled to create the category at all is itself a signal. Crime statistics are conservative instruments; agencies do not redraw twenty-six-year-old reporting taxonomies for a passing fashion. The new line in the ledger is an admission that a tool which barely existed in consumer form three years ago has become a mainstream instrument of theft. Internationally, the picture is larger and worsening. In March 2026, INTERPOL published the second edition of its Global Financial Fraud Threat Assessment, estimating worldwide losses to financial fraud at 442 billion dollars in 2025 — a sum comparable to the entire annual economic output of Denmark. The organisation rated the threat trajectory as escalating and described what it called the “industrialisation of fraud”: the migration of scamming from opportunistic individuals to organised, transnational operations that intersect with human trafficking and cybercrime. Crucially, INTERPOL found that AI-enhanced fraud is roughly four and a half times more profitable than its traditional equivalent, and that so-called agentic AI systems can now autonomously plan and execute entire fraud campaigns, from reconnaissance through to the ransom demand. The economics, in other words, have inverted. For the first time, deception at industrial scale costs almost nothing to manufacture and returns a fortune. Three Seconds Is All It Takes The technical capability at the centre of the grandparent scam is brutally simple to describe. A modern AI voice-cloning system requires as little as three seconds of audio to produce a synthetic voice that is, for practical purposes, indistinguishable from the original. Three seconds is the length of a voicemail greeting, a snatch of a podcast, the audio under a birthday video posted to a public Instagram account. The raw material is not stolen from a secure database; it is volunteered, every day, by the ordinary act of living a recorded life. A grandchild who appears in a single TikTok clip has supplied everything a fraudster needs to manufacture their own kidnapping. What makes the threat acute is not merely that the cloning works but that the tools to do it are cheap, abundant and almost entirely unpoliced. In March 2025, Consumer Reports assessed the voice-cloning products of six companies — Descript, ElevenLabs, Lovo, PlayHT, Resemble AI and Speechify — and concluded that a majority lacked any meaningful safeguard against fraud or misuse. Four of the products, the organisation found, required only that a user tick a box affirming they had the legal right to clone the voice in question. None of those four employed any technical mechanism to confirm that the speaker had actually consented, or to restrict cloning to the user's own voice. Four of the six companies required nothing more than a name or an email address to open an account. The investigation's blunt conclusion, amplified by NBC News and The Register, was that the industry had built a tool capable of impersonating anyone and then placed it behind a self-attestation checkbox. ElevenLabs, one of the most prominent providers, points to a multi-layered safety programme: a prohibited-use policy that bans impersonation, a public AI speech classifier that can identify audio likely to have originated from its system, traceability that links generated content back to the account that produced it, and “no-go voices” safeguards that block the cloning of certain protected figures around election cycles. These are not trivial measures, and they are more than several competitors offer. But they share a structural weakness: almost all of them operate after the fact. They help investigators establish provenance once a fraud has already occurred and a victim has already lost their savings. They do very little to prevent the three-second clone from being generated in the first place, because the thing that would prevent it — robust, mandatory verification that the person being cloned has consented — is precisely the friction that a competitive, fast-moving market is reluctant to impose on itself. When a safeguard costs a company conversions and protects only the customers of its rivals, the market will not supply it voluntarily. It has not. The Forensic Authority Who Went Blind If there is a single moment that captures why detection-based defences are failing, it arrived in a New York Times profile published in June 2026. Its subject was Hany Farid, the University of California, Berkeley professor who is, by broad consensus, the world's foremost authority on deepfake forensics. For