메뉴
BL
TechCrunch AI 27일 전

클라우드플레어, AI 기업의 콘텐츠 무단 수집 제한 예정

IMP
8/10
핵심 요약

클라우드플레어가 2026년 9월부터 검색과 AI 학습을 동시에 수행하는 크롤러의 접근을 기본적으로 차단하여 퍼블리셔의 지식재산권을 보호합니다. 또한 AI가 콘텐츠를 활용해 가치를 창출할 때 퍼블리셔가 수익을 분배받을 수 있는 '페이 퍼 유즈(Pay Per Use)' 모델을 도입했습니다. 이는 웹 생태계의 지속 가능성을 보장하고 AI 기업들에게 콘텐츠 사용에 대한 정당한 대가를 지불하도록 유도하는 중요한 변화입니다.

번역된 본문

클라우드플레어(Cloudflare)가 구글 서치(Google Search)와 같은 전통적인 검색용 웹 크롤러와 AI 에이전트 및 학습에 사용되는 크롤러를 분리하도록 AI 업계에 새로운 최후통첩을 발표했습니다. 이 회사는 수요일에 발표한 내용을 통해 2026년 9월 15일부터 광고가 포함된 페이지에서 '혼합용도(mixed-use)' 크롤러를 기본적으로 차단하겠다고 밝혔습니다. 즉, 사이트 소유자가 설정을 별도로 변경하지 않는 한, 검색, 에이전트 사용, 학습을 혼합하여 수행하는 크롤러는 해당 사이트를 크롤링할 수 없게 됩니다. 회사 측은 이러한 기본 설정 변경은 새로운 클라우드플레어 고객, 기존 고객이 설정하는 새 사이트 및 모든 기존 무료 고객에게 적용된다고 덧붙였습니다.

이러한 조치는 AI 모델 제공자들이 학습 목적이나 에이전트 서비스 구동을 위해 웹 콘텐츠에 접근하는 방식에 큰 영향을 미칠 수 있습니다. 클라우드플레어는 대부분의 웹사이트 소유자들이 자신의 콘텐츠가 검색은 물론 AI 서비스를 통해서도 발견되기를 원하지만, 자신의 지적재산이 무료로 무단 사용되는 것에 대해서는 보호받기를 원한다고 지적했습니다. 클라우드플레어는 특히 '세계 최대의 검색엔진'(명백히 구글을 지칭)이 고객이 AI 학습 없이 검색 노출을 유지하기 어렵게 만듦으로써 다른 AI 기업들보다 약 '2배 더 많은 정보'에 접근할 수 있다고 꼬집었습니다.

구글은 과거에 이러한 일반화된 비판에 반박하며, 자사가 '구글 확장(Google Extended)'이라는 봇을 제공하여 사이트 소유자가 제미나이 앱(Gemini Apps) 및 버텍스 API(Vertex API)와 같은 AI 제품 및 서비스 학습에 콘텐츠가 사용되지 않도록 옵트아웃(opt-out)할 수 있도록 한다고 밝힌 바 있습니다. 이 설정은 구글 서치에 사이트가 노출되는 것에는 영향을 주지 않습니다. 하지만 이 거대 기술 기업의 대표적인 크롤러인 '구글봇(Googlebot)'은 AI 오버뷰(AI Overviews) 및 AI 모드(AI Mode)와 같은 AI 기능을 포함한 검색을 위해 웹을 크롤링합니다.

클라우드플레어의 공동 창립자이자 CEO인 매튜 프린스(Matthew Prince)는 최근 봇(bot) 트래픽이 온라인에서 처음으로 인간 트래픽을 추월한 이정표를 언급하며 다음과 같이 말했습니다. "이제 인터넷 트래픽의 대부분이 인간이 아닌 봇이 되었기 때문에, 지속 가능한 생태계가 구축될 수 있도록 더 강력하고 빠르게 행동해야 합니다." 이러한 변화는 내년이 되어야 일어날 것으로 예상되었던 바로 그 지점이었습니다.

프린스는 "클라우드플레어의 새로운 도구와 파트너십은 웹사이트 소유자에게 향상된 가시성과 상업적 기회를 제공하며, 명확하고 투명한 의도를 가진 봇을 운영하는 AI 회사들에게도 이익이 됩니다. 우리가 제안한 기본 설정 변경이 혼합용도 크롤러로 하여금 검색과 에이전트 사용 및 학습을 분리하도록 장려하기를 바랍니다."라고 덧붙였습니다.

클라우드플레어는 사용자가 자체 AI 시스템을 구축할 수 있도록 다양한 제품을 제공하면서도, 동시에 AI 시대에 퍼블리셔가 자신의 콘텐츠를 더 잘 통제할 수 있도록 돕는 여러 도구도 출시해 왔습니다. 최근 몇 년간 클라우드플레어는 AI 봇에 대한 대항을 위한 도구를 출시했는데, 여기에는 AI 봇이 스크래핑할 때 웹사이트가 비용을 청구할 수 있게 해주는 '페이 퍼 크롤(Pay Per Crawl)'이라는 마켓플레이스도 포함됩니다. 회사에 따르면 후자는 이제 '페이 퍼 유즈(Pay Per Use)'로 발전하고 있으며, 이를 통해 퍼블리셔는 콘텐츠가 단순히 조회(fetched)될 때가 아니라 실제로 가치를 창출할 때 AI 기업으로부터 비용을 청구할 수 있게 됩니다.

클라우드플레어의 데이터에 따르면 AI 크롤러의 크롤링 트래픽 중 50% 이상이 변경되지 않은 페이지를 반복해서 가져오는 데 소비되는 것으로 나타남에 따라, 이러한 변화는 AI 모델 제공자를 위해 퍼블리셔의 대역폭과 컴퓨팅 리소스를 절약하는 데에도 도움이 될 수 있습니다. 이를 실행에 옮기기 위해 클라우드플레어는 초기적으로 세라믹닷에이아이(Ceramic.ai)와 유닷컴(You.com) 두 파트너와 협력하고 있습니다. 퍼블리셔가 옵트인(opt in)하면, 자신의 콘텐츠가 세라믹의 AI 검색 결과에 나타나거나 유닷컴이 자신의 프리미엄 콘텐츠에 액세스할 때 비용을 지급받게 됩니다. 클라우드플레어는 다른 AI 기업들도 자사의 작동 방식에 맞게 이 모델을 커스터마이징할 수 있다고 밝혔습니다.

주제: AI, AI 학습(training), 클라우드플레어, 구글, 미디어 및 엔터테인먼트, 퍼블리셔 기사 내 링크를 통해 구매하시면 소정의 수수료를 받을 수 있습니다. 이는 본 매체의 편집 독립성에 영향을 미치지 않습니다. 사라 페레즈(Sarah Perez), 소비자 뉴스 에디터. 사라는 2011년 8월부터 TechCrunch의 기자로 일해왔습니다. 그녀는 이전에 ReadWriteWeb에서 3년 이상 일한 후 이 회사에 합류했습니다. 기자로 일하기 전에 사라는...

원문 보기
원문 보기 (영어)
Cloudflare has just issued the AI industry a new deadline to separate the web crawlers used for traditional search purposes, like Google Search, from those used for AI agents and training. Starting on September 15, 2026, Cloudflare's default settings will block "mixed-use" crawlers from any pages that host ads, the company announced on Wednesday. That means that the crawlers that blend search, agent use, and training will be blocked from crawling these sites by default, unless the site owner adjusts the settings otherwise. These changes to the defaults will apply to new Cloudflare customers, new sites set up by existing customers, and all existing free customers, the company says. The move could impact how AI model providers are able to access web content for training purposes and to help power their agentic services. Cloudflare points out that most website owners want their content to be discoverable via search and often through AI services as well, but they want protections against having their intellectual property given away for free. Cloudflare specifically calls out the "world's largest search engine" (clearly a Google reference!) as having access to about "2x more information" than other AI companies because the search giant makes it difficult for customers to remain discoverable without being used for AI. Google has pushed back against this generalization in the past, noting that it provides a bot called Google Extended that lets site owners opt out of having their content used for training and AI products and services like Gemini Apps and Vertex API. Its use doesn't impact a site's inclusion in Google Search. However, the tech giant's flagship Googlebot crawls for Search, including AI features like AI Overviews and AI Mode. "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge," said Cloudflare co-founder and CEO Matthew Prince in his announcement of the news, referring to the recent milestone where bots surpassed human traffic online for the first time. That shift was not expected to occur until next year. "Cloudflare's new tools and partnerships give website owners increased visibility and commercial opportunities and benefit AI companies that have bots with clear and transparent intent. We hope that our proposed default changes encourage mixed-use crawlers to separate out search from agent use and training," Prince said. While Cloudflare offers a number of products to help users launch their own AI systems , the company has also released a range of tools to give publishers more control over their content in the AI era. In recent years, Cloudflare launched tools to combat AI bots , including a marketplace that lets websites charge AI bots for scraping , dubbed Pay Per Crawl. The latter is now also evolving into "Pay Per Use," the company said, which will allow publishers to charge AI companies when their content creates value, not just when it's fetched. The change could also help conserve publishers' bandwidth and compute resources for AI model providers, as Cloudflare's data suggested that over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages. To put this into action, Cloudflare is initially working with two partners, Ceramic.ai and You.com. When a publisher opts in, they're paid when their content appears in Ceramic's AI search results or when You.com accesses a piece of their premium content. Other AI companies can customize this model for how they work, Cloudflare says. Topics AI , AI training , cloudflare , Google , Media & Entertainment , publishers When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Sarah Perez Consumer News Editor Sarah has worked as a reporter for TechCrunch since August 2011. She joined the company after having previously spent over three years at ReadWriteWeb. Prior to her work as a reporter, Sarah worked in I.T. across a number of industries, including banking, retail and software. You can contact or verify outreach from Sarah by emailing sarahp@techcrunch.com or via encrypted message at sarahperez.01 on Signal. View Bio November 4 Boston Last chance to save up to $190 on TechCrunch Founder Summit. Join 1,000+ founders and VCs at all stages for real-world scaling insights and connections that move the needle. Savings end June 26, 11:59 p.m. PT . REGISTER NOW Most Popular Flipper Device's new Busy Bar is a customizable display for productivity Ivan Mehta Ford rehires ‘gray beard’ engineers after AI falls short Anthony Ha Govee's smart nugget ice maker makes every iced drink feel like a luxury Aisha Malik Asian AI startups launch Mythos-like models as Anthropic's export ban drags on Kate Park FTC gives Musk the OK to acquire SpaceX alumni startup Mesh Marina Temkin Corgi, the buzzy Y Combinator-backed insurance tech startup, says it didn't steal an open source product Julie Bort Trump administration proposes axing brake-pedal requirement for AVs in a boost for Tesla Sean O'Kane