메뉴
HN
Hacker News • 10일 전

AI 학습 차단하면서 검색 노출은 유지하는 Cloudflare 신기능

IMP
7/10
핵심 요약

Cloudflare가 검색 인덱싱은 허용하면서 AI 학습은 거부할 수 있는 'Disallow AI Training' 설정을 발표했습니다. Apple, Google, Microsoft가 이 설정을 존중하거나 존중하기로 약속했습니다. 내년 초까지 AI 요약에 포함될 콘텐츠 양까지 세밀하게 제어할 수 있게 하는 것이 목표입니다.

번역된 본문

적절한 통제 장치가 없으면 웹사이트 소유자는 오랫동안 어려운 트레이드오프에 직면해 왔습니다. 콘텐츠를 AI 학습에 사용되게 하거나, 검색에서의 발견 가능성을 잃을 위험을 감수하거나 둘 중 하나를 선택해야 했죠. 이런 트레이드오프가 존재하는 이유는 인터넷에서 가장 큰 일부 기업들이 혼합用途 크롤러, 즉 검색과 AI 학습을 동시에 수행하는 단일 크롤러를 사용하기 때문입니다. 하나를 거부하면 다른 것도 거부하게 됩니다.

오늘 Cloudflare는 새로운 'Disallow AI Training(AI 학습 거부)' 설정을 발표합니다. 이를 통해 동일한 크롤러가 콘텐츠로 학습하는 것은 거부하면서 검색 인덱싱은 쉽게 유지할 수 있습니다. Apple, Google, Microsoft는 이 설정을 존중하거나 지정된 기한 내에 존중하겠다고 약속한 상태입니다.

혼합用途 크롤러는 학습 문제에서 어려운 부분이었습니다. 다음은 AI 요약입니다. 사이트 전체에 대한 예/아니요는 너무 무딘 도구입니다. 요약에 콘텐츠가 나타나는지 여부만큼이나 얼마나 많은 콘텐츠가 나타나는지도 중요하기 때문입니다. AI 요약 거부 옵트아웃은 이미 혼합用途 크롤러 운영사에게 요구하는 요건 중 하나입니다. 내년 초까지는 Cloudflare에서 한 번만 설정하면 각 운영사와 별도로 협상하지 않고도 요약에 포함될 콘텐츠의 양을 제어할 수 있도록 하는 것이 목표입니다.

요청만으로는 부족한 이유

대부분의 사이트 소유자는 발견되기를 원합니다. 인간, 에이전트, 그리고 (선의의) 봇에 의해서요. 하지만 개방형 인터넷의 상당 부분은 광고, 구독, 방문자와의 직접 관계로 자금을 조달하며, 이러한 모델은 누군가 실제로 사이트에 도착해야만 수익이 발생합니다.

거의 모든 사이트 소유자가 검색을 유익하다고 생각합니다. Cloudflare 사이트 중 1% 미만만이 검색 봇을 차단하도록 선택했습니다. 하지만 학습은 다른 이야기입니다. 17%의 사이트가 학습을 차단하는 메커니즘을 활성화했습니다. 바로 이 때문에 사이트 소유자에게 획일적인 'AI 차단'이 아니라 더 세밀한 통제가 필요하다고 판단했습니다.

robots.txt 지시어 하나만으로는 이 문제를 해결할 수 없습니다. 누구나 발행할 수 있지만, 누가 크롤링하는지 식별하거나 왜 크롤링하는지 판단하거나 이를 무시하는 크롤러를 막을 수는 없습니다. 하지만 네트워크는 해결할 수 있습니다. 우리는 선호 설정을 발행하고, 크롤러의 신원을 식별하고, 크롤링 목적을 분류하고, 이를 무시하는 크롤러를 차단한 뒤 각 운영사가 실제로 무엇을 하는지 Radar에 보고합니다.

하지만 차단은 크롤러를 제거할 뿐입니다. 크롤러의 행동 방식을 바꾸지는 않습니다. 더 나은 결과는 애초에 사이트 소유자가 선택하도록 강요하지 않는 운영사입니다. 그래서 7월부터 운영사들과 직접 대화해 왔습니다. 반응은 고무적입니다. 거의 모든 운영사가 사이트 소유자가 콘텐츠 사용 방식에 대한 통제권과 투명성, 그리고 자신의 선택이 존중받을 것이라는 확신을 가져야 한다는 데 동의했습니다.

이를 돕기 위해 'Accountable(책임 있는)'이라는 인증을 만들었습니다. Accountable 인증은 현재 제공되는 기능과 이를 제공하겠다는 구체적인 약속을 모두 인정합니다. 인증을 받으려면 봇 운영사가 다음 요건을 충족하거나 충족하겠다고 약속해야 합니다.

robots.txt 또는 유사한 표준을 통해 사이트 소유자가 AI 학습을 옵트아웃할 수 있는 메커니즘 운영사와 직접, 그리고 내년에는 Cloudflare를 통해(아래 섹션 참조) 사이트 소유자가 AI 요약을 옵트아웃할 수 있는 메커니즘 학습에 제공된 페이지를 URL 단위로 확인할 수 있는 가시성과, 검색에서 콘텐츠가 어떻게 나타났는지 보여주는 지표 AI 학습 옵트아웃이 전통적인 검색 결과에 영향을 주지 않는다는 보장

Apple, Google, Microsoft는 모두 Accountable 자격 요건을 충족함을 입증합니다. 각사는 현재 이용 가능한 기능과 아직 개발 중인 기능에 대한 기한이 정해진 약속을 결합했습니다. 각 회사 크롤러의 세부 사항은 아래에 공유합니다.

새로운 보안 설정 옵션

Cloudflare는 행동 기준으로 봇을 분류하며, 단일 봇이 둘 이상의 행동을 보일 수 있습니다. 통제 가능한 세 가지 행동은 다음과 같습니다:

검색(Search) - 검색 인덱스를 구축하기 위한 크롤링 학습(Training) - 모델을 학습하거나 파인튜닝하기 위한 크롤링 에이전트(Agent) - 인간을 대신해 페이지를 방문하는 사용자 지시형 에이전트(챗 페치 봇, 브라우저 사용 에이전트 등)

혼합用途 크롤러는 검색과 학습을 모두 수행하는 단일 크롤러입니다. 통제 장치가 없으면 그...

원문 보기
원문 보기 (영어)
Without proper controls, website owners have long faced a difficult tradeoff: allow your content to be used for AI training, or risk losing discoverability in search. That tradeoff exists because some of the largest organizations on the Internet use mixed-use crawlers: a single crawler serving both search and AI training. Refuse one, and you refuse the other. Today, Cloudflare is announcing a new Disallow AI Training setting that lets you easily stay indexed for search while refusing to let that same crawler train on your content. Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting. Mixed-use crawlers were the hard part of the training question. AI Summaries are next. A site-wide yes or no is too blunt: how much of your content appears in a summary matters as much as whether it appears at all. An opt-out for AI summaries is already one of the requirements we've set for mixed-use crawler operators. By early next year, our goal is to let you control how much of your content is included — set once on Cloudflare, rather than with each operator separately. Why asking isn’t enough Most site owners want to be found: by humans, agents, and (good) bots. But a significant portion of the open Internet is funded by advertising, subscriptions, or direct relationships with visitors, and those models only pay when someone actually arrives. Almost every site owner considers Search beneficial: less than 1% of Cloudflare sites choose to block Search bots. Training, however, is a different story: 17% of sites choose to enable some mechanism to block training. This is exactly why we decided site owners needed more granular controls, rather than a one-size-fits-all “Block AI.” A robots.txt directive alone cannot solve this problem. Anyone can publish one, but it cannot identify who is crawling, determine why they are crawling, or stop a crawler that ignores it. A network can solve it, however: we publish the preference, identify who is crawling, classify why they are crawling, and block the ones that ignore it – then report what each operator actually does on Radar . But blocking removes a crawler. It doesn't change how crawlers behave. The better outcome is operators that don't make you choose at all. So since July, we've been talking to them directly. The response has been encouraging: almost all agreed that site owners should have control and transparency into how their content is used, and reassurance that their choices will be respected. To help site owners understand that, we created a designation: Accountable. The Accountable designation recognizes both capabilities available today and concrete commitments to deliver them. To qualify, a bot operator must meet or commit to meeting the following requirements: A mechanism for site owners to opt out of AI training, through robots.txt or a similar standard. A mechanism for site owners to opt out of AI summaries set with the operator directly, and next year through Cloudflare (see section below for more detail). URL-level visibility into which pages were made available for training, along with metrics showing how content appeared in search. Assurance that opting out of AI training will not affect traditional search results. Apple, Google, and Microsoft all demonstrate that they meet the qualifications to be Accountable. Each combines capabilities available today with time-bound commitments for those still in development. The details of each of these companies’ crawlers are shared below. New security setting options Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior. Three behaviors are available as controls: Search - crawling to build a search index. Training - crawling to train or fine-tune a model. Agent - user-directed agents visiting a page on behalf of a human, such as chat fetch bots and browser-use agents. A mixed-use crawler is a single crawler doing both Search and Training. Without controls, that combination creates the tradeoff described above: site owners cannot refuse one use without refusing the other. To avoid blocking Accountable mixed-use crawlers — the ones that don't force that tradeoff on website owners — we are introducing a new setting: Disallow AI Training. Disallow AI Training is named for the Disallow: directive it publishes in your robots.txt. “Block” setting now means something different Block and “Block on pages with ads” previously did not apply to mixed-use crawlers because blocking them could also affect search discoverability. Now that we have the new Disallow AI Training setting, Block and “Block on pages with ads” apply to all training crawlers, including mixed-use crawlers. Training, Search, and Agent controls are applied at the domain level. With the addition of Disallow AI Training, the available settings are: Allow : All crawlers are allowed, unless blocked by another setting or a WAF rule. Disallow AI Training : Bot Preference Sync publishes the applicable no-training preference in robots.txt. Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search. Disallow AI Training is only available as a setting for Training, not Search or Agent. Block on pages with ads : Crawlers, including mixed-use crawlers, are blocked only on pages detected to be serving an ad. Block : All crawlers, including mixed-use crawlers, are blocked. Disallow AI Training works by publishing a preference in robots.txt. An ads-only preference cannot be expressed that way: Cloudflare can detect which pages serve ads, but that list is too large and changes too frequently to enumerate in robots.txt. That's why there's no Disallow AI Training on pages with ads. Agents do not create the same search-discoverability tradeoff as mixed-use crawlers, and the Internet does not yet have a well-established directive for expressing Disallow preferences to agents. For now, we’re not including a Disallow setting for Agents. As standards such as ai-prefs mature, we will revisit this approach. What changes on September 15? We are making the following changes to Bot Management and AI Crawl Control: Block and Block on pages with ads now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training. To stop training and keep search, use Disallow AI Training. “Block AI Bots” will be deprecated in favor of the more granular Search, Training, and Agent controls. Managed Robots.txt will be deprecated in favor of Bot Preference Sync. Customers who enabled Managed Robots.txt will migrate to the new system. Disallow AI Training will become part of the recommended configuration for certain new domains. Existing customers will have their preferences migrated to the new controls as described below. What you need to do Nothing, in almost every case. Your current settings carry over on their own. If you want mixed-use crawlers gone entirely, you now have to say so. Select Block. It will stop Applebot, Bingbot, and Googlebot from reaching your site — search included. Existing domains that never used the Search/Training/Agent controls Site owners that never configured the more granular controls will be migrated to the new settings based on their legacy Block AI Bots setting: Existing domains that previously configured the Search/Training/Agent controls For domains that previously configured the granular controls, we will preserve the practical effect of their selections under the new definitions. Previous Training selections of Block or Block on pages with ads will migrate to Disallow AI Training. Recommendations for new domains Beginning September 15, customers onboarding a new domain will be offered one of two preset configurations, depending on whether the site earns money from advertising. Ad revenue depends on a human actually see