AI 학습 무단 사용에 반대하는 아티스트 플랫폼 '카라(Cara)'가 2023년부터 약 150만 명의 작가를 모았으나, 8월 한 달에만 세 차례 대규모 스크래핑을 당해 서버 비용 폭증과 작가들의 불안을 초래했습니다. 첫 스크래퍼는 1,200만 점의 작품을 12TB 아카이브로 공개했으나 이후 자신의 행동을 후회하고, 장징나(Zhang) 대표와 함께 작가를 보호하는 새 오픈소스 도구 개발에 협력하기로 했습니다. 장징나는 법적 대응을 위한 크라우드펀딩으로 이미 10만 달러 이상을 모았습니다.
번역된 본문
2023년 초부터 사진가 장징나(Zhang)와 소규모 자원봉사팀은 이미지 공유 소셜 미디어이자 포트폴리오 앱인 '카라(Cara)'를 운영하기 위해 쉬지 않고 노력해왔다. 지금까지 약 150만 명의 아티스트가 이 플랫폼에 모였다. 그들을 끌어들인 것은 무엇이었을까? 자신들의 작품이 AI 모델 학습에 무단으로 사용되는 것에 대한 공동의 반대, 그리고 빅테크에 착취당하지 않으면서 자신의 예술을 알리고자 하는 열망이었다.
하지만 카라가 AI 이미지를 걸러내고 보호 기능을 제공함에도 불구하고—여기에는 스크래퍼가 수집하는 이미지의 스타일을 왜곡해 AI 모방을 방해하는 도구인 '글레이즈(Glaze)'도 포함된다—스크래핑 자체를 막는 것은 사실상 불가능하다. 그리고 바로 이번 달, 8월 13일부터 카라는 세 차례의 대규모 스크래핑을 당했는데, 이로 인해 서버 비용이 폭등했고, 인스타그램 같은 플랫폼(모든 콘텐츠가 메타(Meta)의 학습 데이터로 명시적으로 제공되는 곳)에서 이주해온 창작자들을 불안에 떨게 했다.
첫 번째 사건은 가해자가 카라의 1,200만 점 작품—공개된 이미지 라이브러리 사실상 전부—을 12TB 아카이브로 레딧 서브레딧 r/DefendingAIArt에 올리면서 드러났다. '즐거운 프로젝트였다'고 해당 레딧 사용자 만다린던포피994(MandarinDawnPoppy994)는 이후 삭제된 게시물에서 이 과정에 10달러도 쓰지 않았다고 썼다. '사용자들이 우리를 태그하면서 알게 되었다'고 장징나는 와이어드(WIRED)에 말했다. 스크래퍼가 '레딧에서 자랑하며 데이터셋으로 뭔가 하려는 사람들을 모집하고 있었기' 때문이다. 이는 AI 관련 포럼 전반에서 그가 저지른 행위의 윤리성에 대한 격렬한 논쟁을 촉발했다. '표적이 된 것 같아 매우 상처받았다'고 그녀는 덧붙이며, '법이 이런 데이터 수집에 대한 보호를 따라잡지 못하고 있다'고 지적했다. 즉, 스크래퍼들은 종종 기술적으로 합법하다고 정당화할 수 있다는 뜻이다. (장징나는 별도로 시각 예술가들이 제기한 두 건의 집단소송에 참여 중인데, 하나는 Stability AI와 미드저니(Midjourney) 등을 상대로, 다른 하나는 구글을 상대로 한 것으로, 이 회사들의 이미지 생성 도구가 자신들의 저작권 있는 작품으로 학습되었다고 주장한다.)
그러나 놀랍게도 카라의 작품 전체를 긁어간 이 사람은 자신의 행동을 후회하게 되었고, 아티스트를 보호하는 새로운 오픈소스 도구 개발에 장징나와 협력하기로 합의했다.
그와는 별개로 안타깝게도 다른 스크래퍼들은 계속해서 카라의 취약점과 부족한 자원을 악용했다. 상당수의 AI 지지자들이 카라를 노리는 것에 반대했지만, 소수는 만다린던포피994에 고무되어 장징나가 '모방' 공격이라고 보는 행위를 감행한 것으로 보인다. 두 번째 스크래퍼는 카라에서 약 850만 개의 링크와 사용자명, 제목, 태그 같은 메타데이터를 추출해 AI 개발자 플랫폼인 허깅페이스(Hugging Face)에 업로드했다. 허깅페이스는 삭제 요청이 쇄도하자 성명을 통해 사용자 '캡티브드리머(CaptiveDreamer)'에게 개인 메타데이터 삭제 통지는 보내겠지만, URL은 삭제할 수 없다고 밝혔다. '예술작품 사본이 여기에 호스팅되어 있지 않고' 링크는 '아티스트가 카라에 게시한 사본을 가리키기' 때문이라는 것이다. 회사는 '같은 근거의 추가 저작권 신고도 이 결과를 바꾸지 못할 것'이라고 결론지었다.
마지막으로 8월 22일, 세 번째 스크래퍼가 카라에서 12만 3천 장의 이미지와 개인정보가 포함된 텍스트 게시물, 사용자 소개를 확보해 아카데믹 토렌츠(Academic Torrents)라는 사이트에 모두 공유했다. 장징나는 이에 법률 비용을 위한 고펀드미(GoFundMe)를 시작하고 12만 달러의 목표액을 설정했으며, 이 돈은 사이버 법과 저작권법을 통해 카라를 방어할 수 있는 모든 전략을 모색하는 데 사용될 것이라고 설명했다. 목요일 기준으로 그녀는 10만 달러 이상을 모금했으며, 카라는 추가 법률 지원을 적극적으로 찾고 있다고 말한다. 장징나는 스크래핑 자체만이 아니라, 악의적인 행위자로부터 아티스트를 보호하기 위해 자신과 카라 팀이 실제로 무엇을 할 수 있는지에 대한 혼란 때문에도 좌절감을 느낀다. 일부 사용자는 이미 포트폴리오를 삭제하고 사이트를 떠났다고 한다.
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Since early 2023, photographer Jingna Zhang and a small crew of volunteers have worked tirelessly to maintain an image-sharing social media and portfolio app called Cara . So far, it has attracted about 1.5 million artists. What drew them to the platform? A shared opposition to the unauthorized use of their work to train AI models and a desire to publicize their art while avoiding exploitation by Big Tech. But while Cara filters out AI images and offers protective features—including Glaze , a tool meant to mask the style of the images picked up by scrapers in order to disrupt AI mimicry—preventing scrapes themselves is nearly impossible. And just this month, beginning on August 13, Cara was subjected to three major scrapes, which spiked its server fees and alarmed creators who had migrated there from platforms like Instagram, where all content is explicitly available to Meta as training data. The first of these incidents came to light when the individual responsible posted a 12-terabyte archive of 12 million works from Cara—more or less its entire library of publicly available images—on the subreddit r/DefendingAIArt. “It was a fun project,” wrote the redditor, MandarinDawnPoppy994, in his since-deleted post, saying the process cost him less than $10. “We actually found out about it through our users tagging us,” Zhang tells WIRED, since the scraper was “gloating and looking for other people to join him to do something with the dataset on Reddit,” sparking a fierce debate across AI-related forums about the ethics of what he had done. “I just feel it's targeted and very hurtful,” she adds, noting that “laws are not caught up on” protections against such data harvests, meaning that scrapers can often justify it as technically legal. (Zhang is separately part of two ongoing class actions brought by visual artists, one against Stability AI, Midjourney, and others and the second against Google, alleging that the companies’ image generator tools were trained on their copyrighted work.) In a surprising turn of events, however, the person who grabbed all the art off Cara would turn out to regret his stunt and agree to collaborate with Zhang on a new open-source tool to protect artists. In the meantime, unfortunately, other scrapers continued to take advantage of Cara’s vulnerabilities and minimal resources. While a number of AI proponents objected to going after Cara, a few were apparently emboldened by MandarinDawnPoppy994 to carry out what Zhang sees as “copycat” attacks. A second scraper pulled about 8.5 million links from Cara, as well as metadata like usernames, titles, and tags, and uploaded these to Hugging Face, the AI developer platform. After Hugging Face was bombarded with takedown requests, it responded in a statement that while it would issue a notice to the user, “CaptiveDreamer,” to remove the personal metadata, it could not do the same for the URLs, since “no copies of the artworks are hosted here,” and the links “point to the copies the artists published on Cara.” The company concluded that “further copyright reports on the same basis will not change this outcome.” Finally, on August 22, a third scraper obtained 123,000 images from Cara, along with text posts and user bios that included personal information, sharing it all on a site called Academic Torrents . Zhang then launched a GoFundMe for legal fees, setting a goal of $120,000, explaining that the money would go toward exploring any and all strategies of defending Cara through cyber and copyright laws. As of Thursday, she has raised more than $100,000, and she says Cara is actively looking for any additional legal assistance. Zhang is frustrated not only by the scraping but also by the confusion around what she and the Cara team can realistically do to shield artists from malicious actors, saying that some users have already deleted their portfolios and abandoned the site. “We have done the right things within limits without making it horrible to use,” Zhang says of the app’s current safeguards, including some new temporary measures like login gates—which, she adds, aren’t really a solution to an ongoing, internet-wide problem. And Zhang worries that people blaming her for the string of attacks may not understand that Cara has almost certainly been scraped before, like any other site; it cannot guarantee complete security. “If it makes them feel better, deleting your work and leaving Cara, I support that,” Zhang says of those leaving. “But I don't want to give people the misconception that if they go somewhere else, they are safer, because they're not. Bigger platforms get scraped more, so that makes me feel worse.” Yet Zhang, who is quick to remind WIRED that she is not a tech founder except by accident, has a newfound ally in this battle: the person who kicked off this month’s scraping frenzy, whom she confronted and eventually persuaded to apologize and delete his dataset. “He felt very bad to see how hurt people were,” Zhang says. “So he decided to help us.” “Heft” is a student in North America with a background in software and an interest in digital preservation and archival projects. (He requested that we refer to him only by one of his screen names due to doxing and death threats he says he received over the Cara situation.) In a conversation over Discord, Heft tells WIRED that scraping Cara was originally nothing more than a technical project and that he had no intention of making the data public. Heft says he nevertheless “made a foolish decision to attempt to ragebait with the dataset on Reddit” and “was carried away by trolling in the comments.” He knew it would provoke artists on Cara but did not anticipate the sheer anguish in that community. He saw people “sharing how they were having panic attacks over the scrape, how they deleted their entire portfolios from the internet.” In direct conversations with artists, he gained a greater appreciation of how personal their work was to them and how much they valued their ownership of it. “In retrospect, not only deliberately targeting Cara but presenting it the way I did in the post was cruel and thoughtless,” Heft says. “I missed the consequences that this would have beyond causing a bit of anger.” He notes that while he was initially curious as to whether the data could be successfully used in AI training, he never believed it would, since 12 million images “is not a lot to train an image model” and “most commercial AI labs train from large-scale web scrapes” available from open-data organizations like LAION. Random users may dump troves of data on Hugging Face, says Heft, but he considers it “highly unlikely that OpenAI or Anthropic is scanning every new Hugging Face dataset to train on.” Once shaken out of this theoretical mindset and moved to make amends, Heft joined Cara’s Discord server as a troubleshooter and has been identifying why a number of proposed fixes are unlikely to be successful. “He’s just helping us, you know, taking the time to explain” the structural weak spots that allow for scraping, Zhang says. When a certain tool or design comes up in the conversation, Heft lays out how he can “just break through in like a few minutes, literally,” she says. Because Heft is “of the belief that no site can be made truly ‘unscrapable,’” as he puts it, he and Zhang are collaborating on a countermeasure that can make a difference after the fact. The tool, Lantern , allows artists to create a “one-way fingerprint” for their images without storing them on the platform. It “regularly scans new publicly available AI image datasets,” Heft says, and if an artist’s work appears in one, “the artist receives a notification and a link to the dataset so they can request removal or potentially submit a takedown notice.” Lantern is just getting off the ground, and it’s the kind of imperfect workaround inspired by a regulatory vacuum. But Z