메뉴
HN
Hacker News 7일 전

레딧, 단순 HTML을 '불안전'하다며 구버전 로그인 강제

IMP
7/10
핵심 요약

레딧이 기존 old.reddit.com의 비로그인 접속을 차단하며, 그 이유로 악성 스크래핑과 자동화 트래픽으로부터 플랫폼을 보호하기 위함이라고 밝혔습니다. 하지만 신버전은 여전히 비로그인 상태로 접속할 수 있어, 진짜 목적은 AI 데이터 스크래핑 방어와 데이터 수익화에 있다는 비판이 제기되고 있습니다. 오래되고 단순한 HTML 기반 UI가 기술적으로 보호하기 어렵다는 핑계로 사용자 선택권이 제한되는 사례입니다.

번역된 본문

레딧이라는 기업 레딧을 모르신다면, 기본적으로 수많은 인기 포럼을 호스팅하는 플랫폼이라고 생각하시면 됩니다. "고양이를 보러 왔다가 공감대를 느끼며 머문다"는 슬로건을 내세우는 다른 회사들과 마찬가지로, 레딧 역시 포럼을 완전히 파괴하지 않는 선에서 커뮤니티에서 최대한의 가치를 창출하는 사업을 하고 있는 것 같습니다. 결국 단순히 커뮤니티를 육성하는 것만으로는 새로운 기술 시대의 기업에게 더 이상 숭고한 목표가 아니며, 레딧에게 다행스럽게도 진짜 사람이 만든 데이터는 LLM(대형 언어 모델) 시대에 금과도 같습니다. 레딧이 커뮤니티들에서 비유적인 의미의 간을 떼어가기로 결정했던 지난 일에 대해 읽어보셔도 좋습니다.

레딧-검색 결과 저는 더 이상 레딧에 참여하고 싶지 않지만, 특히 LLM 시대에 가끔씩 방문합니다. 검색어에 site:reddit.com을 추가하면 진짜 사람이 쓴 결과를 찾을 수 있는 확실한 방법이기 때문입니다. 이 점을 아주 명확히 말씀드리자면, 저는 여전히 사람이 직접 쓴 글을 선호합니다.

이제 로그인을 해야만... 어느 정도는 어제 그런 검색을 했더니, 제 기쁨을 뒤로하고 다음과 같은 화면이 반겨주었습니다. (아마도 이것이 제가 늙었다는 것을 의미하겠지만) 저는 레딧의 오리지널 프론트엔드인 old.reddit.com을 사용합니다. 왜 그것이 더 나은 프론트엔드인지 설교하려고 애쓰지는 않겠지만, 그저 그것은 레딧을 위한 하나의 디자인일 뿐입니다! 알겠습니다. 저는 제가 소수파 사용자라는 것을 압니다. 저는 Firefox를 씁니다. 저는 noscript를 사용합니다! 여전히 그것을 사용하는 두 명의 사람들을 더 이상 지원하지 않기로 결정했기 때문에 완벽하게 잘 작동하던 제품에서 쫓겨나는 일에 익숙합니다.

저의 8년 된 스마트폰에게도 이런 일이 일어났습니다. (그래도 아주 잘 작동하지만, 운영체제가 너무 구버전입니다. 작은 영광 속에서 평안히 잠들길 바랍니다.) 10년 된 태블릿에게도 같은 이유로, 제 폰보다 훨씬 더 잘 작동하고 배터리도 짱인데 말이죠. 앱 스토어에서 오래전에 삭제된 구버전 앱이 사용하던 Stack Exchange의 API에도 일어났습니다. 이건 저를 가장 아프게 했습니다. 1

제가 개인적으로 받아들이는 부분은, 레딧이 이 모든 것이 '레딧을 안전하게 유지하기 위해서'라고 말하고 있다는 것입니다.

도대체 누구로부터 레딧을 안전하게 지킨다는 건가요? 저로부터요?? 제 지식에 대한 열망으로부터요?? 그렇다면 안전이라는 건 정확히 무엇일까요? 자, 더 이상 레딓을 읽지 않기 때문에 당연히 못 봤지만, 이 문제에 대해 그들이 어떻게 말하는지 공지사항 페이지(https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/)를 보며 확인해 보겠습니다. (물론 접속하려면 로그인이 되어 있어야 합니다.)

구버전 레딧의 비로그인 환경은 플랫폼에서 악성 스크래핑(scraping) 및 자동화 트래픽의 주요 원인입니다.

음, 알겠습니다. 그런데 왜 신버전 레딧은 여전히 로그인 없이 접속할 수 있을까요? 아, 누군가가 그 점을 물어봤네요.

[질문]: 사람들이 신버전 레딧을 스크랩하지 않는 게 무엇이 다르길래 그런 건가요? 저에게는 구버전 레딧에도 그걸 적용하는 것이 더 나아 보입니다. 게다가, 이렇게 하면 사람들이 신버전 레딧을 스크랩하려고 하지 않을까요?

[관리자 답변]: 답변을 적으려던 참이었는데, 마침 u/Nestramutat-님이 다른 댓글에서 정말 훌륭한 답변을 남기신 것을 보았습니다!

[해당 댓글 (요약)]: ... 첫 번째 질문에 대해 말씀드리자면, 악성 트래픽의 형태는 항상 변합니다. 하나의 방법을 차단하면 새로운 방법이 개발되는 끊임없는 술래잡기 게임이 될 것입니다. 악성 트래픽은 사후에 보기에는 쉽지만, 선제적으로 차단하는 것은 더 어렵습니다. 그들이 구버전 레딧에 모던 보안 스택(modern security stack)이 없다고 주장하는 것을 보아, 이것은 훨씬 더 큰 도전이 되고 있는 것이 분명합니다...

그렇다면 구버전 레딧에는 "모던 보안 스택"이 없다는군요. 저는 풀스택 개발자로서의 명성을 얻은 것은 아니지만, 마우스 우클릭 후 '검사(Inspect Element)' 정도는 할 줄 알기 때문에 그 이유를 알아볼 만한 자격은 있다고 생각합니다.

구버전 vs 신버전 old.reddit.com이 대체 뭐가 그렇게 불안전한지 확인해 보기 위해 마지못해 로그인을 해보겠습니다. 테스트용으로 그들의 공지 스레드를 사용하겠습니다. 아, 훨씬 낫네요. 구버전 레딧이 무슨 짓을 하길래 그렇게 불안전한지 확인해 보겠습니다.

제가 추측할 수 있는 가장 그럴듯한 이유는, 그들 귀중한 사용자 생성 콘텐츠가 너무나도 쉽게 평문 HTML(plain HTML) 형태로 노출되어 있기 때문입니다.

원문 보기
원문 보기 (영어)
Reddit-The-Company If you don’t know Reddit, it basically is the host of many popular forums. And like any company which encourages you to “come for the cats [and] stay for the empathy,” Reddit seems to be in the business of extracting as much value as it can from said forums without completely destroying them. After all, simply fostering community is not a noble enough goal for the New Tech, and fortunately for Reddit, genuinely human-generated data is now gold in the LLM Age. You are welcome to read about the last time they decided to pluck the metaphorical liver from their communities . Reddit-The-Search-Results While I no longer wish to engage with Reddit, I still visit it occasionally, especially in the LLM Age. This is because appending site: reddit.com to a search query is basically a surefire way to find results written by genuine humans. Which, just to be extremely clear, I still find desirable. Now behind a login… sort of After doing such a query yesterday, to my absolute delight I was greeted with I guess this makes me old, but I use the original frontend for Reddit, old.reddit.com . I’m going to try really hard not to preach about why it’s a better frontend, but that’s all it is! A design for Reddit. Look, I know I’m a fringe user. I use Firefox. I noscript! I am no stranger to being forced off a product that worked just fine because someone decides to no longer support the two people who still use it. It happened to my phone of 8 years, which works fine by the way, but is on too outdated of an OS. May it rest in peace in its tiny glory. It happened to my tablet of 10 years, which works even better than my phone and holds a charge like champ, for the same reason. It happened to the API for Stack Exchange used by my copy of their outdated app long removed from the app store. This one hurt me the most. 1 Where I feel like things get personal here is that Reddit is saying that this is all To keep Reddit safe Keep Reddit safe from whom exactly? Me? My desire for knowledge?? Safety is what exactly? So let’s see what they have to say on this matter by going to the announcement, which of course I didn’t see because I don’t read Reddit anymore: https://old.reddit.com/r/modnews/comments/1ujtebf/logging_in_to_use_old_reddit/ . Hope you’re logged in. Old Reddit’s logged-out experience is a significant source of abusive scraping and automated traffic on the platform. Hmmm, OK. But then why is New Reddit still accessible logged out? Oh, someone asked that. [Question] : What’s so different about new reddit that people don’t try to scrape that? Seems to me like it would be better to just implement that on old reddit too. Besides, won’t this just cause people to try and scrape new reddit? [Admin reply] : I was about to type an answer but just saw u/Nestramutat- gave a really eloquent answer in another comment! [The comment (snipped)] : … To your first question, the shape of malicious traffic is always changing. It’s going to be a constant cat and mouse game as you ban one method, a new one gets developed. It’s easy to see abusive traffic in hindsight, but it’s harder to pre-emptively block it. Given that they’re claiming Old Reddit doesn’t have the modern security stack, this is likely proving to be an even greater challenge… So it doesn’t have “the modern security stack.” Now I may not have really earned my Full Stack stripes, but I can right click and select Inspect Element so I’d say I’m qualified enough to see why. Old versus New Old Reddit Let’s start by — begrudgingly — logging in to see what is so insecure about old.reddit.com . I’ll use their announcement thread to test. Ahhh, so much nicer. Let’s check what Old Reddit is doing that makes it so insecure. The best I can guess is that their precious, precious user-created content is available in plain HTML, since that’s basically all Old Reddit does: you don’t even need JS unless you want to load more comments (ask me how I know). It DLs about 1 megabyte and sends about half a megabyte. It’s not shown there, but the page’s HTML itself comprises most of the response. I’ve certainly seen worse, but what’s with the load time? GitHub loads its massive payload about 4x as fast (relative to size). Oh… 2 whole seconds of waiting for a reply. Smells of rate-limiting. Or was Reddit always this slow? Let’s load some more comments. (Reddit never loads all of the comments initially) Well that was nice and lean, and pretty snappy. I don’t see my secrets being sniffed. All I got here is that Old Reddit is a pretty normal webpage, which I guess makes it insecure in comparison to… New Reddit Let’s see why New Reddit is so much better. In case you have unrealistic expectations, let me right them: New Reddit will not load anything more than the post itself without Javascript (JS). That’s probably what makes it more secure. There’s a lot loading here (about 5x Old Reddit), and this is why my analysis gets rather unscientific. Rather than try to get around the “security things” (whatever that means), I instead tried to do the bare minimum necessary to fetch the content. In browser — I did not feel like writing a scraper. This led to me basically blocking all requests to domains (including reddit.com ) except for www.redditstatic.com/js/concat www.reddit.com/svc/shreddit/more-comments/ www.reddit.com/svc/shreddit/comment/ When you do this, the page loads a lot less, but it does load. When you click to load more comments, it spins forever, but I inspected the request fired off and it did get a response with comment text. So as far as I can gather, simply running the JS on the page is sufficient to get enough information to get comments. So I guess that’s what’s stopping the scrapers? Executing Javascript? Just for the heck of it, let’s load some more comments without the request filter. Well, that’s certainly less lean than Old Reddit. In my (again, unscientific) experimenting, I reloaded the page several times and tried to load comments and replies and didn’t get any failures. One time, when I had the request filter off, I got redirected and saw a captcha field sent as a query param, but I couldn’t reproduce that. I don’t know whether you get a captcha if you just gun it directly for the comments. But I did notice this helpful heartbeat sent back to Reddit every time I scrolled or moved my cursor. I guess that makes me feel safer? So what’s safer? Let’s dispel any notion that there are safety issues arising from a frontend, because that’s pure PR crap. Instead, if we read between the lines, Reddit doesn’t want people scraping (because it’s their gold, dammit!) and they think that shafting a few Old Reddit users will disrupt the scraping enough. Does this actually stop scraping? I don’t know! I won’t claim to have proven you can still scrape New Reddit: surely it must be harder, but it seems to just be that scrapers — surprise, surprise — prefer using a leaner form of Reddit. Which, by the way, loads 8x more comments by default (200 vs 25; yes, New Reddit really only loads 25 comments initially, going up to like 35 automatically if you scroll some). Cynically, it seems like Reddit discovered they can make scraping harder by making your browser load more and work more, which they had already done by rewriting a nice piece of HTML into a 5x more bloated mess of web components or whatever. (I gave up trying not to editorialize, sorry) The thing is that I wouldn’t have beef with Reddit if they had just quietly issued a 40X/30X error for old.reddit.com . Or if they had said “no one uses this, we’re removing it.” This