메뉴
HN
Hacker News 36일 전

AI가 우리가 아는 학계를 이미 파괴했는가?

IMP
8/10
핵심 요약

AI의 발전으로 인해 과제 제출물과 연구 논문의 '양적 생산'이 사실상 무한대에 가까워지면서 기존 학계의 운영 방식이 근본적으로 붕괴하고 있습니다. 이제 AI를 고도로 활용하는 학생과 연구자가 성적과 업적 면에서 독자적 연구를 하는 이들을 압도하며, 이는 결국 인간의 진정한 지식과 지적 노동을 평가하는 시스템 자체의 실패를 의미합니다.

번역된 본문

이 글을 작성하는 데 AI는 전혀 사용되지 않았습니다.

만약 학계가 하나의 게임이라고 가정한다면, 저는 그 게임에서 이긴 셈입니다. 종신교직권(Tenure), 기금을 지원받는 연구 석좌, 수상 경력, 리더십 직책, 제가 설립하는 데 도움을 주어 지금은 제가 편집장(Editor-in-Chief)으로 있는 국제 저널, 각자의 성공을 이끌어낸 지도 학생들, 준수한 h-지수(h-index) 등, 성공의 고전적인 척도들은 모두 갖추고 있습니다. 이는 자랑하기 위함이 아니라, 제가 이 게임에서 이기긴 했지만 이 게임이 더 이상 말이 안 통하게 되었다는 점을 지적하고자 함입니다. 대부분의 학자들이 실천해 온 학계는 소위 '양으로 승부하는(Maximalism)' 방식으로 굴러갑니다. 가장 많은 연구 기금, 가장 많은 논문, 가장 많은 학생, 가장 많은 상, 가장 많은 언론 보도가 그것입니다. 비록 요즘은 연구의 영향력과 기여도를 강조하는 데 훨씬 나아졌지만, 근본적인 엔진은 여전히 '물량'입니다. 그리고 그 물량은 항상 인간의 독립적인 글쓰기(연구비 신청서, 논문 투고고, 추천서, 보고서, 칼럼 기고, 보도자료 등)를 통해 생산되었습니다.

문제는 AI가 이러한 물량을 사실상 무한대로 만들어버린다는 것입니다(지구가 멸망하기 전까지야 가능하겠지만, 이는 별개의 이야기입니다).

과제가 가장 뚜렷한 희생양입니다. 대중에게 이미 보이는 부분부터 이야기해 보겠습니다. 학생이 집으로 가져갔다가 제출하는 모든 과제는 실질적으로 AI가 생성했거나 AI로 다듬어졌을 확률이 극도로 높습니다. 지금까지 우리는 종종 이러한 AI 사용을 적발할 수 있었는데, 이는 일부 학생들이 아직 AI를 서투르게 사용하기 때문입니다. 그들은 전형적인 Chat GPT 서식, 모든 문장에 들어가는 쉼표로 구분된 3가지 항목의 목록, 환각(Hallucination) 현상으로 인한 잘못된 인용, 특유의 과장된 표현, 단락 들여쓰기 누락 등 눈에 띄는 쓰레기 같은 글(Slop)을 제출합니다. 우리는 그런 학생들을 적발하며 여전히 통제권을 쥐고 있다고 착각합니다.

하지만 진짜 심각한 문제는 우리가 결코 알아채리지 못하고 이미 적발망을 빠져나가는 것들입니다. 예를 들어, Claude와 ChatGPT라는 두 개의 유료 계정을 가진 학생이 있다고 가정해 봅시다. 한 AI에게 과제 초안을 작성하게 하고, 다른 AI에게 이를 비판하고 다듬게 하는 과정을 글이 완벽해지고 논리가 탄탄해질 때까지 반복합니다. 또한 AI를 통해 참고문헌을 두세 번 확인하고 모든 서식과 문장 부호를 완벽하게 맞춥니다. 이 학생이 만들어낸 결과물은 적발할 수 없을 뿐만 아니라, 대부분의 다른 제출물보다 훨씬 훌륭하므로 더 높은 성적을 받게 됩니다. 이런 AI 극대화 학생들은 AI 사용과 성적 간의 명백한 연관성을 보기 시작하면서, 이른바 '게으른' 학생이나 '부정직한' 학생이 아니라 철저히 '합리적인' 존재로 변모하게 됩니다. 가장 심각한 문제는 현재의 시스템이 다음의 두 가지를 동시에 한다는 것입니다. 즉, 자연스러운 결함과 한계를 지닌 온전히 인간이 쓴 에세이를 제출한 학생에게는 불이익을 주고, 우리가 적발한 서투른 AI 사용자에게는 0점을 주면서도, 고도로 숙련된(그리고 돈을 더 많이 쓴) AI 사용자에게는 보상을 준다는 것입니다. 만약 여러분의 수업에 학생들이 독립적으로 수행해서 성적을 매기는 학기말 논문이 있다면, 십중팔구 여러분(혹은 여러분의 조교, 또는 조교의 AI)은 실제 내용에 대한 진정한 지식과 무관한 성적을 부여하고 있는 셈입니다.

하지만 연구 분야의 문제야말로 제 개인적으로 가장 큰 충격을 줍니다. 우리 섹터에서는 AI를 둘러싼 교육 및 학습 문제에 대해 많은 논의를 해왔습니다. 하지만 지난주 제 연구팀원들에게 말했듯이, 연구 및 전반적인 학문적 성공이라는 차원에서 이것이 어떤 의미를 갖는지에 대해서는 여전히 현실을 외면하고 있는 듯합니다. 대량 생산된 출판 가능한 콘텐츠는 이미 우리 곳에 와 있습니다. 문헌 고찰(Review articles), 방법론 논문, 이론적 종합, 보고서, 질적 데이터의 2차 분석 등; 오늘날의 연구자는 Consensus나 Claude 같은 도구에 프로(Pro) 구독을 결합하여 이러한 결과물을 대량으로 생성해 낼 수 있으며, 이 중 상당수는 출판되기에 충분할 만큼的质量을 갖추게 됩니다. 물론, 일부 심사자들은 일부 투고 논문이 너무 가볍고 부피만 키운 것처럼 보인다는 것을 알아챌 것입니다. (하지만 다시 말하지만, 이는 단지 도구를 최적으로 사용하지 못했기 때문이라고 생각합니다. 과장된 표현이나 빈약한 전제들은 AI에서 얼마든지 훈련을 통해 없앨 수 있습니다.) 만약 소화기에서 물을 뿜어내듯 논문을 쏟아낸다면 상당수가 심사를 통과할 것입니다. 이런 방식으로 일할 의향이 있는 사람은 하루에 한 편에 가까운 논문을 작성할 수 있습니다. (다만, 다소 불편한 온라인 제출 시스템 때문에 속도가 약간 느려질 수는 있습니다.) 그리고 그 사람의 이력서(CV)는 독립적인 지적 노동을 하는 그 누구의 것보다도 빠르게 압도적으로 성장할 것입니다. 이건 바로...

원문 보기
원문 보기 (영어)
No AI was used in writing this post. If academia was a game, I've won it. Tenure, an endowed research chair, awards, leadership positions, an international journal I helped to found and now serve as the Editor-in-Chief, students I have supervised to their own successes, a good h-index, all the classic marks of success. This isn't meant as bragging but rather to point out that while I've won this game, the game no longer makes sense. Academia, as most of us have practiced it, runs on maximalism. The most grants, the most papers, the most students, the most awards, the most news coverage. While we are doing much better these days in highlighting impact and contributions, the underlying engine is still volume, and the volume has always been produced by independent human writing (applications, submissions, letters of support, reports, Conversation articles, press releases, etc., etc.). The problem is that AI makes volume essentially infinite (until the world burns up, but that's a parallel discussion). Assignments are the most obvious casualty I'll start with the part that is already visible to the general public. Any assignment a student takes away and brings back is, for all practical purposes, extremely likely to be AI generated or AI refined. To date we've often been able to detect this use and this is because some students still use AI badly. They submit the obvious slop with classic Chat GPT formatting, comma-separated three item lists in every sentence, the hallucinated citation, the tell-tale hyperbole, lack of paragraph tabs, etc. We catch those students and we feel like we're still on top of things. But the real obvious problems are the ones we'll never notice and that are already passing by detection. Take a student with two paid accounts, say Claude and ChatGPT, who has one AI draft the work and the other critique and refine it, looping until the prose is clean and the argument is tight. The have AI double and triple check references, they nail every bit of formatting and punctuation. That student produces work that is not only undetectable, it is better than most of what gets submitted, and it will therefore earn a higher grade. These AI-maximizing students become the rational ones rather than being 'lazy' or 'dishonest' because they start to see the obvious connection between AI use and grades. Most egregiously, the system now does two things: it penalizes the student who wrote their own merely human essay with natural flaws and limitations, and it hands zeros to the unsophisticated AI users who we catch, while rewarding the sophisticated (and higher spending) ones. If your class has a term paper that students do on their own and submit for grading, chances are that you (or your TA (our your TA's AI)) are assigning grades unrelated to real knowledge of the content. But it's the research issues that really hit me personally We've been talking as a sector a lot about the teaching/learning issues around AI but as I told my research team last week, it seems like we're still 'head in the sand' about what this means in terms of research and overall academic success. Mass produced, publishable content, is ALREADY HERE. Review articles, methodology pieces, theoretical syntheses, reports, secondary analyses of qualitative data; a researcher today can generate these in volume by combining a couple of pro subscriptions to tools like Consensus and Claude, and a significant share of these will be good enough for publication. Sure, some reviewers will spot some article submissions as being too fluffy (but again, I still think that's just not using the tools optimally, you can train AI away from all the hyperbole and empty premises) but if you're blasting them out like a firehose, a lot will get through. Someone willing to work this way can produce something close to a paper a day, slowed down a bit by online submission system clunkiness, and their CV will quickly eclipse anyone doing independent intellectual work. It's the same issue with grant submissions, restrained only a bit by limits on how many a single researcher can submit or hold simultaneously. Picture a team of five colleagues running ten applications into a single CIHR Project Grant cycle by rotating which member sits as nominated principal investigator (each can submit 2 per cycle). The odds of landing at least one are high based on volume alone, before you even account for the fact that AI is genuinely good at some of the common critical errors that sink applications: budget flaws, a highly relevant paper the team missed citing, the eligibility criterion that was maybe flagged so late in final review they decided they didn't have time to fix it. The careful, error-free, comprehensive application used to be the outcome of several failed submissions, now it's just someone who knows how to use multiple AIs or use a cowork/agent system. What's CIHR even going to do when the number of applications triple? What are they going to do when AI submissions are better than human developed ones? So far, the discussion about dealing with this volume is thinking about AI pre-screening of applications. So your AI is now checking my AI...cool, cool, cool. I don't want this to sound like sour grapes like I'm worried that junior scholars are going to outpace me. Rather, I'm worried that academia as a whole careens into nonsense because we haven't adjusted our reward systems to match the current reality. We will pretend this isn't happening for a while The institutional response has been reasonable in terms of coursework and assignments. Due to the complexities of academia, including academic freedom, de-centralized structures, unionized contracts, etc., there won't be rapid, centralized responses about course assignments. Rather, universities are providing supports and guidance to redesign assessments, redesign syllabi, and providing cheating prevention software for certain remote assessments. Many scholars have written more eloquently than I can about processes to ensure learning is occurring and evaluation is meaningful. Yes, going back to paper and pencil strains our current resources, but is a likely necessity. On the research side, the response has seemed far slower. From Tri-Councils initially banning AI use to then allowing it, and most journals having very limited responses beyond perhaps self-declarations, it seems we are already 2 years behind the reality. Indeed, we continue to run on our former processes and metrics while an entirely new system is in place that essentially negates these metrics. The version of academia whereby you submit written content and are rewarded for how much of that written content is taken up in formal venues is already dead in terms of meaning. We just haven't gotten around to holding a funeral yet.