메뉴
BL
The Decoder 25일 전

학생 2만 6천 명 연구: AI 사용의 숨겨진 학습 부작용은 2년 만에 드러난다

IMP
8/10
핵심 요약

중국의 2만 6천 명 학생 대상 연구에 따르면, AI를 활용해 숙제를 빨리 끝내고 성적을 받은 학생들은 실제 시험 성적이 최대 24% 하락했습니다. 특히 AI가 독립적인 사고를 대체할 경우 학업 성취도 하락이라는 심각한 부작용이 2년에 걸쳐 서서히 나타나는 것으로 확인되었습니다. 이는 교육 현장에서 AI 도입 시 발생할 수 있는 장기적인 부작용에 대한 경각심을 일깨워줍니다.

번역된 본문

학생 2만 6천 명을 대상으로 한 연구에 따르면, AI가 가져오는 숨겨진 학습 비용은 온전히 드러나기까지 무려 2년이 걸립니다.

Jonathan Kemper가 THE DECODER가 제작한 프롬프트를 바탕으로 작성한 Nano Banana Pro 이미지

AI를 사용한 학생들은 숙제를 더 빨리 끝내고 더 좋은 평가를 받았습니다. 하지만 시험에서는 점수가 최대 24% 하락했으며, 입시 시험에서 나타나는 학업 성취도 격차의 전체 규모는 약 2년이 지나서야 완전히 드러났습니다.

중국 중부에서 진행된 새로운 연구는 AI를 사용하는 중·고등학생들이 겪는 학업 손실을 기록했습니다. 연구진은 100만 명 이상의 인구를 가진 군(county)에 위치한 7학년부터 12학년까지 26,000명 이상의 학생에 대한 30개월의 패널 데이터를 분석했습니다. 이 데이터에는 월간 시험, 숙제 점수 및 완료 시간, 고교 및 대학 입학을 위한 고위험 입학 시험 결과가 포함되어 있습니다.

연구 기간 동안 학생들의 자가 보고에 따른 AI 사용량은 0%에 가깝던 수치에서 약 80%로 증가했으며, 2024년 9월 DeepSeek V2.5와 2025년 1월 DeepSeek R1이 출시되면서 사용량이 급증했습니다. 가장 많이 사용된 도구는 Doubao, DeepSeek, ChatGLM, Ernie Bot, Qwen이었습니다.

이 연구는 학생들이 각자 다른 시기에 자율적으로 AI를 접하게 되었다는 점을 활용했습니다. 저자들은 이중 차분법(Difference-in-Differences)이라는 연구 방법을 사용했는데, 이는 특정 개입(여기서는 AI 사용) 전후에 실험군의 결과 변화를 측정하고 같은 기간 동안 비교군(미사용 학생)의 변화를 차감하는 방법입니다. 연구진은 각 학생이 AI를 사용하기 시작하기 전후의 성적 변화를 추적한 다음, 이 추세를 아직 AI를 사용하지 않는 학생들과 비교했습니다. 최초 사용 시점은 학생들의 자가 보고 데이터를 기반으로 하며, 인과적 주장은 두 그룹이 AI가 없었다면 비슷하게 발달했을 것이라는 가정에 근거합니다.

숙제는 향상, 시험 점수는 하락 AI를 처음 사용한 지 6개월 후, 숙제 점수는 18% 상승했고 과제당 평균 소요 시간은 64분에서 45분으로 줄었습니다. 동시에, 매월 치러지는 철저한 필기시험(closed-book exams)의 점수는 20% 하락했습니다.

고위험 입학 시험에 미치는 영향도 매우 컸지만 그 정도가 서서히 쌓이는 형태였습니다. 정기 시험의 성적은 반년 안에 떨어졌지만, 입학 시험에 미치는 전체 영향은 18%에서 24% 하락하는 수치로 나타나기까지 약 2년이 걸렸습니다. 따라서 연구진은 단기적인 연구만으로는 학습에 미치는 장기적인 비용을 파악할 수 없다고 지적했습니다.

장기 사용자 5명 중 4명은 업무를 AI에 위임하는 양상 5개월 이상 AI를 사용한 학생 중 약 81%는 50분 미만으로 숙제를 끝냈는데, 이는 AI를 사용하지 않는 가장 빠른 학생들보다도 빠른 속도였습니다. 이들은 숙제 성적은 높았지만 시험에서는 참패했습니다. 저자들은 이처럼 과제 완료 시간이 짧고 숙제 성적은 높으며 시험 점수는 낮은 이러한 조합이 학생들이 자신의 과제를 AI에 외주화(outsourcing)하고 있음을 시사한다고 분석했습니다.

반면, AI를 사용하지 않는 급우들과 비슷한 양의 시간을 숙제에 할애한 AI 사용자들은 더 좋은 숙제 성적을 받으면서도 시험에서도 동등하게 높은 점수를 받았습니다. 이 집단은 이전 성적을 기준으로 한 긍정적 선택(우수 학생만 모인 그룹)의 징후가 전혀 없었습니다. 즉, 이들이 애초에 공부를 더 잘하는 학생들이었던 것이 아니며, AI가 기본적으로 해로운 것도 아닙니다. AI는 주로 독립적인 사고를 대체할 때 학습에 악영향을 미칩니다.

사회과학 분야가 가장 큰 타격 정치와 지리 같은 사회과학 과목은 평균 27%의 하락세를 보였고, STEM(이공계) 과목은 22%, 영어는 17%, 중국어는 9% 하락했습니다. 이는 이전의 대부분의 실험이 수학, 프로그래밍 및 외국어에만 초점을 맞추었던 점을 고려할 때 의미 있는 데이터입니다.

또한 학생 그룹에 따라 효과 차이도 뚜렷했습니다. 저학년인 중학교 저학년 학생들이 고학년보다 더 큰 손실을 입었고(24% 대 17%), 남학생이 여학생보다 더 큰 타격을 받았습니다(21.6% 대 18.4%). 이는 주로 남학생들의 더 높은 AI 사용량 때문인 것으로 연구는 분석했습니다.

성적이 최상위권인 학생들이 가장 큰 피해를 입었는데, 상위 1/3의 학생들은 -24%의 효과를 겪은 반면 하위 1/3의 학생들은 -16%의 하락에 그쳤습니다. 용량-반응(dose-response) 패턴 역시 나타났습니다. 일주일에 1시간까지 AI를 사용하는 학생들은 약 5%의 하락을 보인 반면, 5시간 이상 사용하는 학생들은 그 이상의 큰 타격을 입었습니다.

원문 보기
원문 보기 (영어)
A 26,000-student study shows AI's hidden learning cost takes two full years to surface Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jul 4, 2026 Nano Banana Pro prompted by THE DECODER Students who used AI finished assignments faster and got better grades. On exams, though, their scores dropped by up to 24 percent, and the full scale of the learning gap on entrance exams didn't show up until about two years later. A new study from central China documents learning losses among secondary school students who use AI. The researchers analyzed 30 months of panel data from more than 26,000 students in grades 7 through 12 in a county with over one million residents. The data covers monthly exams, homework scores and completion times, and high-stakes entrance exams for high school and college. Self-reported AI usage rose from near zero to about 80 percent over the study period, with a big jump coinciding with the releases of DeepSeek V2.5 in September 2024 and DeepSeek R1 in January 2025. The most popular tools were Doubao, DeepSeek, ChatGLM, Ernie Bot , and Qwen . The study takes advantage of the fact that students discovered AI on their own at different times. The authors use a difference-in-differences design, a method that measures the change in outcomes for a treated group before and after an intervention and subtracts the change over the same period for an untreated comparison group. Here, they track how each student's performance shifted before and after they started using AI, then contrast that trend with students who weren't using AI yet. The timing of first use comes from self-reported data, and the causal claim assumes both groups would have developed similarly without AI. Better homework, worse test scores Six months after first using AI, homework scores rose by 18 percent while average time per assignment fell from 64 to 45 minutes. At the same time, scores on monthly closed-book exams dropped by 20 percent. The effect on high-stakes entrance exams was just as large but built up more slowly. Regular exam performance fell off within half a year, but the full impact on entrance exams took about two years to appear, ranging from an 18 to 24 percent decline. Short-term studies therefore miss the long-term cost to learning, according to the researchers. Four out of five long-term users show signs of outsourcing After more than five months of AI use, about 81 percent of students finished their homework in under 50 minutes, faster than even the quickest non-users. They got high homework grades but bombed exams. The combination of short completion times, high homework grades, and low exam scores suggests these students were outsourcing their work to AI, the authors write. AI users who spent a similar amount of time on homework as their non-AI classmates, on the other hand, scored just as well on exams while also earning better homework grades. This group showed no sign of positive selection based on prior performance, meaning they weren't simply better students to begin with, and AI isn't harmful by default. It causes damage mainly when it replaces independent thinking. Social sciences take the biggest hit Social science subjects like politics and geography saw an average decline of 27 percent, STEM subjects 22 percent, English 17 percent, and Chinese 9 percent. That matters because most previous experiments have focused on math, programming, and foreign languages. The effects also varied sharply across student groups. Younger students in lower secondary school lost more than older ones (24 versus 17 percent), and boys were hit harder than girls (21.6 versus 18.4 percent), which the study attributes mainly to heavier AI use among boys. Top performers suffered the most, with the top third seeing a minus 24 percent effect compared to minus 16 percent in the bottom third. A dose-response pattern showed up as well. Students using AI for up to one hour per week lost about 5 percent, while those using it five hours or more lost 30 percent. Why almost no one is pushing back The estimated learning penalty fell from about 25 percent in early 2023 to 16 percent by June 2025. The decline also showed up in a fixed group of early adopters, suggesting some degree of adaptation by students and teachers, but the losses haven't gone away. The study explains why the reaction has been muted. Teachers typically see students in only one subject, where a 20 percent grade drop isn't unusual on its own. The aggregate effect on the county average didn't reach about minus 10 percent until June 2025 because few students had been using AI long enough for the damage to accumulate. Students themselves often don't connect the dots, mistaking the mental effort of independent learning for a sign that they're learning poorly. As countermeasures, the study suggests giving students credible information about the long-term costs of outsourcing, putting more weight on in-person exams, and tracking completion time instead of homework grades. AI erodes the value of homework as a signal, and among AI users with above-average homework scores, higher homework grades actually predict worse exam results. Anthropic researcher Andrej Karpathy has argued that schools should stop trying to police AI-generated homework and instead shift the majority of grading to in-class work. His reasoning aligns with what this study found. When students know they'll be tested without AI, they stay motivated to actually learn the material. The pattern lines up with recent findings from other settings. An Anthropic study recently showed that participants who learned new programming skills with AI help scored 17 percent worse on follow-up knowledge tests than the control group, without saving any real time. The results depended on how people used the tool. Those who simply copied AI answers performed worse, while those who used AI to better understand the tasks didn't see the same decline. A study by the Swiss Business School found a negative link between AI use and critical thinking. A separate study by researchers at several American and British universities showed that people who treat AI mainly as an answer machine lose cognitive skills the fastest . A UC Berkeley study analyzing more than 500,000 grades also showed that the share of top A grades in writing- and programming-heavy courses has risen by 13 percentage points since ChatGPT launched. There, too, the effect was concentrated on unsupervised homework, while proctored exams showed no comparable gains. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->