메뉴
BL
The Decoder • 27일 전

최고 성적을 받는 역량일수록 AI가 가장 잘 위조한다

IMP
7/10
핵심 요약

이탈리아 보코니 대학 신입생 1,000여 명 대상 무작위 실험에서 GPT-4o 사용 학생들이 과제 점수를 크게 높였지만, 실제 학습이 개선됐는지는 검증되지 않았습니다. 현재 채점 기준이 완성도와 구조를 중시해 AI를 활용한 부정행위에 취약하다는 점이 핵심 문제로 지적됩니다.

번역된 본문

최고 성적을 받는 역량일수록 AI가 가장 잘 위조한다

핵심 요점

보코니 대학교 신입생 1,000여 명을 대상으로 한 실험에서 GPT-4o가 경영 과제 성적을 크게 향상시킨 것으로 나타났다. 인과 추론(causal reasoning)에 관한 짧은 강의는 기존 방식의 점수를 높이지는 않았지만, 학생들이 더 다양하고 독창적인 해결책을 내놓도록 이끌었다. AI 사용이 실제 학습도 개선했는지 여부는 여전히 답이 없다. ChatGPT 없이 진행된 후속 테스트가 없었기 때문이다.

ChatGPT는 학생 과제의 품질을 크게 향상시켰고, 교육적 개입은 더 다양하고 독창적인 해결책을 장려했다. AI가 학습 자체도 향상시키는지는 열린 질문으로 남아 있다.

무엇이 좋은 학생 과제를 만들며, 그중 어떤 부분을 AI가 개선할 수 있을까? 보코니 대학교 신입생 1,053명을 대상으로 한 무작위 실험에서 GPT-4o는 경영 과제에서 유의미하게 높은 성적을 받도록 도왔다. 인과 추론에 관한 짧은 강의는 기존 점수를 높이지 못했지만, 학생들이 더 다양한 해결책과 원인·가정·제안의 실제 작동 방식에 대한 깊은 사고로 이끌었다.

GPT-4o는 채점 성과에 큰 향상을 가져왔다

2025년 11월, 경영 입문 강좌 13개 분반이 통제군, 인과 추론 강의군, GPT-4o 접근군, 또는 두 가지 모두의 네 그룹으로 무작위 배정되었다. 학생들은 대학 기념품 매장을 위한 마케팅 제안을 180단어 이내로 작성해야 했다. 강의는 일관된 인과 논리, 반증 가능성(falsifiability), 그리고 제안된 행동이 어떻게 원하는 결과로 이어질 수 있는지를 다뤘다.

GPT-4o를 사용한 학생들은 1~5점 척도에서 거의 1점 가까이 높은 점수를 받았다. 그들의 답안은 평균적으로 약 2개 더 많은 아이디어를 담고 있었으며, 논리적으로 더 일관성 있었고, 세 명의 해당 분야 전문가들의 권고사항과 더 가까웠다. 연구진이 논증 품질, 아이디어 수, 아이디어 다양성, 텍스트 특성을 통제한 후에도 측정 가능한 GPT-4o의 우위는 남아 있었다. 저자들은 이를 학생 지식의 증가가 아닌 더 높은 콘텐츠 품질 때문으로 돌렸다.

인과 추론 강의는 기존 점수를 높이지 못했다. 오히려 이 그룹 학생들의 과제는 평균적으로 약간 낮은 점수를 받았지만, 제안된 행동이 왜 효과가 있어야 하는지, 어떤 조건에서 실패할 수 있는지를 더 자주 설명했다. 또한 다른 학생들이 쓴 내용과 달리 더 다양한 아이디어를 만들어냈다. 강의와 GPT-4o를 결합해도 기존 점수에는 추가적인 향상이 없었다. 인과 추론 지표에서는 두 접근 방식이 상호 보완적이었으며, 강의 그룹의 더 큰 아이디어 다양성은 유지되었다.

채점 기준은 관습적인 답안에 유리했다

더 많은 아이디어, 더 일관된 논증, 단일 답안 내 더 큰 아이디어 다양성은 모두 높은 점수와 상관관계가 있었다. 하지만 더 강한 반증 가능성, 제안된 행동의 작동 방식에 대한 더 상세한 설명, 다른 학생들과의 차별성은 오히려 낮은 점수와 상관관계가 있었다. 이것이 독창성이 전반적으로 페널티를 받았다는 의미는 아니다. 이 과제에서 기존 점수는 주로 예상되는 해결 범위 안에 머무는 잘 구조화된 답안에 보상을 주었다. 저자들은 다양성과 독창성이 점수에 반영되어야 한다면 채점 기준에 명시적으로 포함되어야 한다고 결론짓는다.

하지만 현재의 채점 시스템은 세련됨, 구조, 완전성을 측정할 뿐 학습과 이해를 측정하지 못한다. 이로 인해 AI를 부정행위 도구로 활용하기 쉬워진다.

이 연구가 보여주지 않는 것

학생들이 ChatGPT 없이 실제로 무엇을 이해하고 유지했는지 보여줘야 하는 후속 테스트가 없었다. 저자들은 GPT의 우위가 학생이 실제로 얻은 지식에서 비롯된 것인지, 아니면 단순히 AI 도움으로 더 나은 결과물이 나온 것을 반영하는지 불분명하다고 명시적으로 인정한다. 이 연구는 이 과제에서 채점 성과가 개선되었음을 보여줄 뿐, 학습이 개선되었음을 보여주지 않는다. 실험은 단일 대학의 신입생들이 좁은 범위의 마케팅 과제를 수행하는 것만을 다뤘다.

원문 보기
원문 보기 (영어)
The skills that earn top grades are the ones AI can fake best Tomislav Bezmalinović Aug 30, 2026 Nano Banana Pro prompted by THE DECODER Key Points An experiment at Bocconi University with over 1,000 freshmen shows that GPT-4o significantly boosted grades on a business assignment. A short lesson on causal reasoning didn't improve traditional scores but pushed students toward more diverse and unusual solutions. Whether AI use also improved actual learning remains unanswered. There was no follow-up test without ChatGPT. Ask about this article… Search ChatGPT significantly improved student work, while a teaching intervention encouraged more diverse and unusual solutions. Whether AI also improves learning remains an open question. What makes a good student paper, and which parts of that can AI improve? A randomized experiment with 1,053 freshmen at Bocconi University found that GPT-4o helped students earn significantly better grades on a business assignment. A short lesson on causal reasoning didn't raise traditional scores, but it pushed students toward more diverse solutions and deeper thinking about causes, assumptions, and how their proposals would actually work. GPT-4o delivered a major boost to graded performance In November 2025, 13 sections of an introductory management course were randomly split into four groups: control, causal reasoning lesson, GPT-4o access, or both. Students had to write marketing recommendations for the university's merchandise shop in up to 180 words. The lesson covered coherent causal logic, falsifiability, and how a proposed action might lead to a desired outcome. Ad Students with GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their answers contained about two more ideas on average, were more logically coherent, and more closely matched the recommendations of three subject-matter experts. Even after the researchers controlled for argumentation quality, number of ideas, idea diversity, and text properties, a measurable GPT-4o advantage remained. The authors attribute this to higher content quality, not greater student knowledge. Ad The causal reasoning lesson didn't raise traditional scores. Work from those students actually scored slightly worse on average, but they more often explained why a proposed action should work and under what conditions it might fail. They also generated more diverse ideas that diverged from what their peers wrote. Combining the lesson with GPT-4o didn't add any further boost to traditional scores. On causal reasoning markers, the two approaches complemented each other, and the greater idea diversity from the lesson group held up. Ad The grading rubric favored conventional answers More ideas, more coherent arguments, and greater idea diversity within a single answer all correlated with higher scores. But stronger falsifiability, more detailed explanations of how proposed actions would work, and greater divergence from other students' ideas correlated with lower scores. That doesn't mean originality was penalized across the board. For this assignment, the traditional score mainly rewarded well-structured answers that stayed within the expected solution space. The authors conclude that diversity and originality need to be explicitly built into grading criteria if they're supposed to count. Ad But current grading systems measure polish, structure, and completeness, but not learning and understanding. That makes it easy to use AI as a cheating tool. Ad What the study doesn't show There was no follow-up test where students had to demonstrate what they actually understood or retained without ChatGPT. The authors explicitly acknowledge that it remains unclear whether the GPT advantage came from knowledge students actually gained or simply reflected better output with AI assistance. The study shows improved graded performance on this assignment, not improved learning. The experiment only looked at freshmen at a single university working on a narrow marketing task. Randomization happened across 13 class sections rather than individually among all 1,053 students. The main performance score came from human graders who didn't know which group each text belonged to. Several additional measures like causal reasoning and idea diversity were evaluated using models from OpenAI and Anthropic. OpenAI provides the technology being studied and was involved in the research. Several authors work at OpenAI or were employed there during the study. Other studies ask what sticks after the AI goes away Research so far clearly suggests that what matters isn't whether students use AI but whether it supports their own thinking or replaces it. When it replaces it, students suffer. A study covering more than 500,000 US college grades found that top grades after ChatGPT's launch increased most in writing- and programming-heavy courses with a large homework component. Controlled experiments found that after AI was taken away, participants performed worse than a control group. The drop was steepest among those who had mainly used GPT to get direct answers. A 30-month study of more than 26,000 students in China showed a similar pattern over a longer period. Homework got better and faster with AI, but exam scores dropped. On later entrance exams, results were 18 to 24 percent lower over the long term. Students who spent roughly the same amount of time on homework as non-users despite having AI access didn't show comparable declines. Despite this growing body of evidence, OpenAI frames the results mainly as a challenge for grading systems: if AI can produce polished, expert-like work, the final product alone says less about what students actually understand. The authors argue for "rewarding students for producing work that reflects originality, reasoning, and consideration of multiple approaches, not just the most conventional or polished answers." That may be true, but changing the grading system likely means changing the entire education system along with it. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI