메뉴
HN
Hacker News • 1일 전

AI가 모든 과제를 해내버린 후, 내가 수업을 바꾼 방법

IMP
7/10
핵심 요약

카네기멜론 대학의 Christian Kästner 교수는 ChatGPT 출시 전부터 GPT-3가 자신의 수업 퀴즈를 통과하는 것을 보고, 5년 후 AI 에이전트가 모든 과제를 수행할 수 있게 되자 평가 방식을 대폭 재설계했습니다. 학습 목표는 그대로 유지하되, 집에서 수행하는 과제로 이해도를 평가하는 것을 포기하고 시험, 조교와의 상호작용, 비디오 데모 중심으로 전환했습니다. 특히 AI 사용을 금지하는 대신 허용하고, 재제출에 10% 페널티를 부과하는 등 증거 기반 교육법 모범 사례를 포기하는 현실적 트레이드오프를 소개합니다.

번역된 본문

AI가 내 숙제 과제를 전부 해낼 수 있게 된 후, 내가 수업 방식을 바꾼 방법 Christian Kästner, 2026년 9월 23일

ChatGPT가 출시되기 한참 전인 2021년경, Vincent Hellendoorn이 내 수업의 독해 퀴즈(reading quiz)에 GPT-3를 사용해보라고 제안했다. GPT-3는 지정된 논문을 실제로 읽지 않고도 채점 기준(rubric)을 통과하는 그럴듯한 답변을 만들어냈다. 당시 나는 아무것도 바꾸지 않았다. 5년 후, AI 에이전트는 내 과제 전부를 해낼 수 있게 되었고, 학생들이 배워야 할 내용은 거의 바뀌지 않았음에도 나는 그 수업의 평가 방식 대부분을 재설계했다. 전략은 언제나 같다. 집에서 하는 것으로 이해도를 테스트하는 것을 그만두고, 대신 조교(TA)와의 상호작용, 시험, 비디오 데모에 집중한다. 이러한 변화 중 일부는 증거 기반의 교육 모범 사례에 어긋나지만, 그럼에도 나는 그렇게 했다.

최근 몇 년간 나는 주로 '프로덕션 머신러닝(Machine Learning in Production)'이라는 수업을 가르쳤다. 이 수업은 ML 모델을 중심으로 프로덕션급 소프트웨어를 구축하는 것을 다루는 상위 과목으로 MLOps에 중점을 두고 있으며, 보통 100~170명의 학생이 수강한다.

요즘 다른 교육자들과 대화할 때 흔히 나오는 질문은 생성형 AI와 코딩 에이전트 시대에 우리가 어떻게 교육을 바꿨는가이다. 그래서 우리가 한 일을 개략적으로 설명하려 한다.

우리는 AI 혁신과 도구의 변화에 따라 다루는 주제는 조정했지만, 전반적인 학습 목표는 거의 건드리지 않았다. 나는 이것이 입문 과목이 아니라는 점, 그리고 학습 목표가 코드 작성이나 특정 도구 사용이 아니라는 점이 다행이라고 생각한다. 학습 목표는 엔지니어링 트레이드오프, 리스크의 예측과 완화, 팀워크에 관한 것이다. 일부는 모델로 시뮬레이션하거나 위임할 수 있다 하더라도, 나는 이러한 기술이 여전히 익힐 가치가 있다고 생각한다. (입문 과목이나 전통적인 소프트웨어 엔지니어링 과목을 수정한다면 학습 목표가 훨씬 크게 바뀔 것이다.)

어쩌면 중요할 수도 있는 점: 우리는 학생들에게 필기·구술 시험을 제외한 모든 상황에서 어떤 형태로든 출처 표기 없이 AI를 사용할 수 있는 권한을 부여한다. 우리는 오히려 많은 곳에서 AI 도구 사용을 권장한다. 원한다 해도 AI 사용을 단속하는 것은 실현 가능하지 않다고 생각하고, 더 중요하게는 학생들이 어차피 이 기술의 책임 있는 사용법을 배워야 한다고 생각하기 때문이다.

AI가 증거 기반 모범 사례를 포기하게 만들고 있다

실제로 우리가 하는 일보다 이 점이 더 중요하기 때문에 먼저 짚고 싶다. 안타깝게도 AI는 여러 증거 기반 교육 관행을 실질적으로 무너뜨리고 있다 (예: 'How Learning Works', 'The ABCs of How We Learn' 참고).

예를 들어, 연구 증거는 소수의 고부담 평가(시험)보다 피드백이 있는 빈번한 저부담 평가(과제, 퀴즈)를 선호한다. 하지만 AI는 저부담 환경에서의 연습을 무너뜨리고 우리를 시험 쪽으로 밀어붙이고 있다.

마찬가지로 나는 항상 학생들이 실수를 하고 잃은 점수를 되찾기 위해 제한된 수의 과제를 다시 제출할 수 있는 안전망을 제공해 왔다 (학습 결과에 집중하고 과정이 아닌 결과를 평가하라는 스펙 그레이딩과 공평한 채점의 핵심 권고사항이다). 하지만 이 과정이 AI로 악용되는 것을 느꼈다. 생각 없이 AI가 생성한 과제 답안을 먼저 제출하고, 재제출 시에만 채점에서 지적된 문제를 확인하는 식이다 (AI 사용 비용을 외부화하는 전형적인 패턴이다). 이에 대응해 우리는 재제출에 10%의 페널티를 부과하기 시작했다.

또한 수업 중 상호작용은 초기의 저부담 환경에서 학습 자료에 몰입할 수 있게 해주지만, AI가 등장한 후 많은 학생 그룹이 토론 질문을 모델에 떠넘기는 것을 목격했다. 손으로 쓴 제출(pen-and-paper)로 해결할 수도 있지만, 채점 업무량 증가 외에도 학생들의 스트레스를 높이고, 저부담 환경을 해치며, 피드백을 지연시킬 것이다.

일반적으로 이것은 균형 잡기의 문제이며, 나는 악용될 수 있음에도 저부담의 반복적 상호작용을 유지하는 쪽을 택하는 편이다. 그렇다, 일부 학생은 깊은 학습 없이 수업을 통과할 것이다. 하지만 배우고자 하는 학생들에게는 더 나은 환경을 제공한다. 나는 내가 독일에서 직접 경험했던, 대부분 선택 과제였던 교육 모델로 돌아가고 싶지 않다.

원문 보기
원문 보기 (영어)
How I changed teaching after AI managed to do all my homework assignments Christian Kästner Sep 23, 2026 17 4 2 Share Around 2021, well before ChatGPT launched, Vincent Hellendoorn suggested I try GPT-3 on the reading quizzes in my course. It produced convincing answers passing our rubric without actually seeing the assigned paper. At the time, I changed nothing. Five years later, AI agents could do all my assignments and I have redesigned most assessments in that course, even though what I want students to learn has barely changed. The strategy is always the same: No longer test understanding with anything that is done at home and instead focus on interactions with a TA, on exam, and on a video demo. Some of these changes violate evidence-based best pedagogy practices, and I made them anyway. For the last couple of years, I have mostly taught the course Machine Learning in Production , an upper-level course on building production-ready software around ML models with a heavy focus on MLOps, usually with 100 to 170 students. These days one common question when talking to other educators is how we have changed teaching in the age of generative AI and coding agents, so let me outline what we did. We have shifted covered topics with changing AI innovations and tools, but I barely touched the overall learning goals . I am fortunate that this is not an intro course and that the learning goals are not about writing code or using specific tools; they are about engineering tradeoffs, anticipating and mitigating risks, and teamwork. I think these are skills still worth acquiring, even if some can be simulated and offloaded to a model. (Revising an intro course or a traditional software engineering course likely would shift learning goals much more.) Also possibly important: We give students permission to use AI in all settings, in any form, without attribution, except for written and oral exams. We even encourage the use of AI tools in many places. I do not think policing AI is feasible even if we wanted, and more importantly I do think that students need to learn responsible use of these technologies anyway. AI is forcing me to abandon evidence-based best practices Let’s start with this point upfront, since it is more important than what we actually do: Unfortunately, AI is actively undermining several evidence-based teaching practices (e.g., see How Learning Works and The ABCs of How We Learn ). For example, the evidence favors frequent low-stakes assessments with feedback (e.g., homework, quizzes) over few high-stakes ones (e.g., exams) – but AI is undermining practice in low-stakes settings and pushing us more toward exams. Similarly, I always provided a safety net where students can make mistakes and resubmit a limited number of assignments to regain lost points (a core recommendation of specifications grading and grading for equity to focus on learning outcomes, not the process), but we felt that this process was abused with AI: first submit a generated assignment solution without thinking and only look at the issues raised in grading for a resubmission (the typical story of externalizing the cost of AI use). In response, we have since taxed resubmissions with a 10% penalty. Also in-class interactions allow engaging with materials in an early low-stakes setting, but with AI I have seen many student groups offload the discussion questions to a model. Pen-and-paper submissions could fix this, but aside from a higher grading workload, it would also raise stress for students, take away from the low-stakes environment, and delay feedback. In general, this is a balancing act and I tend to err on the side of keeping low-stakes repeated interactions even though it can be abused. Yes, some students will get through the class without much deep learning, but it provides a better environment for those students who want to learn. I don’t want to get back to the model I’ve experienced during my own studies in Germany with mostly optional homework and a single exam at the end of the semester that was responsible for 100% of the grade in the course. This was nice for students who were self-motivated and good at learning for an exam (like me, I guess), but had failure and drop-out rates of 50 to 80%. Thanks for reading! Subscribe for free to receive new posts and support my work. Subscribe Written reflections → 15-minute conversations Now for actual changes in the course: I have given up all parts of assignments that required a written text answer. I still ask for reports that describe a solution and link to the relevant code fragments, but that’s just for navigating their solution and I’m fine with receiving AI generated documents for that. In contrast, reflection documents, like “What were challenging parts?”, “How would you improve teamwork?” or in reading quizzes “For scenario X, identify one plausible data quality problem you might expect that relates to one of the four data cascades discussed in the paper …” have become pointless and can be entirely delegated. Short of hiding the evaluation rubric, I can see no way of stating what I expect in a good answer that cannot be completely offloaded to an LLM. For written reflections, which I used to have as a part of pretty much every assignment, I now shifted to in-person interactions with a TA. After every assignment, each student needs to schedule a 15-minute meeting with a TA to answer a couple of questions in a live conversation ( apparently Stanford cs221 is evaluating the same kind of approach in a controlled experiment this semester). I still share the reflection prompts in the assignment as examples of the kind of questions we ask. Students can still generate an initial answer with an LLM, but they may need to memorize parts of it, and we try to challenge them with follow up questions. The check-in meetings are part of the assignment and currently worth 20% of the assignment points, graded pass/fail. Students can try again if they fail and I encourage my TAs to have fairly high standards – we usually fail quite a few students on their first attempt. There are drawbacks to this design, but overall I am happy with the tradeoffs: Penalty-free retries reduce fairness concerns about TA grading of oral interactions; a more strict TA costs a student time, but not points. Oral check-ins demand more from students with anxiety, but so do written exams, and formal disability accommodations can provide a path in both cases. In fact, professional communication about technical work is a learning goal and oral check-ins train this more than written reflections. Regarding scale: We run the course at a 20:1 student-TA ratio with about 10h of work per TAs per week (fortunate, I know), so the check-ins amount to roughly 300 minutes per TA every two weeks, which is workable. For reading quizzes, I just gave up. I did not think doing in-person pen-and-paper quizzes in class would be worth the stress and the needless memorization work that those would be causing. I actually kept online reading quizzes around for a long time just to signal that I wanted students to look at the paper, fully understanding that most would just ask an LLM. These days, I still assign readings, but only half as many and without any points attached. Instead, I try to integrate lessons from the readings into in-class discussions. Still most students do not do the readings and just ask an LLM when we get to that point in the class (so nothing changed on that front), but those that do might get more out of it. Minor note: We observed that some students used AI during live discussions over Zoom (e.g. Cluely) and we will likely only offer in-person checkins in the future in response. For coding tasks: code + videos + in-person knowledge checks We have weekly labs that are low-stakes small tasks to explore new tools (e.g., Kafka, Grafana, Docker, Weights and Biases). These tasks are necessarily scoped small and need to provide some help to students starting out –