메뉴
BL
The Decoder • 6일 전

현실적인 실수를 만드는 시뮬레이션 학생, AI 튜터 학습 가속화

IMP
6/10
핵심 요약

마이크로소프트와 일리노이 대학교가 개발한 StudentSim은 적은 데이터만으로 개별 학생의 디지털 복제본을 만들어, 실제 학생이 필요한 AI 튜터 훈련의 비용과 시간 문제를 해결합니다. 이 시스템은 학생 특유의 실수를 재현하면서도 튜터의 힌트를 따르는 능력을 갖춰, 체스·영어·수학에서 GPT-5.4보다 우수한 성능을 보였습니다.

번역된 본문

현실적인 실수를 만드는 시뮬레이션 학생, AI 튜터 학습 가속화

마이크로소프트와 일리노이 대학교가 제안한 새로운 시스템은 제한된 데이터만으로 개별 학생의 현실적인 복제본을 만들어냅니다. 이 복제본은 실제 학생에게 피드백을 받기엔 비용이 너무 크고 시간이 오래 걸리는 경우 빠른 피드백을 제공함으로써 연구자들이 AI 튜터를 개선하는 데 도움을 줍니다.

AI 튜터는 각 학생의 강점과 약점에 적응할 때 가장 잘 작동합니다. 하지만 어떤 지도 방식이 각 학생에게 효과적인지 알아내는 것은 실제 학습자가 필요하기 때문에 시간과 비용이 많이 듭니다. 크고 다양한 학생 집단으로 AI 튜터를 훈련하는 것은 "비용과 시간 면에서 감당하기 어렵다"고 저자들은 논문에서 밝혔습니다. 그 결과 이러한 튜터의 개선은 AI 모델 자체의 발전 속도에 뒤처져 왔습니다.

연구진은 학생의 디지털 복제본을 활용해 그 자리에서 빠른 피드백을 제공할 것을 제안했습니다. StudentSim이라 불리는 이 시스템은 해당 학생의 기록이 매우 적은 경우에도 각 학생별로 별도의 복제본을 만듭니다.

학생 복제본은 실수를 하고 지도를 통해 배워야 한다

연구진에 따르면 기존 접근 방식은 필요한 두 가지 능력 중 하나만 다룹니다. 일부 모델은 실제 학생 데이터로부터 학습해 학생의 행동을 신뢰할 수 있게 재현하지만, 튜터의 설명을 활용하지 못합니다. 다른 방식은 학생 역할을 하도록 프롬프트를 받은 언어 모델입니다. 이 모델은 튜터의 힌트를 잘 따르지만, 모방해야 할 학생의 수준과 맞지 않습니다.

StudentSim은 이 두 가지 능력을 모두 측정 가능한 목표로 전환합니다. 복제본이 학생의 답변(전형적인 실수 포함)을 얼마나 가깝게 재현하는지, 그리고 튜터가 도움을 준 후 답변을 얼마나 잘 수정하는지를 측정합니다. 튜터 훈련에는 현실적인 출발점과 지시에 반응하는 시뮬레이션 학생이 모두 필요합니다.

2단계 훈련으로 제한된 학생 데이터 활용 가능

연구진의 가장 큰 장애물은 데이터 부족입니다. 영어 작문 데이터셋에서 중위권 학생은 겨우 3편의 에세이를 작성했고, 3분의 2 이상이 5편 이하를 작성했습니다. 이렇게 적은 예시로 복제본을 직접 훈련하면 모델이 해당 예시에 과적합(overfit)되어 실패한다고 연구진은 설명합니다.

대신 StudentSim은 두 단계로 훈련합니다. 첫째, 기반 모델이 특정 과목의 모든 학생 데이터를 통합해 학습합니다. 이를 통해 흔한 실수와 학생이 튜터의 힌트 후 답변을 수정하는 방식을 배웁니다. 그다음 연구진은 해당 학생의 소수 기록을 활용해 이 모델을 개별 학생에 맞게 조정합니다. 모든 과목에서 이 시스템은 알리바바의 Qwen3-4B-Instruct 언어 모델을 기반으로 사용합니다.

체스, 영어, 수학에서 GPT-5.4 능가

연구진은 체스, 외국어로서의 영어, 수학 분야의 학생 60명을 대상으로 이 방법을 테스트했습니다. 실제 학습자의 기록을 담은 공개 데이터셋을 사용했습니다.

StudentSim은 학생 역할 프롬프트를 받은 더 큰 GPT-5.4 언어 모델을 세 과목 모두에서 능가했습니다. 체스에서 StudentSim은 플레이어의 다음 수를 약 두 배 더 정확히 예측했고 수정 지시를 거의 항상 따랐습니다. GPT-5.4와 전용 체스 모델은 뒤처졌습니다.

연구진에 따르면 각 기존 방법은 서로 다른 약점이 있습니다. GPT-5.4는 힌트를 따르지만 특정 학생의 실수를 재현하지 못합니다. 체스 모델은 플레이어의 행동과 일치하지만 언어적 힌트를 이해하지 못하고 무시합니다. 한 포지션에서 실제 플레이어 세 명이 서로 다른 수를 두었을 때, StudentSim은 각 플레이어의 선택을 재현했지만 체스 모델은 세 명 모두에게 같은 최선수를 예측했습니다. GPT-5.4는 세 경우 모두 틀렸습니다.

시뮬레이션 학생으로 훈련하면 체스 튜터 개선

또 다른 개념 증명으로, 연구진은 학생 복제본을 활용해 체스 튜터를 개선했습니다. 프로 체스 선수들이 세 가지 버전(이 훈련 없이 만든 것, GPT-5.4를 학생으로 훈련한 것, StudentSim으로 훈련한 것)을 평가한 결과, StudentSim으로 훈련된 튜터가 세 가지 평가 항목 모두에서 최고 점수를 받았습니다.

원문 보기
원문 보기 (영어)
Simulated students that make realistic mistakes help AI tutors learn faster Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 20, 2026 Nano Banana Pro prompted by THE DECODER A new system from Microsoft and the University of Illinois builds realistic replicas of individual students from limited data. These replicas provide rapid feedback when getting it from real students would be too expensive and slow, helping researchers improve AI tutors. AI tutors work best when they adapt to each student's strengths and weaknesses. But finding out which guidance works for each student takes time and money because it requires real learners. Training an AI tutor with a large, diverse group of students is "prohibitively expensive and time-consuming," the authors write in their paper . As a result, improvements to these tutors have lagged behind advances in AI models. The researchers propose using digital replicas of students to provide quick feedback in their place. Their system, called StudentSim, builds a separate replica for each student, even when very few records of that person's work are available. Student replicas need to make mistakes and learn from guidance Existing approaches handle only one of two necessary skills, according to the researchers. Some models learn from real student data and reliably reproduce a student's behavior, but they can't use a tutor's explanations. Others are language models prompted to act as students. They readily follow the tutor's hints but fail to match the abilities of the student they're supposed to mimic. StudentSim turns both skills into measurable goals. It measures how closely a replica matches a student's answers, including typical mistakes, and how readily it revises an answer after the tutor helps. Tutor training needs both a realistic starting point and a simulated student that responds to instruction. Training in two stages makes limited student data usable The researchers' biggest obstacle is a lack of data. In the English writing dataset, the median student has written just three essays, and more than two-thirds have written five or fewer. Training a replica directly on so few examples fails, the researchers say, because the model overfits to those examples. StudentSim instead trains in two stages. First, a base model learns from the pooled data of all students in a subject. It learns common mistakes and how students revise their answers after a tutor's hint. The researchers then tailor that model to an individual student using the few records available for that person. Across all subjects, the system uses Alibaba's Qwen3-4B-Instruct language model as its base. StudentSim outperforms GPT-5.4 in chess, English, and math The researchers tested the method on 60 students across chess, English as a foreign language, and math. They used public datasets containing records from real learners. StudentSim outperforms the larger GPT-5.4 language model in all three subjects when GPT-5.4 is prompted to act as a student. In chess, StudentSim correctly predicts a player's next move about twice as often and almost always follows corrective guidance. GPT-5.4 and specialized chess models fall behind. Each existing method has a different weakness, the researchers say. GPT-5.4 follows hints but doesn't reproduce a particular student's mistakes. Chess models match a player's behavior but can't understand verbal hints and ignore them. In one position, three real players chose three different moves. StudentSim reproduced each player's choice, while a chess model predicted the same most likely move for all three. GPT-5.4 got all three wrong. Training with a simulated student improves a chess tutor As another proof of concept, the researchers used a student replica to improve a chess tutor. Professional chess players evaluated three versions, one without this training, one trained with GPT-5.4 as the student, and one trained with StudentSim. The StudentSim-trained tutor scored highest on all three measures. It made the fewest serious factual errors and received the highest scores for explanation quality and adaptation to the individual student. In this case, the student preferred questions that guided them toward a solution rather than direct instructions. The tutor trained with GPT-5.4 scored worse on factual accuracy than the tutor that received no extra training. The researchers say this is only a proof of concept, not a claim to have built the best tutor. Chess works as a test case because an engine can objectively judge whether a move is good in any given position. Essay writing and open-ended math are harder because they lack reliable scoring functions for free-form answers. Next, the team wants to model how students acquire, retain, and forget knowledge over many practice sessions. The code is available on GitHub . Researchers used AI agents to replicate about 1,000 real people in 2024, based on two-hour interviews with each participant. A separate study showed how error-prone these replicas can be. Nine open language models tasked with mimicking user behavior on X, Bluesky, and Reddit became less accurate in their content as they sounded more human . Microsoft is also testing AI tutors with real students. In a pilot project in Nigeria , students worked with Copilot twice a week for six weeks. Their test-score gains were equivalent to nearly two additional years of learning. OpenAI and Google offer their own learning modes through Study Mode and Guided Learning . These rely on system instructions and models fine-tuned for teaching, but neither maintains a model of the individual learner. Without that adaptation, AI assistance can hurt performance. Studies show that users perform worse after brief AI assistance than people who worked on their own from the start. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->