메뉴
BL
MIT Tech Review 9일 전

채용 과정에서 AI가 사람보다 편향될 가능성이 더 높다

IMP
9/10
핵심 요약

최근 연구에 따르면, 채용 과정에서 ChatGPT, Claude, Gemini와 같은 대형 언어 모델(LLM)이 인간보다 훨씬 더 강력한 고정관념과 편향을 형성하는 것으로 나타났습니다. 특히 최신 추론 모델일수록 제한된 경험을 바탕으로 성급하게 일반화하는 경향이 강해집니다. 이는 사용자의 정보를 기억하고 맞춤화하는 최신 AI 에이전트 기술이 의도치 않은 차별과 편향을 야기할 수 있음을 시사하므로 매우 중요합니다.

번역된 본문

경영 요약: 다음에 직장에 지원할 때, 어떤 인사 담당자도 보기 전에 AI가 먼저 귀하의 이력서를 검토할 수 있습니다. 하지만 AI가 귀하를 공정하게 평가할지 의문을 제기할 만한 충분한 이유가 있습니다. 연구자들은 이미 대형 언어 모델(LLM)이 학습 데이터에서 인간의 편향을 물려받는다는 사실을 알고 있습니다. 새로운 연구에 따르면 LLM은 경험을 통해 자체적인 편향을 형성할 수도 있으며, 구직자에 대한 고정관념은 사람보다 더 강한 것으로 나타났습니다. AI 기업들이 사용자에 대한 아주 사소한 세부 사항까지 기억하는 에이전트 모델을 구축하기 위해 경쟁하면서, 이들은 편향을 형성할 수 있는 '학습 자료'를 AI에게 쥐여주고 있는 셈이 될 수 있습니다.

프린스턴 대학교와 시카고 대학교의 연구진은 인간이 어떻게 고정관념을 형성하는지 탐구했던 심리학 연구를 바탕으로, ChatGPT, Claude, Gemini 등 LLM을 대상으로 시뮬레이션된 채용 게임을 진행했습니다. 각 모델은 가상의 도시 시장으로부터 컨설턴트로 고용된 것으로 가정한 뒤, 의사, 변호사, 보육 교사, 청소부 등 20개의 직무에 대한 채용을 돕라는 지시를 받았습니다. 후보자들은 Tufa, Aima, Reku, Weki라는 네 가상 민족 그룹에서 왔습니다. 매 라운드마다 새로운 직무 공석이 생겼고 각 그룹에서 한 명씩 총 네 명의 후보자가 주어졌습니다. 모델이 한 후보자를 채용한 후에는 그 후보자가 직무에서 성공했는지 여부를 알게 되었으며, 다음 라운드로 넘어갔습니다. 모델은 총 40라운드 동안 가능한 한 많은 성공적인 채용을 하라는 지시를 받았습니다. 모델들은 모르는 사이에, 모든 후보자가 모든 직무에서 동일한 성공 확률을 가지고 있었습니다.

모델들은 채용 결과에 대한 초기 관찰을 바탕으로 다른 그룹의 후보자들을 다른 직무로 신속하게 분리하기 시작했습니다. 예를 들어, 모델이 Aima 출신 후보자가 의사(높은 수준의 온기와 능력이 필요하다고 간주되는 직업)로서 실패했다는 피드백을 받으면, 모든 Aima 출신을 의사로 채용하는 것을 피했습니다. 대신 모델은 의사보다 온기와 능력이 덜 필요하다고 분류한 청소부로 Aima 출신을 채용하기 시작했습니다. 이 모델들은 원래 연구의 인간 참가자들보다 인구 통계 그룹별로 사람을 고정관념으로 판단할 가능성이 훨씬 더 높았습니다. 2점은 각 그룹이 완전히 자체적인 직무 틈새에만 국한됨을 의미하는 연구의 분리 척도에서, 인간 참가자들은 0.84점을 받았습니다. 반면 모델들은 약 65% 더 높은 점수를 받았으며, OpenAI의 추론 모델인 o3는 1.83점을 받아 가능한 최대 점수에 가까웠습니다.

이는 LLM이 "제한된 데이터로부터 일반화를 만들어내는 데 매우 열광적이기 때문"이라고 7월 서울에서 열린 ICML(국제 기계학습 회의)에서 논문을 발표한 이 연구의 공동 저자이자 프린스턴 대학교 박사과정 학생인 라이언 류(Ryan Liu)는 말합니다. "이것은 문자 그대로 그들이 최적화된 목적의 많은 부분을 차지합니다." 모든 의사 결정권자는 인간이든 기계든 이전에 효과가 있었던 방식을 고수하는 것과 더 나은 결과를 가져올 수 있는 새로운 방식을 시도하는 것 사이에서 트레이드오프에 직면하는데, 이는 심리학자들이 '탐험-활용 딜레마(exploration-exploitation dilemma)'라고 부르는 현상입니다. 이는 새로운 식당을 시도할지 신뢰할 수 있는 단골 식당을 갈지 선택하는 것과 같습니다. LLM은 수학, 코딩, 과학 문제로 훈련되기 때문에(이러한 작업들은 소수의 예만으로 패턴을 일반화하는 것을 보상합니다) 너무 일찍 직관에 정착할 수 있습니다. 논리 퍼즐을 푸는 데 도움이 되는 동일한 본능이 LLM을 고정관념을 갖도록 만들기도 합니다. 실험에서 OpenAI의 o3 및 DeepSeek의 R1과 같이 추론 능력이 더 높은 최신 모델은 훨씬 더 강력한 편향을 보여주었습니다. LLM이 사회적 상황에서 성급하게 일반화하려 할 때 "문제가 발생하는 경향이 있다"고 류는 말합니다. OpenAI와 Anthropic은 논평 요청에 응답하지 않았습니다.

이 연구에 참여하지 않은 코넬 대학교의 컴퓨터 과학자인 안젤리나 왕(Angelina Wang)에 따르면, 이제 챗봇이 향상된 기억력 및 개인화 기능을 얻고 있다는 점에서 이번 연구 결과는 특히 시의적절하다고 말합니다. 챗봇이 이전 대화 기록을 활용할 때 과거에 경험했던 것과 동일한 종류의 행동에 "과도하게 의존"하여 편향을 형성할 수 있다고 그녀는 덧붙입니다. 사용자가 챗봇에게 자신이 말한 내용을 기억하기를 원하기 때문에 단순히 챗봇이 덜 기억하게 만드는 것은 해결책이 아닙니다. 왕은 "우리는 여전히 너무 많지도 너무 적지도 않은 적절한 수준을 파악하기 위해 노력하고 있다"고 말합니다. 모델에게 공정하라고 지시하는 것만으로는 모델의 행동을 바꾸지 못했

원문 보기
원문 보기 (영어)
EXECUTIVE SUMMARY The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience—and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases. Researchers at Princeton University and the University of Chicago ran LLMs, including ChatGPT, Claude, and Gemini, through a simulated hiring game, adapted from a psychology study that explored how humans can form stereotypes. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs, including doctors, lawyers, child-care aides, and janitors. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each round, there was a new job opening and four candidates, one from each group. After the model hired a candidate, it learned whether they succeeded at their job and moved onto the next round. The model was told to make as many successful hires as possible over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed at every job. The models quickly started segregating candidates from different groups into different jobs on the basis of early observations of hiring outcomes. For example, when a model was told an Aima had failed as a doctor, a job considered to require high levels of warmth and competence, it veered away from hiring all Aimas as doctors. Instead, it started hiring Aimas as janitors, which the model classified as being less warm and competent than doctors. The models were even more likely to stereotype people by demographic group than the human participants in the original study. On the study’s segregation scale, where 2 means every group has been completely confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with OpenAI’s reasoning model o3 scoring 1.83, close to the maximum possible. That’s because LLMs “really are eager to create generalizations from limited data,” says Ryan Liu, a PhD student at Princeton University and a coauthor of the study, which was published in a paper at ICML in Seoul in July. “That’s literally a lot of what they’re optimized for.” Every decision-maker, human or machine, faces a trade-off between sticking with what worked before and trying something new that might work better—a phenomenon psychologists call the “exploration-exploitation dilemma.” It’s like choosing between a new restaurant and your reliable favorite. Because LLMs are trained on math, coding, and science problems—tasks that reward generalizing from just a few examples—they can settle on a hunch too early. And the same instinct that helps LLMs crack logic puzzles also makes them quick to stereotype. In the experiment, newer models with higher reasoning capabilities, such as OpenAI’s o3 and DeepSeek’s R1, showed even stronger biases. When LLMs rush to generalize in social settings, “that’s when things tend to go wrong,” says Liu. OpenAI and Anthropic did not respond to requests for comment. The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang, a computer scientist at Cornell University who did not work on the study. When a chatbot draws on its previous conversation history, it can “over-index on the same kinds of behaviors it’s experienced before” and form biases, she says. Simply having chatbots remember less isn’t a fix, though, because users want chatbots to remember what they say. “We still are trying to figure out just the right amount that isn’t too much or too little,” says Wang. Telling the model to be fair didn’t change its behavior much. “Either it can’t put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires,” says Liu. But promising the models an additional bonus for diverse hiring made them far less biased. The trick, then, is to design goals that “incorporate desirable social values in order to make the large language model act in socially desirable ways,” says Liu. The models also became less biased when they were told more personal information about individuals. In another experiment in the same study, the researchers asked the models to resettle members of different ethnic groups in cities across Canada. When the models were told personal information relevant to the ability to adapt to a new city, such as age and education, they were less likely to segregate people by their ethnicity. But when they were given irrelevant information, such as hair color and tattoo shape, the models largely fell back to sorting people by their ethnicity again. To what extent AI systems will stereotype job applicants in the real world is still an open question. While the models in the experiment immediately learned whether they’d made successful hires, a model screening résumés in the real world doesn’t get an instant report card. Companies can take a long time to find out whether a new hire is any good. But when feedback does trickle in, a model could still read too much into those results when making future hires. As companies increasingly deploy LLMs to screen résumés and even conduct interviews , the finding that models can form biases from their hiring experience “is a really serious implication that they should grapple with,” says Wang. As LLMs learn from experience to make decisions about who gets hired, who gets a loan, or who gets parole, the biases we should worry about may include ones no human ever taught them. “These novel biases—they’re sort of ever present,” says Liu. Deep Dive Artificial intelligence A startup claims it broke through a bottleneck that’s holding back LLMs Subquadratic has now shared more details about its new model. But some are still skeptical. By Will Douglas Heaven archive page A reality check on the AI jobs hysteria What do the numbers really say about the impact of artificial intelligence on the labor market? The answer might surprise you. By David Rotman archive page Anthropic’s Code with Claude showed off coding’s future—whether you like it or not As tools like Claude Code get better, more and more developers are happy to hand off coding tasks to them. The way software gets built has changed for good. By Will Douglas Heaven archive page Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. By Will Douglas Heaven archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.