When an AI screens your resume for a job, the risk of discrimination may go beyond biased training data. New research from Princeton University and the University of Chicago shows that large language models can form their own biases from experience, stereotyping job applicants even more than humans do. As companies race to deploy AI in recruitment, this finding raises serious concerns about fairness in automated hiring.
LLMs learn to discriminate from limited data
The study adapted a classic psychology experiment on stereotype formation. Researchers asked models including ChatGPT, Claude, and Gemini to act as consultants for a fictional mayor, hiring candidates for twenty jobs such as doctors, lawyers, child-care aides, and janitors. Candidates belonged to four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each round, the model picked a candidate and immediately learned whether they succeeded. Unbeknownst to the models, all candidates had equal success probability. The results showed that LLMs quickly segregated candidates by group based on early outcomes. If an Aima failed as a doctor, the model avoided all Aimas for that role, steering them toward less prestigious jobs like janitors.
Sponsored Protocol
On the study's segregation scale—where 2 means complete segregation—human participants scored 0.84. AI models scored about 65% higher, with OpenAI's o3 reasoning model reaching 1.83, near the maximum. Newer models with advanced reasoning capabilities exhibited the strongest biases.
The generalization instinct undermines neutrality
Ryan Liu, a Princeton PhD student and coauthor, explained that LLMs are optimized to generalize from few examples. Trained on math, coding, and science problems, they tend to commit to a hypothesis too early. Angelina Wang, a computer scientist at Cornell not involved in the study, noted that as chatbots gain persistent memory, they may over-index on past experiences, reinforcing biases. Simply telling the model to be fair had little effect, but offering a bonus for diverse hires dramatically reduced bias. Providing relevant personal information (age, education) also mitigated segregation, while irrelevant details (hair color) did not.
Sponsored Protocol
Real-world implications for AI recruitment
Whether these biases will manifest in real hiring depends on feedback loops. In the experiment, feedback was immediate; in practice, companies learn about a hire's success only after months. However, when feedback arrives, models could overinterpret it. Wang warns that companies using LLMs for resume screening must grapple with this implication. Liu concludes that as models learn from experience to decide who gets hired, loans, or parole, the biases we should worry about may include ones no human ever taught them.
Sponsored Protocol
For more on AI safety, see the article on Context Bombing blocking AI hacking agents. Also check the comparison of ChatGPT, Claude, and Gemini picking the best smartphone. Additional background on algorithmic bias is available on Wikipedia.
Source: https://www.technologyreview.com/2026/07/20/1140655/ai-biases-hiring-humans