AI Doesn’t Just Inherit Hiring Bias, It Invents New Ones

By Lee Flanagan

22nd Jul. 2026  |  Last Updated: 22nd Jul. 2026

Picture a fictional city where a mayor hires a recruiting consultant to fill twenty jobs, from doctors to janitors, drawn from four fictional ethnic groups called Tufa, Aima, Reku and Weki. Every candidate from every group has exactly the same chance of succeeding. Early on, one Aima candidate fails as a doctor, purely by chance. From that single data point, the consultant starts steering every other Aima applicant away from medicine and toward lower status work. No one told it to do this, and no historical prejudice was baked into its training. It built the pattern itself, from one unlucky outcome.

That consultant was not a person. It was a large language model, and the city was a hiring simulation built by researchers at Princeton University and the University of Chicago. The study, titled “Costly Exploration Produces Stereotypes With Dimensions of Warmth and Competence,” was published in the Journal of Experimental Psychology, reviewed by MIT Technology Review and covered by The Times of India. The finding that should worry any TA leader evaluating AI screening tools is not that AI can absorb bias from bad training data. It is that AI can manufacture brand new discrimination from nothing, as a side effect of letting it decide, on its own, who advances.

That capacity to invent bias from thin air is, in our view, the strongest argument yet for keeping a human making the final call rather than automating the decision itself. The model was not trained badly. Letting any system decide autonomously, from limited data, is structurally how this kind of bias gets made.

The Study Was Designed to Remove Every Excuse

Researchers tested several leading models, including ChatGPT, Claude and Gemini, placing each in the role of a hiring consultant across forty rounds of decisions. The four applicant groups differed only in name, and every candidate carried identical underlying odds of success in every job. That design mattered: any stereotype the models formed came from the model’s own decision-making, not from real ability differences or historical inequity baked into a resume dataset.

That is the detail commentary on AI hiring keeps skipping past. Not whether a tool learned society’s prejudices from its data, but what happens when it is simply left to keep deciding who gets the next opportunity. The researchers found that it starts inventing categories of its own.

The Models Discriminated More Than the Humans They Were Tested Against

The original human version of this experiment produced an average segregation score of 0.84, a measure of how strongly participants sorted different groups into different occupations. The language models produced segregation scores roughly 65% higher than that human benchmark, and OpenAI’s reasoning model o3 recorded a score of 1.83, close to the maximum possible on the researchers’ scale.

Ryan Liu, a Princeton PhD student and co-author of the study, explained why: “They really are eager to create generalisations from limited data. That’s literally a lot of what they’re optimised for.” That instinct makes a model good at spotting patterns in code or in a spreadsheet. Applied to people, it becomes a machine for turning a handful of unlucky outcomes into a permanent verdict on an entire group.

Fairness Instructions Failed

Newer reasoning models, including OpenAI’s o3, showed stronger stereotyping than earlier, less capable ones. More reasoning power did not produce fairer decisions. It produced faster, more confident ones, which is not the same thing. Instructing the models to be fair had little effect either. The models understood fairness as a concept and kept optimising for successful hires anyway.

That gap between understanding a principle and acting on it is the real design problem. A fairness instruction sits on top of a system already optimised for something else, and instructions do not change what a model is rewarded for. A vendor rarely shows what a system was rewarded for during training, which is why the decision still needs a human check built into the workflow.

What This Means for Every Screening Tool on Your Shortlist

The industry pitch for AI in hiring has always been speed and objectivity: process thousands of applications and find the pattern humans miss. This study undercuts that pitch. Our read is that the pattern a model finds is not always real, and sometimes it overreacts to a handful of chance outcomes and locks an entire group out of a role, with no human able to see it happening.

That mechanism does not need a forty-round laboratory simulation to operate. We think any tool that auto-screens or auto-ranks candidates at scale is exposed to the same kind of failure: it makes repeated calls from limited signal with no person checking the reasoning behind them. In our work with hiring teams, the tools that raise the most red flags quietly accept or reject candidates on thin, early signal, without human review.

This should change how any TA leader evaluates an AI screening or ranking product. An AI that autonomously decides who advances, from limited data, is not a feature to be impressed by. It is the exact failure mode this research documents.

Put one direct question to the vendor in front of you: does the tool show the reasoning behind a ranking, or does it hand over a score and nothing else? Removing the human from the decision does not remove the bias. It relocates it, and gives it a new, less visible source.

Original reporting: The Times of India.

Frequently asked questions

Did the study find AI is always worse than humans at avoiding bias in hiring?

In this specific controlled simulation, yes: the language models tested produced segregation scores roughly 65% higher than the human benchmark from the original psychology experiment, with OpenAI’s o3 model scoring close to the maximum on the researchers’ scale. The study tested several leading models in one simulated scenario, not every AI hiring tool on the market.

Why did the researchers use fictional ethnic groups instead of real demographic data?

Using four fictional groups with identical underlying odds of success meant any stereotype the models formed could not be blamed on real ability differences or on historical bias already present in a training dataset. It isolates bias the model generates through its own decision-making, separate from bias inherited from data.

Does telling an AI hiring tool to ‘be fair’ actually reduce bias?

Simple fairness instructions had little effect, even though the models understood the concept and kept optimising for successful hires anyway. That gap between understanding fairness and acting on it shows instructions alone do not change what a system is optimised to do.

Are newer, more advanced AI models safer to use in hiring decisions?

The study found the opposite: newer reasoning models, including OpenAI’s o3, showed stronger stereotyping than earlier, less sophisticated ones. Greater reasoning capability did not translate into fairer decision-making in this test.

What should TA leaders take from this when choosing an AI screening tool?

Treat any tool that autonomously accepts or rejects candidates from limited signal as a red flag rather than a feature. The safer approach keeps a human making the final call, using structured evidence the tool can show, rather than a score it cannot explain.