Washington’s AI Interview Pilot Will Test the Machine, Not the Decision

By Lee Flanagan

7th Sep. 2026  |  Last Updated: 7th Sep. 2026

The US federal government is about to hand hiring managers something new: a transcript of a job interview they did not conduct, produced by a virtual agent they did not train, for a decision they still have to make. CBS reported Monday, citing US officials, that the federal government will test AI-run interviews on the Tech Force, an eight-month-old initiative recruiting private-sector tech experts into two-year government posts, using a platform called CodeSignal, according to the Jerusalem Post.

The federal government employs 1.9 million workers. Running AI interviews across a workforce this large, if it works as planned, would show that AI can conduct interviews reliably. It would still prove nothing about whether the people receiving the output can make a good decision from it, and that gap is the part that matters.

What CodeSignal Actually Delivers

CodeSignal’s virtual agents will screen applications in early rounds and conduct interviews by phone, audio and text. Hiring managers will then receive transcribed recordings of the candidates. The reporting describes the whole workflow: agent conducts interview, human receives transcript. There is no mention of a scoring framework, a rubric or any standard for what a hiring manager should look for when that transcript lands in their inbox. CodeSignal is the vendor here, and its job in this pilot is to produce the interview. What happens after the transcript arrives is not CodeSignal’s problem to solve, and nothing in the reporting suggests it is trying to.

Human Oversight Is a Policy Line, Not a Process

OPM guidance encourages agencies to use AI in hiring with human oversight, especially for certain crucial personnel decisions. That phrase does real work in a policy document and almost none in the middle of a packed hiring day, when a manager is working through a stack of transcripts against the clock. Oversight only means something as a set of actions a person actually takes: what they read for, how they weigh it, what they compare it against, how their judgment differs from the next hiring manager’s on the same transcript. None of that is specified for the Tech Force pilot.

Scott Kupor, Trump’s nominee to direct OPM, has pushed to make the federal hiring process more efficient and wants agency heads to pull at least a third of new hires from early-career candidates as the workforce ages. Efficiency and oversight are not naturally compatible goals, and nothing in the reporting explains how this pilot resolves the tension between them.

Scale Makes This Everyone’s Preview

OPM’s administrative staff already have government-issued access to Copilot, Claude, ChatGPT and Gemini, and the office has used AI to draft job descriptions since earlier this year. The Tech Force pilot is the next step: AI moving from support tooling into the interview itself. Several officials told CBS that AI use in federal hiring is likely to grow.

A government of this size running the pilot is not a quiet trial. It is a live test case, and other large employers already using AI screening as their first layer will be watching what happens when the AI stops screening and starts interviewing. CBS also reported that the federal government has fallen behind those same private employers on AI interviewing, and the Tech Force pilot reads as an attempt to close that gap. Catching up on screening speed is a different achievement from closing the decision-quality gap this piece is about.

Raw Data Is Not a Decision

The coverage stops short of the real question: whether handing a hiring manager more raw, unstructured data makes a decision fairer or more accurate. It does not. It makes the decision more dependent on whoever happens to be reading that day, in what mood, against what unstated bar. In our work with hiring teams, unscored interview material does not get reviewed evenly. It gets skimmed by a manager under time pressure, and it gets remembered selectively by the time a decision finally gets made. A transcript is not a scorecard. It is just more evidence for someone who was never given a way to weigh it.

CBS’s reporting cannot answer this, because the pilot itself never defines it. OPM’s guidance names the goal, human oversight, without naming the mechanism. That is not a criticism of AI conducting interviews. Virtual agents running structured interviews at scale is a legitimate capability, and, in practice, a steadier one than a rotating cast of hurried human interviewers each asking whatever comes to mind. The problem sits entirely on the other side of the transcript, in the moment a hiring manager has to turn a document into a decision with no shared standard for what a good answer looks like.

The Test Nobody Is Measuring

If this pilot succeeds by the standard CBS’s reporting implies, meaning the interviews run and the hires get made, that tells you AI can conduct an interview reliably. It tells you nothing about whether the decisions on the other end are any better, fairer or more consistent than what untrained human interviewers already produced. The federal government has built the conditions to answer a genuinely useful question and, on the evidence available, built nothing to measure it. If your organization is running a version of this pilot without a defined standard for what human oversight means in practice, you are running the same untested experiment.

Original reporting: jpost.com.

Frequently asked questions

Where does CodeSignal’s responsibility end in the Tech Force pilot?

CodeSignal screens applications and conducts the interview itself by phone, audio or text, then sends hiring managers a transcribed recording. Its role stops at the transcript, and the reporting describes no scoring tool or rubric for what happens once a hiring manager receives it.

Does OPM’s guidance say what human oversight should look like in practice?

OPM guidance tells agencies to use AI in hiring with human oversight, especially for crucial personnel decisions, but it does not spell out what that oversight involves for the Tech Force pilot or any other program.

Why should employers outside government pay attention to this pilot?

The federal government employs 1.9 million workers, and CBS reports it has lagged private-sector employers on AI interviewing. Large employers already using AI to screen candidates have reason to watch what happens once AI moves into the interview and decision stage.

Will AI’s role in federal hiring grow beyond this pilot?

Several officials told CBS that AI use in federal hiring is likely to grow. OPM staff already use tools including Copilot, Claude, ChatGPT and Gemini, and the Tech Force pilot marks AI’s move from drafting and screening into running the interview itself.