Interviewing

How Do You Run a Structured AI-Use Interview That Compares Candidates?

An AI-use interview becomes comparable through three commitments, and the question list is not one of them. Every candidate gets the same questions in the same order, every answer is rated on one scale whose levels are written as observable behavior before the first interview, and the follow-ups are fixed too, since an improvised second question is where comparability leaks. Write the levels in the role's own vocabulary: a marketer's check and an analyst's are different acts. None of it shows whether the checking would happen on the job.

The takeWriting the anchors is the expensive part, and it is expensive because it makes a team say out loud what checking a claim looks like in their own job. Most panels have never had that argument. They discover, an hour in, that two of them would rate the same answer at opposite ends of the scale, and the disagreement was never about a candidate at all. That hour is the round's real product. The ratings are downstream of it. Nine times out of ten, a rubric session that goes smoothly and finishes early has produced nothing specific enough to disagree with, and the interview after it will measure charm.

Where Olive fits

Open a role and see what the work shows

A rubric makes two answers comparable, but both are still accounts of the work rather than the work. Olive holds the case fixed instead: everyone invited to a role meets the same authored assignment with an AI assistant available, and a human reviewer writes six findings, each attached to the moment in the session it rests on.

Rank your shortlist

What makes an AI-use interview comparable?

Comparability comes from three commitments, and the question list is not among them. Every candidate gets the same questions in the same order, every answer is rated on one common scale, and the interviewers agree beforehand on what an acceptable answer sounds like 2. Drop any one of those and you are collecting impressions. The structure is what lets you subtract one candidate's answers from another's.

The payoff is measured. A 2023 review of selection-method validity put structured interviews at the top of its list with a mean operational validity of .42, against .29 for unstructured ones 1. That gap is not about interviewer skill. It is what happens when two conversations are made to differ only in the candidate.

AI-use rounds drift out of structure faster than most, because the subject is genuinely interesting to the interviewer. Someone mentions a tool nobody has tried, the round becomes a conversation about tooling, and the next candidate gets a different interview. Federal guidance is blunt about where that ends: unstructured interviews show low rating consistency between interviewers, low to moderate validity, and greater exposure to legal challenge 2.

So the deliverable here is not a question bank. Good AI questions already exist. Twelve of them, with the answers that should worry you, are a fine place to start. The deliverable is the scale you rate them on, and the discipline of writing it before anyone sits down.

Ask these five questions, in the same order

Five fits inside a 45-minute round with room for one real follow-up each, and the follow-up is where an answer stops being rehearsed. Every question is about one recent task the candidate finished with an assistant, and every candidate answers about that same single task of their own. Name the task at the top of the round and keep the conversation on it.

Each question rates exactly one behavior. Say which, in the interviewer's guide, because a question that rates two behaviors gets one rating and loses both.

1. Framing. "What was the first thing you typed to the assistant on that task (the actual first message, not the tidy version)?" Fixed follow-up: "What would have made your answer wrong, and did you tell it that?" 2. Evidence. "Which specific claim in what it gave you did you check, and where did you go to check it?" Fixed follow-up: "What did that source turn out to say?" 3. Delegation. "What did you keep and do yourself, when the assistant could have done it?" Fixed follow-up: "Why that piece and not another one?" 4. Rejection. "What did you throw out, and on what grounds?" Fixed follow-up: "What did you put in its place?" 5. Effect. "What changed because of the check: a number, a recommendation, a caveat you added?" Fixed follow-up: "Who had already seen the version before it changed?"

The follow-ups are fixed too, and that is the part panels get wrong. Improvised second questions are where structure leaks: two candidates answer the same prompt and then get different pressure, and the ratings stop meaning the same thing. Ask the written follow-up every time, rate, and only then probe freely. A fixed follow-up is what exposes understanding, because a prepared answer has depth of exactly one.

Notice what is missing. Nothing here asks which tools the candidate uses or how often. A tool list rates nothing, and volume of AI use is not the skill. A candidate who judged the model was the wrong instrument and did the step by hand has demonstrated the thing you are testing.

Write the rubric before you write the questions

Write the levels first, because a question you cannot rate is a conversation. Federal guidance is to build at least three proficiency levels and label them, aiming for five to seven 2. Three is enough for this: not demonstrated, partly demonstrated, demonstrated. Each level is a sentence describing what the interviewer would actually hear, and every rating gets the candidate's own words recorded beside it.

Here is the full scale for question two, evidence. Copy the shape rather than the wording.

RatingWhat the interviewer heard
DemonstratedNames a specific claim, says where they went to check it, and says what the source turned out to say, including the times it agreed.
Partly demonstratedNames a check but not a claim ("I read it through"), or names both but cannot say what the source said.
Not demonstratedDescribes reviewing the output as a whole, or says it looked right, or the trail ends at the assistant.

Three habits make a scale like that hold. Record the sentence, not just the level: a rating with no quote cannot be defended in a debrief, and a debrief where nobody can quote anything reverts to who was most likeable. Rate immediately after each answer, before the next question, so a strong opening does not color the rest. And where a panel is used, rate independently first and compare afterwards.

Those habits exist for named failure modes. The same guidance lists six rating errors structure is there to hold off (rater bias, halo, central tendency, leniency, strictness, and rating people who resemble you), and says the countermeasure is comparing what the candidate did against the behaviors written into the levels 2. That works only if the behaviors are written down.

The legal argument runs the same direction. An interview used to make a hiring decision is a selection procedure, and the EEOC's guidance holds that a procedure screening out a protected group must be shown job-related and consistent with business necessity 3. A rubric written in advance, plus a recorded reason beside each rating, is what you have to show. "She seemed more AI-native" is not. Before a scale like this decides anything, it is worth reading on how reliable an AI-collaboration rubric actually is.

Why do the anchors change from role to role?

Because the behavior they describe is different work. "Demanded a source" means opening the filing for a financial analyst and opening the study behind a statistic for a marketer. Different acts, different tells, different ways of going wrong. Dimension names travel across roles. Anchors do not. Write a role's anchor by asking what the check physically looks like on a Tuesday in that job.

Take evidence sourcing, the second question, and read the top level in three roles:

RoleWhat "demonstrated" sounds like
MarketingWent to the study behind a quoted market figure, found its sample and its year, and either kept the number with the qualifier attached or cut it.
Financial analysisWent to the filing or the source workbook instead of the summary, and can say which line the figure came off and what it reconciles to.
Software engineeringRan the generated code against a case the assistant did not write, or read the library source instead of the assistant's description of it.

Same skill, three different Tuesdays. Hand a marketer the engineer's anchor and they rate low on something they do well, which is the quiet way a rubric stops being fair while still looking rigorous.

Building the role version costs one conversation with two people who currently do the job: ask each for the last time an assistant told them something confident and wrong, and what they did in the next ten minutes. Their answers are the anchor text. Federal guidance puts a job analysis first for exactly this reason: interview questions have to reflect competencies derived from the work 2, and an AI-use round is not exempt because the tool is new.

Which behaviors deserve anchors is not a matter of taste either. In a survey of 319 knowledge workers describing 936 first-hand examples of AI-assisted work, the thinking that survived shifted toward verifying answers, integrating them into the task, and stewarding the work rather than executing it 5. Those are the acts to write levels for. Prompt vocabulary is not one of them. See how Olive measures this.

What can a structured interview still not tell you?

Not whether any of it would actually happen on the job. Every answer is a report about work rather than the work, and self-reports about AI use are measurably unreliable: in a randomized trial, sixteen experienced developers took 19% longer to complete issues with early-2025 AI tools and still believed afterwards that the tools had sped them up by about 20% 4. Structure fixes comparability. It does not close the gap between telling and doing.

Close it cheaply in one of two ways. Add fifteen minutes of live work at the end of the same round: unfamiliar material, assistant open, one act rated instead of five accounts. Or run a work sample on real occupational material, which costs more and is the only format that leaves a record you can re-read. The trade-off sits side by side in work sample versus structured interview.

Structure also does not stop an answer from being prepared. Candidates rehearse with the same assistants you are asking about, and the polished STAR answer now arrives at scale. What resists preparation is the fixed follow-up operating on the answer just given: which file, which page, which number moved. A model can write someone a convincing account of a check. It cannot supply the detail of a check that never happened.

One last thing, in the invitation rather than the room: say the round covers how they work with AI, and that using it is expected. An unstated rule gets guessed at, and the guessing tracks how much interview coaching someone has had, which is the opposite of comparable.

See how it works

Common questions

How many questions fit in a structured AI-use round?

Five in 45 minutes, each with one fixed follow-up. The follow-up takes longer than the first answer and is the part doing the work. More questions asked at checklist pace produce more surface and no more signal. If the round is 30 minutes, cut to three and keep the follow-ups rather than keeping five and dropping them. Whatever the number, it has to be the same number for every candidate, or the ratings are not comparable.

Do two interviewers need to rate the same interview?

Not always, but ratings mean more when they do. Two people rating independently and comparing afterwards tell you whether the anchors are legible; when they diverge, the level description is usually vague rather than the candidate ambiguous. Fix the wording, not the disagreement. If only one interviewer is available, have them write the candidate's exact words beside each rating so somebody else can audit the call later.

Can one rubric cover every role?

The dimensions can. The anchors cannot. "Demanded evidence" describes opening a filing for an analyst, a study for a marketer, and a test case for an engineer. Three different observable acts. Keep one rubric skeleton across the company so the dimensions mean the same thing everywhere, then rewrite the level descriptions per job family with two people who do that job. Reusing one role's anchors on another rates familiarity with the wrong work.

What if the candidate says they don't use AI?

That is an answer, not a disqualification. Ask what they decided not to use it for and why. A considered refusal is the delegation judgment your third question rates, and it can rate at the top. Then ask how they check a confident claim from a source that cannot show its work, because that skill predates the tools. Rate the reasoning rather than the usage. A rubric that rewards volume measures exposure.

Is a structured interview safer legally than an unstructured one?

It gives you something to show, which is the point. An interview used to make a hiring decision is a selection procedure, and under the EEOC's guidance on employment tests and selection procedures, one that screens out a protected group must be shown job-related and consistent with business necessity. A rubric written in advance, identical questions, and a recorded reason beside each rating are the artifacts that answer that months later. Jurisdictions differ, several now regulate automated employment decisions specifically, and none of this is legal advice.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology 16(3), 2023. doi.org Structured interviews top the paper's validity list at a mean operational validity of .42 (80% credibility interval .18 to .66); unstructured interviews are estimated at .29.
  2. 2. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Defines structure as the same questions in the same order, a common rating scale, and interviewer agreement on acceptable answers; unstructured interviews show low rating consistency and greater legal exposure. Also specifies at least three labeled proficiency levels (aim for five to seven), a job analysis as step one, and six rating errors countered by comparing behavior against the level anchors.
  3. 3. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A selection procedure that screens out a protected group must be shown job-related and consistent with business necessity.
  4. 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial of 16 experienced open-source developers: 19% longer to complete issues with early-2025 AI tools allowed, while participants believed afterwards that AI had sped them up by about 20%.
  5. 5. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ACM CHI Conference on Human Factors in Computing Systems (CHI '25), 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: the critical thinking that remains shifts toward information verification, response integration and task stewardship rather than execution.

5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.