Interviewing

What Interview Questions Show How a Candidate Works With AI?

Ask a candidate about one task they finished last month with an AI assistant, then work their own answer: what they typed first, what they refused, what they checked outside the chat, what changed because of it. Tool tours and philosophy questions get rehearsed answers. Ask every candidate the same set against a rubric written first, and don't score volume of AI use. Know the limit: an interview reaches what someone can say about a check, never the check itself, so pair it with a live extension or work sample.

The takeTwelve questions below, one question underneath: can you tell when the confident thing in front of you is wrong. That skill is old. Editors, auditors and second-year associates were always paid for it. What the assistant changed is how much of an ordinary working day now runs through it, and the rubric should have moved with that. A round spent admiring how naturally someone talks about AI is measuring exposure and calling it judgment. I'd expect the people who catch the model out to be the ones who sound least impressed by it.

Where Olive fits

Open a role and see what the work shows

These questions reach what a candidate can describe about checking a confident claim; they stop short of the act itself. Olive puts the act in front of them: an occupational assignment, an assistant willing to do all of it, and a human reviewer who writes six findings anchored to the moments they happened.

Rank your shortlist

What makes an AI question worth asking?

Anchor it to one piece of work the candidate actually shipped, and ask about acts rather than opinions. An opinion about AI is free and rehearsable; a decision made at 4pm on a Tuesday has details attached: what the assistant got wrong, what got cut, what was checked. Ask what happened, then ask how they knew.

The reason to ask about acts is that self-report on AI is measurably unreliable. In a randomized controlled trial, sixteen experienced open-source developers working in their own repositories took 19% longer to finish issues when allowed to use early-2025 AI tools, and afterwards still believed the tools had made them about 20% faster 1. They were not lying. Fluent output feels like progress, and the feeling survives the measurement.

That gap is what a good question aims at. In a survey of 319 knowledge workers describing 936 first-hand examples of AI-assisted work, higher confidence in the AI predicted less critical thinking, while higher confidence in their own ability to do the task predicted more, and the thinking that remained shifted toward verifying answers, integrating them, and stewarding the task rather than executing it 2. Verification, integration and stewardship are therefore the behaviors worth interviewing for. Prompt technique is not.

Two popular question shapes fail for the same reason. "How do you use AI in your workflow?" invites a tour of tools, and a list of tools tells you almost nothing about judgment. "Tell me about a time you used AI to solve a problem" invites a story built for exactly this round. Both can be answered well by someone who has never once checked a claim a model made.

Ask these twelve questions

Each one names a moment rather than a capability, and each carries a companion: the answer that should worry you. Run them as follow-ups inside a round you already have rather than as a new stage. Six asked properly, with one real follow-up each, beat all twelve asked as a checklist.

1. "Start with the very first message you typed into the assistant on that task. Not the polished version, the actual first message." Worry when the first message is the deliverable. "Write me the positioning brief" means the framing was handed over before it existed. 2. "What would have made your answer wrong, and did you say that to the model?" Worry when they can name a failure condition now but never named one then. Criteria invented afterwards are criticism, not framing. 3. "Which specific claim in that output did you check, and where did you check it?" Worry when the answer is "I checked everything" or "it all looked right." A real check has a location: a file, a page, a query, a person. 4. "What did the assistant tell you on that task that turned out to be wrong?" Worry when nothing comes. Anyone who has worked seriously with a model for a month has a story; no story means no scrutiny, or no real use. 5. "What part of that task did you deliberately not hand over?" Worry when the boundary is about capability ("it can't do spreadsheets") rather than judgment. The interesting line is work the model could have done and they kept anyway. 6. "What did you hand over that, looking back, you shouldn't have?" Worry when the answer is nothing. This is a calibration question, and everyone using these tools daily has an answer. 7. "What existed between the brief and the final draft?" Worry when nothing did. A plan, an outline, a criteria list, a rough data cut. An intermediate artifact is what makes the deliverable reviewable by anyone, including them. 8. "Show me a prompt you rewrote. What was wrong with the first one?" Worry when the rewrite is about tone or length. A substantive rewrite changes what is being asked, not how politely. 9. "What did you throw away, and on what grounds?" Worry when everything was kept and edited. Editing is tidying. Rejecting a direction and saying why is judgment. 10. "Tell me about a time you decided the model was the wrong tool and did it by hand." Worry when they treat that as a confession. Volume of AI use is not skill, and a candidate who says so unprompted has understood the job. 11. "What changed because of the check: a number, a recommendation, a caveat?" Worry when the check confirmed everything. A verification that never moves anything is a ritual. 12. "Same task tomorrow, no assistant. What would you do differently?" Worry when the only answer is that it would take longer. The useful answer names a step they would now do first regardless.

Every one of these is a doorway, not a score. The value is in the second question, and the follow-up is what exposes understanding. A prepared answer has depth of exactly one.

Which follow-up fits the candidate's field?

The follow-up has to come from the work the candidate does, because that is where the checking actually happens. A marketer's verification is opening the source behind a statistic; an engineer's is the code that ran and was still wrong; an analyst's is the number that reconciles to nothing. Ask the version they would recognize from their own week.

  • Marketing. "You cited a market-size figure in that brief. Where did it come from, and did you open it?" Worry when the trail ends at the assistant. The most on-message statistic is the one least likely to be opened.
  • Engineering. "Describe a time generated code ran clean and was still wrong. How did you find out?" This one has evidence behind it: in a study of AI-assisted programming, participants with an assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure, while the participants who trusted the assistant least and reworked their prompts produced fewer vulnerabilities 3.
  • Data and analysis. "The model explained your result fluently. What would that explanation have looked like if the result were an artifact of the join?" Worry when a plausible story is treated as a finding.
  • Product. "The request was vague. Did the assistant resolve the ambiguity, or did you?" Worry when the spec is confident about something nobody decided.
  • Legal, claims and revenue-cycle work. "The draft was persuasive. What in the file contradicted it?" Worry when persuasiveness is the whole quality bar.

The same principle runs the other way: you can hand the candidate output that contains a known error and watch what happens, which is a cheaper and harder test than any recollection. Testing whether a candidate catches an AI error takes ten minutes and cannot be prepared for, because the error is yours.

Score the answers before you ask them

Set the standard before the first interview: name what a strong, adequate and weak answer sounds like for each question, then ask every candidate the same questions in the same order and record why each answer landed where it did. Without that, twelve good questions produce twelve unstructured impressions, and the round quietly measures who interviews well.

An interview used to make a hiring decision is a selection procedure: the Uniform Guidelines define one as any measure used as a basis for an employment decision, and put informal and casual interviews in that class alongside written tests 4. Under the EEOC's guidance on employment tests and selection procedures, a procedure that screens out a protected group has to be shown job-related and consistent with business necessity 5, and "the panel felt she was AI-native" is not something you can show. The written rubric and the recorded reasons are what survive that question, months later, from someone who was not in the room.

Three things do not belong in the rubric. Volume of AI use, because a candidate who judged the model was the wrong instrument and did the step by hand has demonstrated the thing being tested. Prompt vocabulary, because it is a month of exposure, not a year of judgment. And the candidate's fluency describing AI, which is the specific skill a rehearsed answer already has. Score the acts: what was framed, what was demanded, what was kept, what was built in between, what was refused, what was tested. A worked version of that scale is in a rubric for AI-assisted answers.

What can't an interview show you?

Whether they would actually do it. Every question here captures a candidate describing a check; none of them captures a check. That limit is real rather than rhetorical: the developers in the trial above were sincerely wrong about their own speed by roughly forty points, and an interview only ever reaches what someone can tell you about themselves 1.

Two ways to close the round cost little. The first is a fifteen-minute live extension of their own answer, assistant open, on material they have not seen. You are watching one act instead of hearing about twelve. The second is a work sample on real occupational material, where the assistant is available and the task is built so that the confident answer is wrong in a way only checking reveals.

Both cost more than a question, and both are the only formats that produce a record you can re-read. If you are choosing between instruments rather than adding one, the trade-offs sit side by side in detector, interview or work sample, and the behaviors worth looking for are the same either way, and what good AI use looks like does not change with the format.

One last thing to say out loud in the invitation: that the round will cover how the candidate works with AI, and that using it is expected rather than tolerated. A rule nobody states gets guessed at, and the best guessers are the most-coached candidates.

See how it works

Common questions

How many AI questions should one interview include?

Three to five, inside a round you already run. Each needs room for a real follow-up, and a follow-up takes longer than the first answer, and that is the part doing the work. Twelve questions asked at a checklist pace produce twelve surface answers and no evidence. Pick the ones that match what the role actually does: framing and refusal for judgment-heavy work, verification and delegation for anything where a wrong number ships.

Should you tell candidates the interview will cover AI use?

Write it into the invitation, in one sentence, before they prepare. An unstated rule gets guessed at, and the guessing tracks coaching rather than capability. Saying it also makes the answers scorable: you can only judge how someone used a model if they knew they were allowed to. State whether an assistant will be open during the round, because a candidate who assumed it was banned will describe their work rather than do it.

What if the candidate says they don't use AI at all?

That is an answer, not a disqualifier. Ask what they decided not to use it for and why. A considered no is a delegation boundary, and it is the same judgment you are testing everywhere else. Then ask how they would check a confident claim from a source that cannot show its work, because that skill predates the tools. The worrying version is not the abstainer; it is the candidate who has used AI daily for a year and cannot name one thing it got wrong.

Is asking about AI use legally risky?

Asking about job-related behavior is ordinary interviewing. The risk lives in how the answers are used: an interview that informs a hiring decision is a selection procedure, so ask every candidate the same questions, score against a rubric written beforehand, and record the reason for each rating. Avoid questions that reach toward disability, language proficiency or personal circumstances, and check anything jurisdiction-specific with counsel.

Can a candidate rehearse these answers with AI?

The generic ones, yes, and they will. What resists preparation is specificity plus a follow-up drawn from the answer just given: which file, which page, which number changed, what the assistant said before you cut it. A model can write a candidate a convincing story about verification; it cannot supply the detail of a check that never happened. If an answer stays smooth at three levels down, that itself is the finding.

Does Olive replace these interview questions?

No. An Olive assessment is bought by the employer: the candidate works a 40-to-60-minute task from their own occupation with an AI assistant available, and a human reviewer writes six findings on what happened: how the problem was framed, what evidence was demanded, what was kept, what was built in between, what was refused, what was tested. Each finding carries the moment it rests on. There is no composite score and no ranking, and the candidate is granted the same report the employer reads. It answers a different question than an interview: what the person did, not what they say they do.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial, 16 experienced open-source developers on their own repositories: 19% longer to complete issues with AI tools allowed, while developers still believed afterwards that AI had sped them up by about 20%.
  2. 2. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ACM CHI Conference on Human Factors in Computing Systems (CHI '25), 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in GenAI is associated with less critical thinking, higher task self-confidence with more, and the remaining thinking shifts toward information verification, response integration and task stewardship.
  3. 3. Do Users Write More Insecure Code with AI Assistants? arXiv (Stanford University; published at ACM CCS 2023), 2023. arxiv.org Participants with access to an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure; participants who trusted the assistant less and reworked their prompts produced fewer vulnerabilities.
  4. 4. 29 CFR 1607.16 — Definitions (Uniform Guidelines on Employee Selection Procedures) Code of Federal Regulations (Legal Information Institute, Cornell Law School), 1978. law.cornell.edu Defines a selection procedure as any measure, combination of measures, or procedure used as a basis for any employment decision, covering the full range of assessment techniques from written tests through informal or casual interviews and unscored application forms.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A selection procedure that screens out a protected group must be shown job-related and consistent with business necessity.

5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.