Interviewing
What Does Being Good at Using AI Look Like in an Interview?
Good AI use in an interview shows up as three moves. The candidate frames the problem in their own words before generating anything, demands a source for the one claim the AI's answer rests on and opens it, and tests something against the world outside the chat: the code run, the figure recomputed, the person called. That check belongs to the candidate's own field, so ask what they opened. Tool lists and speed are not evidence. And an interview gets you the account of the work, never the act.
The takeThe failure mode here isn't the candidate. It's the panel. Talking well about AI is the easiest thing in hiring to rehearse, and a 45-minute round quietly rewards whoever narrates their own work best, which is a different skill from doing it. The survey behind this piece found higher confidence in the tool tracking with less critical thinking, not more. So the most persuasive AI talker in your loop is, I suspect, often the one carrying the least evidence. A panel that can't tell those two apart will keep hiring fluency and calling it judgment.
Where Olive fits
Open a role and see what the work shows
Description is as far as a round reaches, so Olive hands the candidate the work instead: a 40-to-60-minute occupational assignment with an AI assistant available. A human reviewer writes six findings, each anchored to the timestamped moment it happened, and the candidate is granted the same document.
Rank your shortlistWhat are the three behaviors worth watching for?
Framing, evidence, and verification. A candidate who is good with AI states the problem in their own terms before asking for output, demands a source for the one claim the answer rests on, and checks something against the world outside the chat. Everything else (tool names, prompt syntax, how fast the draft arrived) is a proxy for those three, and a weak one.
- Framing before generating. The first move goes after understanding rather than output. It names what is being decided and what would make an answer wrong, in the candidate's own words. The failure shape is asking the assistant for the deliverable in the first message and letting the framing arrive with it.
- Evidence for the claim that matters. Not "add citations": a source demanded for one specific load-bearing claim, and actually opened. The failure shape is confident assertions flowing into the deliverable as facts, with nothing asked for behind them.
- A test against something outside the chat. A number recomputed, a script run, a page opened, a person called. The failure shape is treating fluent description as a result, which is the single most expensive habit a new hire can bring.
Those three are not a taste judgment. A survey of 319 knowledge workers, describing 936 real tasks they had done with generative AI, found the thinking does not disappear when a model arrives. It moves, from gathering information toward verifying it, integrating a partial response, and stewarding a task the model is doing part of 1. What is left after the drafting is automated is exactly what an interview should be asking about.
If you want the same idea as a scored instrument rather than a question list, the four Ds rubric is the closest widely-used framing, and it maps onto these three moves cleanly enough to borrow.
Why does verification look different in every field?
Because the thing you check against belongs to the job, not to the chat. A financial analyst ties a number back to the filing. An engineer runs the code and reads the failure. A marketer opens the campaign data the model summarized. A paralegal pulls the case. The move is the same; its object is the field's, which is why one generic definition of AI skill cannot be assessed.
| Field | What gets checked, and against what | What a real check leaves behind |
|---|---|---|
| Financial analysis | A figure, against the filing or the source packet | A number that changed, or a claim withdrawn |
| Software engineering | The generated code, by running it | A failing test, a rewritten function |
| Marketing | A summarized result, against the campaign data | A statistic cut, a claim re-scoped |
| Legal and contracts | A cited authority, by opening it | A position softened, a citation dropped |
| Data and analytics | A result, by re-deriving it | A different coefficient, a caveat added |
| Recruiting and people ops | A drafted requirement, against the actual role | A line removed as unjustifiable |
The stakes scale with the domain. Domain-specific legal research tools (not general chatbots, but products sold to lawyers on the promise of grounded answers) hallucinated between 17% and 33% of the time in a 2024 Stanford evaluation 3. Nothing in the interface distinguished a usable memo from a sanctionable one. Somebody opening the cited case did.
So write your question against your own material. "How do you check AI output" gets you a philosophy; "walk me through the last time you caught something wrong in a model's answer, and what you opened to catch it" gets you the field-specific act or nothing. If your team spans several functions, the check differs per function, which is a design constraint on any competency framework you write, not a detail to smooth over.
Ask what they threw away, not what they use
The most informative question in the round is about a rejection. Ask for one thing an assistant produced that the candidate did not use, then ask how they knew. A candidate who opened the source, re-ran the number or called the person is describing verification. A candidate who says it sounded off is describing taste, which is a much weaker piece of evidence.
Four questions, in this order, do most of the work:
1. "Take me through the last time you used it. What were you deciding?" You are listening for a task with a decision attached, not a tool name. A four-product inventory with no task behind it is where thin claims stop, and the pattern is common enough to have its own failure mode. 2. "What did you ask it before you asked it for the output?" Framing shows up here or nowhere. Silence is an answer. 3. "Show me something it gave you that you didn't use. How did you know?" The rejection question. The follow-up is the whole question; the first half is just setup. 4. "What did you deliberately keep and do yourself?" Someone who judged the model was the wrong instrument for a step, and says why, has demonstrated the thing you are hiring for.
One caution about who scores well here. In that same survey, higher confidence in the tool was associated with *less* critical thinking, while higher confidence in their own ability was associated with more 1. The most fluent AI talker in your pipeline is not automatically the one doing the checking, and the two are easy to confuse in a 45-minute round. Push past the first answer every time. The follow-up is what exposes understanding, because a rehearsed account survives one question and rarely survives three.
What doesn't count as being good at AI?
Volume, speed, tool inventory and prompt syntax. None of them survives contact with the job. Sixteen experienced open-source developers in a 2025 randomized trial took 19% longer to complete real issues in their own repositories when AI tools were allowed, and still believed afterwards that the tools had sped them up by 20% 2. If practitioners misread their own throughput that badly, a candidate's account of their gains is not evidence.
- Volume of use is not a virtue. Heavy use with no refusals is worse evidence than occasional use with a clear account of what got thrown out.
- Speed is not judgment. The candidate who produced the deliverable fastest may simply have skipped the checking, and in a 45-minute conversation those two look identical.
- Certificates and tool lists price in nothing. They say a course was completed, not that a wrong answer was ever caught.
- Polish is a distractor. Every candidate's answers are more fluent than they were two years ago, which is why the same well-shaped STAR answers keep arriving and why fluency has stopped carrying signal.
The useful inversion: stop asking what the candidate can produce with a model and start asking what they refuse from one. Production is the part that got cheap. Hiring for verification rather than production is the same argument applied to the job description, and it changes what the interview is for.
Where does an interview stop?
At description. An interview can capture a candidate telling you how they checked a confident claim last quarter; it cannot capture them checking one. What you get is a reconstruction, filtered through how well someone narrates their own work under time pressure. Narration is a separate skill, unevenly distributed, and one plenty of strong practitioners have less of than the people they outperform.
That ceiling is worth stating plainly, because the fix is cheap and most teams skip it. Write the three moves down as a rubric before the first interview, ask every candidate the same four questions in the same order, and record what the answer contained rather than how it felt. Even an informal or casual interview counts as a selection procedure once it decides who advances 5, and the federal guidance on those procedures is to administer them the same way for every candidate 4.
Past that, the only instrument that closes the gap is one where the work happens in front of you: a short task on occupational material, with an assistant available and a rubric written beforehand. That is a work sample rather than an interview, and it costs an afternoon to build. Replacing an online assessment that AI has already solved usually funds it. Multiple-choice AI literacy tests, code-collaboration graders, unwatched take-homes and live AI-open rounds are all real approaches with different trade-offs; what none of them can do is tell you what happened, moment by moment, unless the session was recorded and read by a person.
Common questions
What if a candidate says they barely use AI?
Ask what they do instead, and what they check. Someone who works in a regulated environment, on air-gapped systems, or under a tool ban may have deep verification habits and no model in the story at all. The three moves survive translation: how the problem was framed before work started, what evidence was demanded for the claim that mattered, what got tested against something outside their own head. Judge the moves, not the tool. Low usage plus rigorous checking is a stronger hire than daily usage with nothing ever refused.
Is a tool list ever useful in an interview?
Only as a starting point for a real question. Knowing someone uses a specific assistant tells you what to ask next (which task, how often, what it got wrong) and nothing on its own. Tool names are the cheapest thing on a resume to add and the first thing candidates tune to a job posting. Treat the list as an index into their work, spend zero minutes grading it, and move to the task and the rejection as fast as the conversation allows.
Should candidates be allowed to use AI during the interview itself?
If the job uses AI, an AI-open round is closer to the work than a banned one. The trade-off is that watching someone prompt live measures composure under observation as much as judgment, and a 30-minute window rarely leaves room for the checking step that carries the signal. An open round works best when the task has a wrong answer the model will confidently produce, and the interviewer is watching for whether it gets caught rather than for how the prompt was worded.
How do you keep this fair across candidates?
Same questions, same order, same rubric, written before the first interview. Score what the answer contained (was a check named, was a source opened, was anything refused) rather than how confident the delivery sounded, since fluency in describing your own work is unevenly distributed and mostly unrelated to doing it well. Give every candidate the same accommodation options and the same amount of time. Consistency is also what makes the step explainable if anyone asks later how the round worked.
Does this work for non-technical roles?
It transfers, and the questions barely change. A recruiter checks a drafted job requirement against the actual role and cuts what cannot be justified. A support lead checks a generated macro against a real ticket history. A communications manager opens the statistic the model put in the second paragraph. The verification object is different in every function; the move (testing a claim against something outside the conversation, and changing the work when it fails) is identical, and it is just as interviewable.
Can an assessment show these moves instead of describing them?
Yes, because nothing is being narrated. Each of the three moves gets a finding of its own, and so does the rejection: a human reviewer writes six findings covering problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification, each carrying the timestamped moment it rests on. Olive is employer-purchased: a 40-to-60-minute assignment authored for the candidate's occupation, with an AI assistant available. Each finding lands as demonstrated, partly demonstrated or not demonstrated, with no number standing for the person and no hiring recommendation. The candidate reads the identical document.
References
- 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ microsoft.com 319 knowledge workers described 936 real tasks; higher confidence in the tool was associated with less critical thinking, higher self-confidence with more, and the effort shifted toward information verification, response integration and task stewardship.
- 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers, 246 randomized real issues: 19% slower with AI tools allowed, while believing afterwards they had been 20% faster.
- 3. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools ✓ arxiv.org Domain-specific legal research tools hallucinated between 17% and 33% of the time, so the field's own verification step is what separates a usable answer from a wrong one.
- 4. Employment Tests and Selection Procedures ✓ eeoc.gov The EEOC's guidance on employment tests and selection procedures: administer them the same way for every candidate, without regard to protected class.
- 5. 29 CFR 1607.16 — Definitions (Uniform Guidelines on Employee Selection Procedures) ✓ law.cornell.edu A selection procedure is any measure used as a basis for an employment decision, explicitly including "informal or casual interviews" and unscored application forms.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.