Interviewing
How Can You Tell If a Candidate Really Uses the AI Tools They List?
A resume listing six AI tools can't tell you whether the candidate actually uses them, so ask. Sort the six into categories (coding assistant, research assistant, general chat, data copilot) and keep the two or three the job runs on. For each, ask for one time the tool was wrong, how they found out, and what they changed. Daily users answer with an event. A blank isn't dishonesty: bans, NDAs and plain nerves all produce one. And what you get is still an account of the catch, reconstructed weeks later.
The takeThe six-tool list is a genre convention by now, and reading it as a claim about skill hands it a weight nobody who typed it intended. What this probe really measures is how well someone narrates old work under time pressure, which is a genuine talent and not the one in the job description. The strongest AI users I've watched interview flatly: they describe the correction with no moral attached to it. Nobody has measured whether fluent self-description predicts careful checking, and until someone does, a thin answer is evidence about the interview at least as much as about the candidate.
Where Olive fits
Open a role and see what the work shows
A probe about a wrong answer gets you the candidate's account of catching one; it can't show you the catch. Olive puts that in front of them as work: an occupational assignment, an assistant that will overreach, and a human reviewer who writes six findings against the moments in the session, with the candidate granted the same report.
Rank your shortlistWhat does a six-tool list actually tell you?
That the candidate reads the same feeds as everyone else in their field. Adoption is close to universal in software: 84% of respondents to Stack Overflow's 2025 developer survey use or plan to use AI tools, and 51% of professional developers use them daily 1. A list of six separates nobody. What a product name can't tell you is whether the person noticed when the tool was wrong.
Self-report doesn't close that gap either. In a randomized trial run by METR, 16 experienced open-source developers worked 246 real issues from their own repositories, and allowing AI tools made them 19% slower. They had forecast a 24% speedup beforehand, and after the fact they still believed the tools had sped them up by 20% 2. Sincere people were wrong about their own AI use by roughly forty points. Asking "how good are you with Cursor" samples the same faulty instrument.
So ask about events instead of ability. Events are checkable, they have a date and a consequence, and a candidate who worked the way the resume implies has dozens to choose from. Outside software the list is also a stronger claim than it looks: in late 2024, 23% of employed US respondents had used generative AI for work in the previous week and 9% used it every workday 3. Six tools on a marketing or finance resume is a minority claim, which makes it worth ten minutes. The shorter version of this check is ten minutes with no assessment budget.
Ask one question per tool category, not per tool
Six names usually collapse into three or four categories, and the probe belongs to the category. Ask what it got wrong, how they found out, and what they changed. Run it once per category rather than six times, or the interview becomes a product quiz, and tool trivia is not the thing you're hiring for. Pick the two categories the job actually runs on.
| Category | Names you'll see | The probe | The answer that ends it |
|---|---|---|---|
| Coding assistant or agent | Copilot, Cursor, Claude Code, Windsurf | "Describe a change it wrote that you rewrote. What was wrong with it?" | "It's usually right, I just tidy the formatting" |
| Research and synthesis | Perplexity, NotebookLM, a deep-research mode | "Which source did it cite that didn't say what it claimed?" | Never opened a cited source |
| General chat and drafting | ChatGPT, Claude, Gemini | "What did it assert confidently that turned out false, and how did you find out?" | "It sounded off, so I rewrote it" |
| Spreadsheet and data | Excel Copilot, a notebook assistant, a BI copilot | "Which figure did you re-derive by hand, and did it match?" | Nothing was ever re-derived |
| Meeting notes and automation | Otter, Fireflies, an agent wired into a workflow | "What did a summary get wrong about a decision, and who caught it?" | Summaries forwarded unread |
| Image and design generation | Midjourney, Firefly, a built-in generator | "What did it produce that you couldn't ship, and why not?" | Only cosmetic objections |
The list is field-coded, so read it that way. Copilot and Cursor on an engineering resume are duty tools and their absence would be the odd thing; a research assistant on a consulting or marketing resume sits closer to the analysis itself, and the honest probe there is about sources rather than about diffs. In finance and revenue operations the tool usually touches a number somebody will act on, which changes the question again, and screening for AI judgment outside engineering runs on the same three-part probe with different artifacts behind it.
Why does the disqualifying answer change by tool?
Because each tool invites a different shortcut. A coding agent invites shipping a diff nobody read; a research assistant invites citing a source nobody opened; a notetaker invites forwarding a summary nobody corrected. The wrong-answer question stays the same, but what counts as a failing answer follows the tool's own failure mode. Decide before the call which answer ends the line of questioning.
- Coding assistants fail plausibly. The top frustration in Stack Overflow's 2025 survey was "AI solutions that are almost right, but not quite," reported by 66% of respondents, with 45.2% naming debugging AI-generated code as more time-consuming than writing it 1. A candidate who uses one daily has been burned by a near-miss and can name it. "It saves me hours" with no near-miss attached is the thin answer, because near-misses are the dominant experience of the tool.
- Research assistants fail invisibly. The output reads finished and the citations look real. So the only useful probe is whether a source was opened: which claim did you check, what did the page actually say, what changed in the deliverable. A candidate who has never opened a cited source has been handing on someone else's confidence.
- Data copilots fail arithmetically. A formula that runs is not a formula that's right. Ask which figure was re-derived by hand and whether it matched. "I sanity-check the outputs" is a posture; "the growth rate was compounding when it should have been average, and I caught it re-adding the column" is an event.
- Notetakers and workflow agents fail silently. Nobody reads the summary against the meeting. Ask who caught the last error and what it cost.
One pattern runs under all four. The useful answer names something outside the conversation that settled the question: a page opened, a number recomputed, a colleague called. "It seemed off" is taste, and taste is not verification. That distinction is the whole content of testing whether a candidate catches AI errors, and it survives being asked of a senior and a new grad in the same words.
What if they can't name a time it was wrong?
Press once, then write down what you got. Ask for the most recent thing they threw away rather than the most interesting one, because recency is easier to retrieve and harder to invent. If the second attempt still produces nothing, the finding is that the list is thinner than it reads, which is what the question was for. None of this is evidence of dishonesty, and it shouldn't be recorded as if it were.
Three honest reasons a good candidate blanks. Their last employer banned the tools, or ran on air-gapped systems where nothing was allowed near a model. Their work is under NDA and they've been trained not to describe it. Or they narrate their own process badly under time pressure, which is a separate skill from doing the work and is not evenly distributed across candidates. Offer the same fallback to everyone: describe the shape of the correction without the client, the file or the number.
A candidate who says they barely use these tools is answering a different question, not failing this one. What to do when your strongest candidate doesn't use AI is worth settling before the interview rather than during it.
Ask every candidate the same probes, in the same order, and record the answers the same way. A step that decides who advances is a selection procedure, and consistency is what makes it explainable later 4. It's also the only way six answers become comparable, and running a structured interview about AI use is the version of this you can put in front of a panel.
Where does this probe stop?
At a story about work you didn't see. The probe gets a candidate describing a correction made weeks ago, reconstructed under time pressure and filtered through how well they talk about themselves. That's better evidence than a list of product names, and it's still description. Nobody in the room watches the candidate catch a confident wrong answer while it's happening.
The ceiling matters because the behavior you're buying is a live one. On a task the candidate has never seen, with a deadline and an assistant that answers everything fluently, what gets framed first, what gets demanded a source, what gets kept by hand, what gets refused. None of that is visible from an account of it. Interview answers are also the most coached surface in the process now, which is why the next step is usually a work sample rather than another conversation.
If you go that way, write the answer key before you write the task, and let the assistant be open rather than banned, because a banned tool just moves the work off-screen. Take-home assignments where candidates use AI still carry signal when the grading is about what was refused and checked rather than about polish. Olive is one instrument in that category and says so; multiple-choice AI literacy tests, code-collaboration graders, unwatched take-homes and in-house rounds are all real approaches with different trade-offs. See how Olive measures this.
Common questions
Is listing six AI tools on a resume a red flag?
No. It's a claim about exposure, and in most technical fields it's an accurate one, since daily use is now the norm among professional developers. Treat the list as a menu of questions rather than as a credential or a warning sign. The only thing worth reacting to is a candidate who can't attach a single concrete event to any of the six, and even that is a thin-claim finding rather than a character judgment.
Should I ask them to demo a tool live in the interview?
Only if every candidate gets the same task and the same time. A live demo mostly measures fluency with a user interface and comfort being watched, neither of which is the skill in question. If you do run one, grade what the candidate refuses and what they check, not how fast they type or how polished the output looks. A short prepared task beats "show me how you use Cursor," which rewards performance.
How many of the six should I actually probe?
Two or three, chosen by the role rather than by the resume order. Pick the categories the job runs on: the coding assistant for an engineer, the research assistant for a consultant or marketer, the data copilot for an analyst. Probing all six turns a 45-minute interview into an inventory and tells you less, because the sixth tool is usually the one they've touched twice.
What if the candidate's employer banned AI tools?
Take it as a normal answer and switch the question. Ask how they decided what to check in work that had no assistant in it: which claim they verified, what they found, what changed. The underlying behavior is the same one, and it predates the tools. Offer that substitute to every candidate rather than only to the ones who bring it up, so the round stays comparable.
How does Olive fit alongside this interview question?
It's the stage after. This question gets you a candidate's account of catching a wrong answer; Olive puts the catch in front of them as work. They take an assignment built for their occupation, with an AI assistant available that will overreach, and a human reviewer writes six findings, each anchored to a moment in the session. The candidate is granted the same report you read.
References
- 1. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co 84% use or plan to use AI tools and 51% of professional developers use them daily; 66% name "almost right, but not quite" as their top frustration and 45.2% name debugging AI-generated code as more time-consuming.
- 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, after forecasting a 24% speedup and while still believing afterwards that they had been sped up 20%.
- 3. The Rapid Adoption of Generative AI ✓ nber.org 23% of employed respondents used generative AI for work in the prior week and 9% used it every workday as of late 2024.
- 4. Employment Tests and Selection Procedures ✓ eeoc.gov A step that decides who advances is a selection procedure and has to be applied consistently across candidates.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.