Interviewing
How Do You Get Managers Who Don't Use AI to Judge AI-Assisted Work?
A manager who doesn't use AI can still judge AI-assisted work in an interview without training. Hand them three questions their own domain judgment already scores: what did you ask for first, which specific claim did you check and where, and what did you throw out. The manager writes the answer key, because the tell separating careful AI use from careless is field-specific and only they know it. All three answers are self-report, so plant an error in a short piece of work and watch what the candidate does.
The takeThe weakest reader on the panel isn't the manager who has never opened the tool. It's the one who uses it daily and has stopped being surprised by what it produces. The research here is about people trusting a model, but the mechanism travels: a confident explanation buys acceptance whether or not it is right, and an interviewer impressed by fluent output will be impressed by a candidate narrating fluent output. Nobody has measured that split inside a real loop, so take it as a bet. Even so, the suspicion an AI course is meant to train out of the non-user is the one thing the round cannot run without.
Where Olive fits
Open a role and see what the work shows
Three questions reach what a candidate can tell a manager about checking a confident claim; they stop short of the check itself. Olive puts the act in front of them: an occupational assignment, an assistant willing to do all of it, and a human reviewer who writes six findings anchored to the moments they happened.
Rank your shortlistWhat is the manager actually missing?
Not AI knowledge. An answer key. A manager who has run a marketing team for eight years already knows which statistic in a positioning brief is the one nobody opened, and which one sinks the launch. What they lack is a written standard for what a candidate should have done with an assistant, and the confidence that their own domain judgment is the right instrument for reading it.
The risk is not that they know too little about AI. It is that fluent, well-explained output is more persuasive than correct output. In a study of people making decisions alongside an AI system, explanations increased the chance a person accepted the AI's recommendation regardless of whether it was right, and did not improve the team's accuracy beyond what the AI's recommendations without explanations already produced 1. That finding is about people trusting a model, but the mechanism, a confident explanation buying acceptance, is the same one running on a panel listening to a candidate narrate AI-assisted work.
Encouragingly, the evidence has a second half. In a survey of 319 knowledge workers describing 936 first-hand examples of AI-assisted work, higher confidence in the AI predicted less critical thinking, while higher confidence in one's own ability to do the task predicted more 2. Domain confidence is the asset the manager already has. Build the round to use it rather than to replace it with a month of tool exposure.
What happens without a written standard is predictable. Managers fall back on the one thing they can pattern-match, which is prose style, and rejecting candidates for sounding like AI produces confident decisions with nothing recorded behind them.
Ask these three questions
Three, in this order, about one piece of work the candidate finished in the last month. What did you ask for first. Which specific claim did you check, and where. What did you throw out, and why. Each is a question about an act with a time and a place attached, which is why a manager can score it without ever having opened the tool.
1. "Say the first thing you typed to the assistant, the real message rather than the tidy one." Strong: the opening move states the problem, a constraint, or what would make an answer wrong. Weak: the opening move is the deliverable. "Write the Q3 forecast" hands over the framing before the framing exists. Scoring this needs no knowledge of prompting, because the question is whether the candidate had thought about the problem before anyone typed. 2. "Which specific claim in that output did you check, and where did you check it?" Strong: a claim, a location, and a result. "The churn figure: I opened the filing and it was annual, not quarterly, so the number was off by four." Weak: "I reviewed all of it" or "it looked right." A real check has a place, and the candidate can name it. 3. "What did you throw out, and on what grounds?" Strong: a direction refused, with the reason stated. Weak: everything kept and edited. Editing is tidying; refusing a framing is judgment. Be more worried, not less, when a candidate treats "I decided the model was the wrong tool and did that part by hand" as a confession, because how much AI someone used is not the measure.
Ask the same three, in the same order, of every candidate, and have each interviewer write the reason for the rating rather than the rating alone. Panels drift within two candidates otherwise, and getting a panel to judge AI use the same way is mostly a matter of everyone holding the same sheet of paper.
What does a strong answer sound like in each field?
The words change, the structure doesn't. Each family has one artifact an assistant produces fluently and one check only someone in the field would think to run. A strong answer names that check and says what it moved. A weak answer describes the output as good (well-structured, on-brand, clean), which is a judgment about the surface, not about whether anything underneath is true.
- Software engineering. Strong: generated code that ran clean and was still wrong, and how they found out: a test they wrote themselves, an input they tried, a dependency they read. Weak: "it compiled and the tests passed," when the tests were generated too. There is evidence behind this one. In a controlled study, participants with an AI code assistant wrote significantly less secure code than those without and were more likely to believe their code was secure, while the participants who trusted the assistant least and reworked their prompts produced fewer vulnerabilities 3.
- Marketing. Strong: a statistic in the brief traced back to the primary source, and what happened when the source turned out to say something narrower. Weak: "the copy was on-brand and I fixed the tone." The most on-message number in any AI-written brief is the one least likely to have been opened.
- Finance and accounting. Strong: a figure recomputed by hand or against the filing, and the recommendation that moved because of it. Weak: a model whose logic is described confidently but reconciles to nothing. Ask which cell they rebuilt, and what it changed.
- Legal operations. Strong: a citation pulled and read before it went into a draft. Weak: trusting a research tool because the vendor said it does not make things up. Legal research tools sold on that promise were measured producing hallucinated output between 17% and 33% of the time 4.
None of the four is a question about AI. Each is the check the manager would already run on a junior's work, asked at the point where an assistant makes it skippable. That is also why a list of tools on the resume tells you almost nothing: the tools are common to all four families and the tell is common to none of them.
Build the answer key in thirty minutes
Sit the manager down with one task their team shipped last quarter and ask them the same three questions about it. Their own answers are the key. Write each one as three lines (strong, adequate, weak) in the manager's words, and hand it to every interviewer on the loop before the first candidate. That is the whole training, and it takes half an hour.
Three lines each, anchored to acts. Strong: names a specific claim, a source, and what changed. Adequate: names a check but not what it moved. Weak: describes the output rather than any act performed on it. If two interviewers cannot score the same answer the same way, the wording is carrying too much, and a rubric two reviewers score the same comes down to anchoring every level to something the candidate did.
Write it before the round rather than after. An interview used to make a hiring decision is a selection procedure, and the Uniform Guidelines count even an informal or casual interview as one 5. Under the EEOC's guidance on employment tests and selection procedures, a procedure that screens out a protected group has to be shown job-related and consistent with business necessity 6. "The panel felt he wasn't AI-native" cannot be shown. A dated sheet with three anchors and a written reason per candidate can.
Keep three things off the sheet. Volume of AI use, because a candidate who judged the model was the wrong instrument for a step has demonstrated the thing being measured. Prompt vocabulary, which is a month of exposure rather than a year of judgment. And fluency describing AI, which is precisely what a coached candidate arrives with. What good AI use looks like is a set of acts, and acts are what a non-user can score.
What can't three questions tell you?
Nothing about follow-through. Every answer here is a report on a check, given by someone with an obvious interest in the report, and self-report about AI work is unreliable even when nobody is being interviewed. Sixteen experienced developers in a randomized trial took 19% longer to finish issues with AI tools allowed, and still believed afterwards that the tools had sped them up by about 20% 7.
Two cheap closures. Extend the round by fifteen minutes and hand the candidate a short piece of AI-assisted work with one error planted in it (your error, in your field, so no preparation reaches it), then watch what they do. Or move the whole thing to a work sample, where an assistant is available and the confident answer is wrong in a way only checking reveals. What a work sample should test now changed for exactly this reason.
Say in the invitation that the round covers how the candidate works with AI and that using it is expected rather than tolerated. An unstated rule gets guessed at, and the guessing measures interview coaching rather than judgment. If the manager still wants a shortcut, scoring an answer produced with AI is a shorter document than the AI-fluency curriculum they were about to sit through.
Common questions
Should a manager who doesn't use AI take a course first?
No. A short course teaches tool vocabulary, which is the part a coached candidate already has, and it does nothing for the answer key. Thirty minutes writing down what a strong, adequate and weak answer sounds like for one real task from their own team is worth more than a day of tool training. If the manager wants exposure, the useful version is doing one of their own tasks with an assistant open and noticing where it overreached. That moment is the one they will recognize in a candidate's answer.
What if the manager thinks using AI at all is a red flag?
Settle it before the round rather than inside it. Decide what the role's work actually involves and write the standard down, because two interviewers holding opposite views produce a rating that measures which one the candidate drew. A considered abstention is still an answer worth hearing: ask what they decided not to use it for and why, then ask how they check a confident claim from any source that cannot show its work. The worrying candidate is not the abstainer, but the daily user who cannot name one thing a model got wrong.
Can a recruiter run this round instead of the hiring manager?
A recruiter can ask all three questions. Scoring the field-specific half needs someone who knows what a check costs in that discipline: whether opening the source was hard, whether recomputing the figure took two minutes or two days. The workable split is a recruiter asking and recording answers close to verbatim, with the manager reading them against the key they wrote. That fixes calendar problems and produces a record two people can disagree over, which is the point of writing anything down.
How much time does this add to the loop?
About twelve minutes inside a round you already run: three questions with one real follow-up each. Nothing new gets scheduled. The cost that matters is the half hour with the manager beforehand, and that is paid once per role rather than once per candidate. If the loop is already long, drop a behavioral question instead of adding a stage, since these are behavioral questions with a sharper subject.
What if the role isn't in one of those four families?
Ask the manager one question: what does this job get wrong when someone competent is rushing? The answer names the check that separates careful work from careless work in that discipline, and a strong answer to question two is a candidate who ran it. Underwriting, revenue cycle, supply chain, journalism, teaching: each has one artifact that reads finished before it is, and someone who has done the job for five years can name it in a sentence.
Does an assessment replace the manager's judgment here?
No. Olive is employer-purchased: the candidate works a 40-to-60-minute task from their own occupation with an AI assistant available, and a human reviewer writes six findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each carrying the moment it rests on. There is no number standing for a person, and the candidate is granted the same report the employer reads. It answers a different question than an interview does: what someone did on a real task, not what they can describe having done.
References
- 1. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance ✓ arxiv.org Explanations increased the chance that a person accepted the AI's recommendation regardless of its correctness, and did not improve team accuracy beyond what the AI's recommendations without explanations already produced.
- 2. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in GenAI is associated with less critical thinking, while higher confidence in one's own task ability is associated with more.
- 3. Do Users Write More Insecure Code with AI Assistants? ✓ arxiv.org Participants with an AI code assistant wrote significantly less secure code and were more likely to believe it was secure; those who trusted the assistant less and reworked their prompts produced fewer vulnerabilities.
- 4. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools ✓ arxiv.org Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time, despite vendor claims of eliminating hallucination.
- 5. 29 CFR § 1607.16 — Definitions (Uniform Guidelines on Employee Selection Procedures) ✓ law.cornell.edu A selection procedure is any measure, combination of measures, or procedure used as a basis for an employment decision, including informal or casual interviews.
- 6. Employment Tests and Selection Procedures ✓ eeoc.gov A selection procedure that screens out a protected group must be shown job-related and consistent with business necessity.
- 7. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Randomized controlled trial with 16 experienced open-source developers: 19% longer to complete issues with AI tools allowed, while developers believed afterwards that AI had sped them up by about 20%.
7 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.