Screening
Two AI Questions Worth Their Minutes in a Fifteen-Minute Screen
In a fifteen-minute phone screen, two AI questions are worth the time. Which part of your last job did AI take over, and what did you stop doing by hand because of it. Then: what did you try with AI that did not work. Neither can be answered in the abstract, and both end in something specific. Write down the one claim the next round will test rather than rating the answer. Fifteen minutes cannot verify AI skill, but it can decide what to verify.
The takeThe AI slot in most phone screen scripts is an inventory question, and a list is the one answer a candidate can give without having done any of the work, so the minute buys nothing that separates anybody. A screen that hands the next stage one testable claim per candidate is worth more than a screen that produced a rating nobody reopens. Treat the fifteen minutes as routing rather than as evidence, and it starts earning its place in the loop.
Where Olive fits
Open a role and see what the work shows
Olive is where a claim recorded in a screen gets checked as work: a role-grounded assignment on the candidate's own clock, a think-aloud spoken or typed, and a short written debrief, returned as six findings with the evidence excerpt behind each one. A person writes every word of that report.
Rank your shortlistAsk These Two Questions
Which part of your last job did AI take over, and what did you stop doing by hand as a result. Then: what did you try with AI that did not work. Both need a real job behind them, they take about four minutes together with follow-ups, and both end in something specific enough to check later. Everything else in the AI slot can go.
The first question has a tell. A candidate who actually moved work names a task, a cadence and a consequence: the weekly variance commentary, drafted by a model since March, and the two hours a week that went somewhere else. A candidate who has not will describe a capability rather than a change, and the follow-up that separates them is one sentence long: what do you do now that you did not do before.
The second question is the one worth protecting when the call runs short. A failure has to be lived through before it can be described, so the answers are hard to fabricate and easy to follow up. Ask what the model got wrong, how they found out, and what they changed afterwards. "I stopped asking it for citations and started checking them myself" is a small answer that carries a lot of the signal you were hoping for.
One caution on both: do not reward the story, reward the specificity. A polished narrative about an AI transformation is exactly what a candidate can prepare, while a boring answer about a spreadsheet macro and a hallucinated figure usually is not.
Why Does the Tool Inventory Question Fail?
Because a list can be recited by anyone and checked by nobody. Naming a tool records exposure, not judgment, and the two do not travel together: when 288 teachers took both a self-report and a knowledge test built on one AI literacy framework, correlations between the two ran from 0.07 to 0.24 1. A tool list is a self-report with brand names in it.
That study covered teachers in Taiwan rather than job candidates, so read it for the shape rather than for a rate. Two instruments, built on the same four-part framework, largely failed to agree, and the profiles the researchers found included people who underrated themselves as well as people who overrated themselves. Both groups are misread by a question about tools.
An unscored conversation can also make the read worse rather than neutral. In a controlled test, undergraduates predicting a classmate's semester grades did worse after conducting an unstructured interview than they did from prior grades alone, and a majority chose to interview someone answering questions at random over conducting no interview at all 2. That is undergraduates and grades rather than recruiters and jobs, so what travels is the mechanism: low-diagnostic information dilutes good information, and an interviewer will make sense of anything. Six AI tools on a resume is the same problem in written form.
If a tool question has to stay in the script, make it about a choice rather than a list. "Which one did you stop using, and why" can only be answered by someone who used two, and it produces a comparison instead of a recitation. It is still weaker than either of the main questions, so give it the minute left over rather than the four that matter.
Write Down the Claim, Not a Rating
One line per candidate, in the candidate's own words, naming what the next round will check. "Rebuilt the weekly variance commentary with a model and now reviews rather than writes it" is a claim with a task inside it. "Strong AI user, four out of five" is a number nobody can reopen and nobody will argue with. The screen's real output is a question for the next stage.
That line has somewhere to go. It belongs in the scorecard field the hiring manager reads before the next conversation, and it belongs in the brief if a work sample follows, because a task aimed at the candidate's own claim is more informative than a generic one and takes no longer to grade. The manager arrives knowing exactly what is unresolved.
It also protects the candidate from the screen. A rating produced in fifteen minutes travels through the process as though it were measured, and it usually reflects fluency on a phone call. A written claim travels as what it is: something a person said, which the process intends to check. Verifying an AI claim in ten minutes with no assessment budget is the cheapest way to close it if the loop has no room for a task.
Keep the Screen the Same for Every Candidate
Same two questions, same order, same follow-ups, notes in the same three fields. That is what structure means in the research literature and in federal practice guidance: identical questions asked in the same order, a common rating scale, and agreement in advance on what an acceptable answer contains 3. At fifteen minutes it costs nothing to run this way.
It is also the largest single improvement available at this stage. In the 2022 re-analysis of selection research, structured interviews estimate at .42 against .19 for unstructured ones, roughly double the predictive value 4. Those are research codings of interview format rather than products, and none of the underlying studies involved questions about AI, so read the gap as an argument for structure rather than as a number attached to this script.
Structure at this stage has a second payoff that shows up later. When every candidate answered the same two questions, the difference between a strong remote screen and a flat onsite is visible as a difference in evidence rather than as a mood. When the screen and the onsite disagree is a much easier conversation when the screen was the same for everyone.
What structure does not require is a script read aloud without warmth. Ask the two questions in the same words, follow up as far as the answer goes, and stop. Whether AI skill can be screened from a resume at all is the upstream question that decides how much this screen has to carry.
Common questions
What if the candidate's job did not involve AI at all?
Then the question about what AI took over has an honest answer and the screen has learned something real. Ask what they would move onto a model first and what they would keep by hand, which tests the same judgment without requiring a history. A candidate in a role or a sector where the tools were blocked is not a candidate without judgment, and treating an absence of exposure as an absence of skill filters on employer policy rather than on the person.
Should the AI questions be sent in advance?
Sending them makes answers more considered and less anxious, and it costs the element of surprise, which was never worth much. Since neither question can be answered without a real job behind it, preparation improves the quality of the account rather than manufacturing one. Send them with the confirmation email and keep the follow-ups unsent.
How should the answers be recorded?
One claim per candidate in the candidate's own words, plus the follow-up that was asked and what came back. Three sentences is the right size. Avoid a number: a rating produced in a fifteen-minute call is read downstream as though it were measured, and it mostly records how comfortable someone is on the phone.
Does this replace asking about the rest of the job?
No. These two questions occupy roughly four minutes of a fifteen-minute call and belong alongside the usual ground on scope, motivation, availability and compensation range. The point is that the AI slot in the script stops being an inventory question, not that the screen becomes an AI screen.
What if every candidate now gives a rehearsed answer?
The opening question about what AI took over can be rehearsed, and the follow-ups cannot, which is why they matter more than the opener. Ask what they checked before sending, how they found out something was wrong, and what they do differently now. Rehearsed material runs out in about two exchanges, and what happens after that is the part worth writing down.
References
- 1. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arxiv.org Supports the claim that self-reported and tested AI literacy correlate weakly (0.07 to 0.24 across 288 teachers), so a tool inventory does not report capability.
- 2. Belief in the unstructured interview: The persistence of an illusion sjdm.org Supports the claim that an unstructured conversation dilutes better evidence, and that interviewers will construct meaning from answers given at random.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the definition of a structured interview as the same questions in the same order, a common rating scale, and agreement in advance on an acceptable answer.
- 4. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the claim that structured interviews estimate at .42 against .19 for unstructured interviews in the current re-analysis.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.