Screening

How Do You Verify 'Uses AI Daily' on a Resume in Ten Minutes?

Verifying 'uses AI daily' on a resume takes ten minutes: ask for an artifact the work already produced and one output the candidate threw away. Two minutes on which tool, which task, how often. Three on the artifact: a revision history, a pull request diff, an analyst's assumption tab. Five on the rejection, and how they knew it was wrong. If nothing is shareable, which is a normal answer at a serious employer, put all eight remaining minutes on that rejection. Daily use leaves a trail. A resume line doesn't.

The take'Uses AI daily' is heading for the same fate as 'proficient in Excel', and not because anyone is lying. The sentence costs nothing to write, every applicant reads the same posting, and it looks set to sit on most of the pile within a hiring season. What you actually wanted was the moment somebody stopped and checked. Nobody has measured whether days per week predicts anything about that moment, and I'd be surprised if the phrase survives its own popularity. The claim is not being faked so much as retired.

Where Olive fits

Open a role and see what the work shows

Ten minutes gets you an artifact and a story about it; it cannot show you the candidate deciding. Olive is an employer-purchased assignment for the occupation (40 to 60 minutes with an AI assistant), returned as six findings a human writes, each anchored to the moment it happened, and the candidate is granted the same report.

Rank your shortlist

Why is 'uses AI daily' worth ten minutes?

Because daily use is still uncommon enough to be a real claim. As of late 2024, 23% of employed respondents had used generative AI for work in the previous week and 9% used it every workday 1. Pew found roughly one in ten workers using AI chatbots at work daily or a few times a week, with 55% rarely or never 2. That line on the resume is a minority claim, and minority claims are testable.

Two things follow. The claim is cheap to type and expensive to fake in detail. Someone who opens a model once a month has no revision history, no saved project, and no ready answer about the last output they threw out. And the artifact costs the candidate nothing to bring, which keeps the check inside a screen call instead of pushing it into an assessment you haven't bought.

What won't work is checking the resume itself. Text detectors misclassify writing by people who learned English later as machine-written, and the same study showed simple prompting defeats them 3. Sorting applications by suspected authorship is a different question anyway, and mostly a worse one. See whether an AI-written resume should disqualify anyone. What you're testing is how the person works, not who typed the bullet.

What artifact proves the claim, by field?

The cheapest artifact changes with the job, and asking for the wrong one gets you a shrug. For a writer it's a prompt-and-revision history; for an engineer, a pull request whose diff shows what was accepted and what was rewritten; for an analyst, the assumption tab of a model an assistant helped build. Ask for the thing their work would leave behind anyway.

DisciplineWhat to ask forWhat sixty seconds with it tells you
Writing, content, commsThe prompt thread behind one published pieceWhether the brief was framed before generating, and which draft got cut
Software engineeringOne pull request an assistant wrote part ofWhat was accepted verbatim, what was rewritten, what review caught
Data and analysisThe assumption tab or notes sheet of a model built with helpWhether any figure was re-derived by hand, or taken as returned
Sales, support, opsA saved prompt, macro or template still in useWhether it survived contact with real customers, and what got edited out
Recruiting and peopleA rubric or job description drafted with a modelWhat language was cut, and on what grounds
DesignA generated concept beside the shipped fileWhich parts survived, and what the model couldn't be talked into

One rule across every row: ask for something that already exists. A candidate who has to make an artifact for you is doing an unpaid take-home, and what comes back is a performance rather than a record. Provenance is its own question, and a portfolio piece raises who actually built it before it tells you anything about daily practice.

If nothing is shareable, that's a normal answer from anyone working somewhere serious. Skip to the refusal question and stay there for the full eight minutes.

Run the ten minutes: two, three, five

Two minutes on the claim, three on the artifact, five on a refusal. Keep that order, because the artifact is what makes the refusal question answerable and the refusal is where the ten minutes earns its keep. Write the questions down before the first call and ask them the same way of every candidate, because a screen that decides who advances is a selection procedure whatever you call it 4.

  • Minutes 1–2, the claim. Which assistant, on which recurring task, and roughly how many days a week. You're not grading the tool. You're finding out whether there's a specific task attached to the word 'daily', because a tool name with no task is where thin claims stop.
  • Minutes 3–5, the artifact. Have them open it while you're talking and narrate it: what the first prompt asked for, what came back, what's still in the final version. Ask one question the artifact itself can answer ("which of these paragraphs is the model's and which is yours"), so the record does the corroborating, not the anecdote.
  • Minutes 6–10, the refusal. "Tell me about something it gave you that you didn't use." Then the follow-up that actually matters: how did you know it was wrong? A candidate who checked something outside the conversation (opened the source, ran the number, called the person) is describing the behavior you're paying for. A candidate who says it "sounded off" is describing taste.

The refusal question does most of the work because it can't be answered from a tool list. Someone who uses a model daily has thrown out dozens of outputs and remembers at least one; someone who used it twice for the resume itself has nothing to reach for. The same probe scales past this call. It's the backbone of a fuller check on resume AI claims when you have more than ten minutes.

What answers should end the conversation?

Three. A tool list with no task attached to it. A story where the model produced the finished thing and nothing was changed. And a rejection example that turns out to be a formatting complaint. None of the three proves dishonesty; all three mean the claim is thinner than the resume line implied, which is exactly what ten minutes was meant to find out.

  • The tool list. Four product names, no task, no frequency. Usually it means the candidate read your job posting carefully, which is not nothing but is not this.
  • The unedited output. "It wrote the whole deck and the client loved it." Ask what they'd have changed with another hour. If the answer is still nothing, the delegation boundary is the finding, and it's a real one.
  • The cosmetic refusal. "I fixed its tone." Fine, but tone is not a claim. Push once: was anything it asserted actually wrong, and how did you find out.

One caution. Talking well about your own work is a separate skill from doing it, and it isn't evenly distributed: some strong practitioners describe their process badly under time pressure. Give the quiet answer a second question before you close it out, and hold every candidate to the artifact rather than to the fluency. What good AI use actually looks like is a shorter list than most rubrics assume: framed before generating, evidence demanded for the claim that mattered, something tested against the world.

Where does a ten-minute check stop?

At description. Ten minutes gets a candidate telling you about a decision made weeks ago, with an artifact as corroboration. It never gets you the decision itself. You're hearing a reconstruction, filtered through how well someone narrates their own work, and the artifact confirms an outcome rather than the reasoning that produced it.

That's a real ceiling, not a hedge. The screen can tell you the claim is substantiated and give your hiring manager three specific things to open in the next round. It cannot tell you what happens when a confident, wrong answer arrives on a task the candidate has never seen, which is the moment the job is actually made of. For a small team, the honest move is to stop the free version here and put the remaining question in front of a work sample rather than a longer conversation. The shape of the smallest defensible process is one screen, one sample, one decision.

Budget matters less than most vendor pages suggest at this size. A work sample you write yourself, run identically for every finalist and graded against something written down beforehand, beats an unstructured second interview and costs an afternoon. Olive is one instrument in that category and says so; multiple-choice AI literacy tests, code-collaboration graders and unwatched take-homes are all real approaches with different trade-offs.

See a sample report

Common questions

What if the artifact is confidential?

Ask for the shape instead of the content. Someone under NDA can describe how a prompt thread was structured, say which draft got cut and why, or show a redacted diff with the client name removed. If nothing at all can be shown, spend the whole ten minutes on the rejected output, because the reasoning is what you were buying and it travels without the file. Treat "I can't share that" as a normal answer from anyone working somewhere serious, not as a dodge.

Is asking for a prompt history fair to every candidate?

Only if you ask everyone and accept the same substitutes. Some people work under tool bans, on air-gapped systems, or in roles where nothing exportable exists, and none of that says anything about judgment. Write the request down, offer the same fallback to each candidate, and grade the reasoning rather than the polish of the artifact. Consistency is also what keeps the step defensible if anyone asks later how the screen worked.

Can an AI detector do this faster?

No, and it answers a different question. A detector guesses whether text was machine-written; the resume claim is about how someone works. Detectors are also evaded by simple prompting, and they misclassify writing by people who learned English later as machine-written, which points a bias problem straight at your own pipeline. Ten minutes on a real artifact costs nothing, produces something you can write down, and survives being explained to the candidate.

How often does 'daily' actually need to be true?

Ask what the role needs before you ask what the candidate does. Most jobs need someone who reaches for a model at the right moments and leaves it alone at the wrong ones, which is not the same as opening one every morning. Volume is a weak signal in both directions: heavy use with no refusals is worse evidence than occasional use with a clear account of what got thrown out. Set the bar as a behavior and the frequency stops mattering.

How does Olive fit into a ten-minute screen?

It doesn't; it's the stage after. Olive is an employer-purchased assessment: the candidate works a 40-to-60-minute task built for their occupation with an AI assistant available, and a human reviewer writes six findings, each attached to a timestamped moment in the session. There is no composite number and no hiring recommendation, and the candidate is granted the same report the employer reads. The free tier covers ten attempts a month, so a shortlist of three costs nothing.

References

  1. 1. The Rapid Adoption of Generative AI National Bureau of Economic Research (Bick, Blandin and Deming), 2024. nber.org 23% of employed respondents used generative AI for work in the prior week; 9% every workday.
  2. 2. U.S. Workers Are More Worried Than Hopeful About Future AI Use in the Workplace Pew Research Center, 2025. pewresearch.org About one in ten workers use AI chatbots at work daily or a few times a week; 55% rarely or never.
  3. 3. GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu and Zou (arXiv), 2023. arxiv.org Detectors misclassify non-native English writing as machine-generated, and simple prompting evades them.
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A step that decides who advances is a selection procedure and has to be applied consistently.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.