Screening
How Do You Verify 'AI-Proficient' on a Resume?
An 'AI-proficient' line on a resume can't be verified from the page: nothing on it is evidence, and a detector answers only whether a model touched the text. Verify it like any other claimed skill: one short task drawn from the role, an assistant allowed in writing, the working record graded instead of the deliverable. What was framed, what evidence was demanded and opened, what was refused, what was checked outside the model. If your team teaches the tooling in week one, verify judgment and don't gate on the tool.
The take'AI-proficient' is heading where 'proficient in Microsoft Word' already went: a line everyone writes and nobody reads. That's not a candidate honesty problem. The phrase names a tool, and tools don't come with levels, so the claim has nothing to be true or false about. What does have a level is the refusal, the moment someone reads a fluent wrong answer and declines it. Screen on the claim and you may well be screening for whoever writes the most assured resume, which is the one thing an assistant now supplies free. It's the most commoditized line on the page, and still, I'd bet, the first one screens read.
Where Olive fits
Open a role and see what the work shows
A resume claim can't be checked from the document, so Olive assesses the person instead: a 40-to-60-minute assignment built on one occupation's own work, done with an AI assistant, returned as six findings a human wrote with the timestamped excerpt behind each. The candidate is granted the identical report.
Rank your shortlistWhy can't you verify AI proficiency from the resume?
Because nothing on the page is evidence. "AI-proficient," a prompt-engineering certificate, a portfolio of unusually clean work: all of it is self-reported, and the same tools that produced the work will happily write the claim about it. Detection doesn't rescue this, either: what a detector answers is whether a model touched the text, which is not the question you asked.
That distinction carries the whole problem. Detection asks whether an artifact was AI-written; verification asks whether this person does good work with AI in the room. Only the second is a hiring signal, and only the second survives a challenge. The first puts you in the position of making an employment decision on a probabilistic guess about authorship.
The useful part is that verification is an old problem with a settled answer. You already verify claimed skills you can't take on faith, by asking the candidate to do a piece of the job. Nothing about AI changes that mechanic. What changes is which part of the work you have to watch, because the finished deliverable is no longer the part that separates people.
What does 'AI-proficient' mean for this role?
It means something different in every job, which is why a single screening question is either unfair or uninformative. For a software developer, the observable behavior is testing generated code against something the model didn't write (a test suite, a linter, a staging run). For a financial analyst, it's demanding the source behind the number that moves the recommendation, and then opening it. Same habit, different act.
The occupational task lists make this concrete. The O\*NET summary for Software Developers (SOC 15-1252.00) carries 17 tasks, among them developing or directing software system testing or validation procedures 4. The summary for Financial and Investment Analysts (SOC 13-2051.00) carries 26 different ones, built around advising clients on capitalization, assessing companies as investments, and recommending remedies for firms in difficulty 5. A verification question written for one of those roles tells you close to nothing about the other.
So do the small piece of thinking first. Name the two or three moments in this specific role where a confident AI answer could be wrong in a way that costs money, credibility or a customer. Those moments are the verification target, and they are also the only defensible basis for setting a proficiency bar at all.
Design one task that makes the claim testable
Hand over a short piece of the actual job, an assistant they may use however they like, and a deliverable with a real constraint attached. Forty to sixty minutes is enough. The task must contain at least one thing a model will get confidently wrong (a stale figure, a plausible claim with no source, a library function that doesn't exist), because that is the moment proficiency becomes visible.
Three properties do the work:
- It comes from the role. The Uniform Guidelines on Employee Selection Procedures are explicit that a content-validity case starts with a job analysis of the important work behaviors, and that "the closer the content and the context of the selection procedure are to work samples or work behaviors, the stronger is the basis for showing content validity" 2.
- AI use is permitted, in writing. A task that quietly bans the assistant measures compliance and tests something you will never observe again after the hire. Say the tool is allowed, and say the transcript is part of the submission.
- The trap is real, not cruel. A wrong figure planted in the source packet is fair. A task nobody can finish in the time given tests stamina.
One caveat the same guidelines put in writing: content validity is not an appropriate strategy for a knowledge, skill or ability an employee is expected to learn on the job 2. If your team teaches the tooling in week one, verify judgment rather than tool familiarity, and don't gate on the tool at all.
Read the working record, not the finished deliverable
Two candidates can hand back the same document, and everything that separates them sits upstream of it. Ask for the assistant transcript alongside the deliverable, or watch the session live, and read for six acts: what was framed before anything was generated, what evidence was demanded, what was kept by hand, what existed between the brief and the answer, what was refused, and what was tested against something outside the conversation.
Each one either happened or it didn't, which is what makes them gradeable:
- Framing. Did the first move go after understanding the problem, or straight at the deliverable?
- Evidence. Was a source demanded for a specific claim and actually opened, or was "cite your sources" typed once and forgotten?
- Delegation. What did the candidate keep? A check run by hand, or a judgment made without asking, is worth more than volume of prompting.
- Structure. Was there anything the deliverable can be read against: a plan, a criteria list, an outline?
- Refusal. Was any model output rejected on stated grounds, or was all of it accepted with the edges tidied?
- Verification. Was a claim tested outside the chat, and did the result change a number, a recommendation, or a stated limit?
How much AI got used is not on that list. A candidate who decided the model was the wrong instrument for a step and did it by hand has shown you the thing you were trying to see. Write the standard down as a rubric before the first session runs, or the first impressive deliverable will set it retroactively, which is one route to the candidate who aces the loop and struggles in the first quarter.
Common questions
Can an AI detector tell me whether a candidate is genuinely AI-proficient?
No. A detector estimates whether a model produced a piece of text; proficiency is about what the person did with the model, and the artifact doesn't record that. The two questions come apart even when the detector is right. A strong analyst who drafted with an assistant and then checked every figure reads as AI-generated; a weak one who pasted output untouched reads the same way. The output can't separate them, so verify with a task.
Should candidates be told the AI assistant is allowed?
Yes, in writing, before the task starts. A task that quietly bans AI and then judges people for using it measures compliance, not capability, and it produces a result you can't defend or reuse. Saying the tool is permitted also lets you ask for the transcript as part of the submission, which is the only part of the work that shows judgment. Candidates who ask which tools count are asking a fair question. Answer it the same way for everyone.
How long should a verification task be?
Under an hour for most roles. A task long enough to be a project selects for people with free time and for candidates who can absorb unpaid work, and completion falls as it grows. An hour is still long enough to contain one confident model error and to force a real decision about it, which is all the task has to do. If a role genuinely can't be sampled in an hour, sample one slice of it rather than stretching the clock.
What if a candidate chooses not to use AI on the task?
That's a result, not a refusal to take part. Note what they did instead and whether the judgment held up. Deciding a model is the wrong instrument for a step, and doing that step by hand, demonstrates the same skill you were trying to see. Treat heavy use and light use as neutral facts and grade the reasoning around them. Volume of AI use is not what the job pays for, and scoring it rewards the candidate who prompts the most.
Is a prompt-engineering certificate evidence of anything?
It evidences that a course was completed. It doesn't show what the person does when a model states something confidently wrong in the middle of real work, which is the behavior the job depends on. Treat a certificate the way you'd treat any other coursework line: a reason to ask a sharper question, not a substitute for the answer.
Can a structured interview verify this instead of a work sample?
Partly. A structured interview captures a candidate describing how they would check a confident claim, which is useful and far better than an unstructured chat. It cannot capture them checking one, because the act happens in the work rather than in the retelling, and rehearsed answers about AI habits are easy to produce now. Use the interview to probe the record a task produced, rather than in place of it.
References
- 1. Employment Tests and Selection Procedures eeoc.gov Supports the job-related-and-consistent-with-business-necessity standard, and lists work samples and simulations among covered selection procedures.
- 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR Part 1607 govinfo.gov Section 1607.14(C) requires a job analysis of important work behaviors, states that closer resemblance to work samples strengthens content validity, and excludes abilities learned on the job.
- 3. Automated Employment Decision Tools (Updated) rules.cityofnewyork.us The adopted Local Law 144 rule, effective July 5, 2023, carrying the bias-audit and candidate-notice obligations for automated employment decision tools.
- 4. 15-1252.00 - Software Developers onetonline.org Source for the 17-task occupational profile, including developing or directing software system testing or validation procedures.
- 5. 13-2051.00 - Financial and Investment Analysts onetonline.org Source for the 26-task occupational profile and its capitalization, company-assessment and remedy-recommendation tasks.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.