Screening
AI Skill Doesn't Survive the Resume Stage
You can't screen for AI skill from a resume, because a resume carries a claim about the skill and never the behavior behind it: the behavior is a sequence of decisions, and the document is only the outcome of one. So give the resume stage a single job: mark the claim you intend to test, then route everyone through the same short task rather than pre-filtering on who wrote the better paragraph about their AI workflow.
The takeEvery published answer to this question is a checklist: named tools, a prompt engineering certificate, a portfolio of AI projects. Two of those three are claims a model will write for anyone who asks, and the third certifies exposure to one vendor's console. Screening for AI skill on paper is not difficult, it is structurally impossible, and treating it as a hard problem awaiting a better checklist delays a funnel change that gets cheaper the earlier it happens.
Where Olive fits
Open a role and see what the work shows
Where a resume carries the claim, Olive puts the behavior in front of a reviewer: a role-grounded assignment done with an AI assistant that will do the whole thing if nobody stops it, returned as six findings a person writes, each anchored to a timestamped excerpt. Ten attempts a month are free, so it can run beside the round already in place.
Rank your shortlistWhat Can a Resume Tell You About AI Skill?
Which claims a candidate is prepared to make, and nothing about whether they hold up. A tool list, a certificate line and a project bullet are assertions, and each one is now cheaper to produce than to check. Read them the way you read a reference to a former employer: a place to point a question later, not a finding you can act on.
A claim with a task inside it ("rebuilt the weekly variance commentary with a model and now reviews it instead of writing it") names something a later stage can test. An inventory ("ChatGPT, Claude, Copilot, Midjourney, Notion AI") names nothing, because the same five words appear on most of the pile.
So the resume stage is doing routing work rather than evaluation work. It tells you the occupation, the seniority, whether the previous job was in AI-exposed work, and which one sentence is worth returning to. That is a useful stage. It is not a screen for skill, and asking it to be one is where the funnel breaks. When a candidate lists six AI tools the useful move is choosing which of the six to ask about, not counting them.
Two claims are worth marking on almost any resume for an AI-exposed role. The first is the task the candidate says moved, because that is where a follow-up question has purchase. The second is any number attached to it, since a number invites the question of how it was measured and who checked. Mark them and move on. The stage has done its job in under two minutes.
Why Certificates and Tool Lists Measure Exposure
Because exposure is what they are built to measure, and the exam guides say so plainly. The AWS Certified AI Practitioner names a target candidate with up to six months of exposure who uses but does not necessarily build AI solutions, and lists developing or coding models as out of scope 1. It is a recognition test under time pressure, with no performance task anywhere in it.
The other credential that shows up constantly is stranger still. Microsoft's Azure AI Fundamentals exam page asks for knowledge of Python syntax and programming techniques and familiarity with Azure resources, and the majority of the exam weight sits on implementing solutions inside one vendor's platform 2. A non-technical manager reading "AI Fundamentals" on a resume is not reading what they think they are reading.
Neither observation is an argument against certificates. Somebody who passed one sat down and learned a body of material, which is worth something. It is an argument against treating the line as evidence of judgment, because the exam did not ask for judgment and does not claim to have measured it. Whether a prompt engineering certificate beats a work sample is the sharper version of the same question.
Stop Pre-Filtering on the Claim
A pre-filter on an unverifiable claim does not remove weak candidates. It removes the ones who described themselves less fluently. Self-reported and tested AI skill barely move together: when 288 teachers took both a self-report and a knowledge test built on the same four-part framework, correlations between the two ran from 0.07 to 0.24 3. Sorting on the self-report sorts on writing.
That study is teachers in Taiwan, not job candidates, so the rate in it belongs to that cohort and stays there. What travels is the shape of the finding: the two instruments were built on one framework and still did not agree, and the profiles included people who underrated themselves as well as people who overrated themselves. A pre-filter loses both groups in the same pass.
The replacement is cheap. Set the requirement at the level of the work rather than the claim, let everyone who clears the occupational bar through to the same short task, and spend the screening effort on reading the task instead of ranking the paragraphs. Going straight from application to work sample is the fuller version of this move, and it turns on arithmetic: applicant count times review minutes against the hours the requisition actually has.
Test the Behavior the Claim Points At
The behavior worth testing is what a person does when the assistant is confidently wrong. In the BCG field experiment, on a task deliberately chosen to sit outside the model's competence, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer than the control group 4. Nobody in that condition could see which side of the line their task was on.
One sentence covers the skill, and it decomposes into moments you can watch: framing the problem before generating anything, asking for a source on the claim that matters, keeping the judgment that should not be handed over, throwing away output that looks finished, and checking a result against something outside the conversation. Each of those is visible in a forty-minute task and invisible on any document.
Researchers building an AI literacy assessment for a US Navy robotics programme reached the same design conclusion: a scenario task simulating AI use on the job outperformed the knowledge tests they had adopted or written, and prevailing assessments over-weight technical foundations against practical work 5. That is one programme's comparison rather than a published margin, so treat it as a design argument.
Run the task after a human has expressed interest rather than in front of an open posting, keep it inside an hour, and write the answer key first. Verifying an AI claim in ten minutes with no assessment budget is the low-cost floor if that is all the process can carry this quarter, and the two questions worth asking in a phone screen decide which claim the task should aim at.
A usable task looks ordinary. Hand the candidate a real artifact from the job with two defects planted in it, an assistant that will happily agree with whatever they propose, and a deliverable a colleague would actually receive. Then read for what they refused. The candidate who ships the assistant's second draft and the candidate who threw out the first three are both visible within ten minutes of reading.
Common questions
What actually counts as an AI skill?
A set of decisions rather than a set of tools: framing a task before generating, requiring a source for the claim that carries weight, keeping the judgment that should not be delegated, rejecting output that looks finished but is wrong, and checking a result against something outside the model. Tool names change every few months and these do not, which is why a requirement written in behaviors ages better than one written in product names.
Should the posting still ask for AI experience?
Ask for the work, not the years. "Has moved a recurring analysis task onto an assistant and can describe what changed" is checkable in a conversation and a task. "Three years of generative AI experience" is checkable against a calendar and excludes people who did the work last year. Write the requirement so that a candidate could tell whether they meet it.
Is a portfolio of AI projects better evidence than a certificate?
Slightly, and for one reason: a project has specifics a follow-up question can reach. It carries the same authorship problem as everything else on a resume, so treat it as a claim with more surface area rather than as proof. The useful question is not whether the project is impressive but what the candidate rejected while building it, and why.
Can a resume be checked for whether AI wrote it?
No, and building a stage on that idea creates a problem rather than solving one. Text-authorship tools are unreliable enough that a result cannot support a decision about a person, and a candidate who used a model to tighten their phrasing is doing what the posting will ask them to do at work. Move the evidence question to a task and the authorship question stops mattering.
Does this mean the resume screen should be dropped entirely?
Not necessarily. It still routes: occupation, seniority, location, work authorization, and the one claim worth returning to. What it cannot do is grade AI capability, so remove that job from it rather than removing the stage. Whether the task can then go to everyone is arithmetic: applicant count times review minutes against the hours the requisition has, which is where a high-volume req runs out first.
References
- 1. AWS Certified AI Practitioner (AIF-C01) Exam Guide, Version 1.4 d1.awsstatic.com Supports the claim that a foundational AI credential targets up to six months of exposure, expects the holder to use rather than build AI solutions, and puts model development out of scope.
- 2. Exam AI-901: Microsoft Azure AI Fundamentals learn.microsoft.com Supports the claim that a vendor 'AI Fundamentals' credential asks for Python knowledge and weights most of the exam on implementing inside one platform.
- 3. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arxiv.org Supports the claim that self-reported and tested AI literacy correlate weakly (0.07 to 0.24 across 288 teachers), so a self-description does not report capability.
- 4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the claim that on a task outside model capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer than the control group.
- 5. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arxiv.org Supports the claim that a realistic scenario task measured applied AI literacy better than the knowledge tests the same researchers adopted or built.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.