Screening
A Resume Stopped Being a Writing Sample
A resume drafted with a model is still evidence of two things, and neither of them is the writing. It is a list of claims a third party can check, and a statement of priority: what the person put first, what got three lines, what they left off. Read it for those and it still earns its place in the process. Read it as a sample of thinking or care and you are grading whichever model produced the bullets.
The takeThe generic-phrasing tell should be retired rather than refined. Uniform bullets identify a cheap tool and a hurried afternoon, and they carry nothing about the person's judgment. Every quality the tell keys on, plain phrasing, a template structure, tidy parallel verbs, belongs to somebody who never used a model: a second-language writer, a graduate working from a careers-office template, a person who last applied for a job a decade ago. Anyone confident they can read authorship off a page is measuring their own confidence.
Where Olive fits
Open a role and see what the work shows
A resume is a claim about the work and never the work itself. Olive returns the other half to the employer who bought it: one role-grounded assignment done with an AI assistant on the candidate's own clock, read by a human reviewer who writes six findings, each quoting the timestamped moment it rests on.
Rank your shortlistWhat Is a Resume Still Evidence Of?
Claims and priorities. A claim is any line with an issuer standing behind it, so somebody other than the candidate can settle it and there is a cost if it turns out false. A priority is the arrangement: what sits at the top, what got three lines, what is missing. Neither one is prose, which is why neither one moves when the prose is regenerated.
Three things it stopped being evidence of, and they are the three most screening advice still rests on:
- Writing ability. The document is now a collaboration by default, and there is no way to separate the halves.
- Care and attention to detail. A tidy layout and consistent tenses cost one prompt. An untidy one might mean a phone, a screen reader, or a plain-text export.
- Effort or interest. A tailored resume takes seconds, so tailoring no longer signals that anyone wanted the job in particular.
What replaces them is duller and more useful. A checkable claim gives the later stages something to act on. An ordering choice cannot be settled by anybody, and it still tells you what this person thinks their strongest work is, which is exactly what a first conversation should open on.
Why Grading the Prose Grades the Tool
Because writing is the task that got cheapest first. In a pre-registered experiment, 444 college-educated professionals completed occupation-specific writing tasks, and those given ChatGPT finished 37% faster with grades 0.45 standard deviations higher, with the largest gains going to the weakest writers 1. A document that compresses the distance between a strong writer and a weak one cannot be used to tell them apart.
Those were paid online tasks with no colleagues, no revision cycle and no consequences, and the published version of the same study reports the effects differently, so the decimals are not the part to carry. The direction is, and the direction is enough. Prose quality on a resume now measures access and habit.
The fallback, reading the document and forming a private view, measures less than it feels like it does. Asked to tell GPT-3 text from human writing across stories, news and recipes, untrained evaluators performed at chance, and three quick training methods lifted them only to about 55% 2. That was an older model generation, and the evaluators were crowdworkers rather than hiring managers reading in their own field. It sets a floor under the problem and settles nothing about any particular reader.
The cost of acting on the feeling anyway is not evenly distributed. Running detectors over 72 journal abstracts, one tool that was highly accurate on the clean human-versus-machine split over-flagged AI-assisted text from non-native authors at 25% against 11% for native authors 3. Small sample, academic abstracts, and detector versions from late 2024. The asymmetry is the part that transfers, and the mixed case, a person's own work polished by a model, is the case that actually arrives in a hiring inbox. When every document in the pile looks equally good, what to screen on once the resumes all look perfect is the practical version of this problem.
Read It as a Claims Index and a Priority Statement
Read it twice, for different things. First pass: mark every claim with an issuer behind it, so employers, dates, titles, degrees, licences, named products shipped. Second pass: read only the order and the space. What sits at the top, what got three lines, what is missing between two dates. The first pass produces a verification list and the second produces the questions worth asking.
The order is more informative than it looks, with one correction. A resume tailored to the posting will echo the posting's own priorities, so read the ordering against the job description rather than in isolation. Where the two match exactly, that is the posting talking. Where they differ, where someone led with the smaller project or gave the prestigious employer one line, a person made a choice and can explain it.
The omissions are the other half of the same reading. A gap, a role held for five months, a degree with no dates: none of those is a finding, and each is a good first question. The claims themselves go into a different pipeline entirely, because only some of what an application asserts can ever be checked, and knowing which is which stops a screen from spending its attention on the parts that will never resolve.
What this changes on Monday is small and specific. Stop scoring the summary paragraph. Stop counting quantified bullets as evidence of rigour, since a number is the easiest thing in the document to generate. Start writing one line per candidate naming the claim you would check and the question you would ask.
Which Questions Go to the Conversation?
The ones the document cannot settle, which is anything about judgment, ownership and choice. Send the checkable claims to verification, and carry everything else into a structured conversation where every candidate gets the same questions rated the same way. That format is worth roughly double what an unstructured chat about the same resume is worth, and the gap between two interview formats is larger than the gap between two methods 4.
Three question shapes do most of the work:
- Ownership. Which part of this did you decide, and what did you hand off. Works on any bullet, and it is a normal question rather than an accusation.
- Consequence. What broke, what did it cost, what changed afterwards. A resume rarely carries any of that, drafted by hand or not, so the answer has to come from the person.
- Rejection. Name something you produced and threw away. The answer is unavailable to anyone who was not there.
None of that requires knowing who typed the resume, which is the point. The three questions work identically on a hand-written document and a generated one, so the authorship question stops mattering the moment the conversation starts.
One thing worth checking separately is whether the screen itself is still earning its keep. If a large share of applicants is being cut on a document that now says less than it used to, finding out whether the resume screen is throwing away the wrong people matters more than tuning how it is read. The intake side has the same problem, and it is fixable in an afternoon: which application fields still produce signal determines what you are reading before any of this begins.
Common questions
Is an AI-written resume a reason to reject someone?
On its own, no. Rejecting for it means rejecting on a guess about authorship, and nothing in the document can support that guess. The question worth asking is whether the claims are true and whether the person can do the work, and neither of those is answered by how the document was drafted. Decide the rule in advance so it is not made case by case.
Do typos still mean anything?
Less than they did, and not what people think. A clean document costs one prompt, so an absence of typos says nothing about the person who submitted it. A typo tells you a line went untidied, which is a fact about the drafting and not about the candidate. If precision genuinely matters in the role, test it on work that resembles the role instead of on a formatting artifact.
How do you compare two resumes that look identical?
Stop comparing at the document level. Two equally polished resumes carrying different claims are different candidates, so compare the claims: what was shipped, over what period, with what scope. If the claims are also similar, the document has genuinely run out of information and the next step is a short piece of real work or a structured conversation, not a closer reading of the same page.
Should the resume screen still be the first stage?
It works as a cheap first pass on checkable claims, and it fails as a judgment about people. Keeping it means narrowing what it decides: eligibility, obvious mismatch of scope, whatever the role genuinely requires on paper. Anything past that belongs to a stage with more evidence in it. Where volume is high and the roles are uniform, put the first real decision on a short piece of work instead.
What should candidates be told about how their resume is read?
Tell them plainly, in the posting. Say that the resume is read for claims and priorities rather than as a writing sample, that AI is permitted on it, and what the next stage actually assesses. It cuts the guessing that produces both dishonest documents and self-eliminating applicants, and it costs one paragraph in a template you already maintain.
References
- 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) economics.mit.edu Supports the claim that written output no longer separates strong writers from weak ones: 444 professionals, 37% faster with 0.45 standard deviations higher grades, largest gains to the weakest writers.
- 2. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that reading a document and forming a private view about authorship is unreliable: untrained evaluators performed at chance, and brief training lifted them only to about 55%.
- 3. The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication pmc.ncbi.nlm.nih.gov Supports the claim that the cost of guessing at authorship falls unevenly: over-flagging of AI-assisted text ran at 25% for non-native authors against 11% for native authors, on the mixed case that hiring actually sees.
- 4. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the claim that adding structure roughly doubles what an interview predicts, at .42 against .19, which is why resume questions belong in a structured conversation.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.