Screening
What Do You Screen On When Every Resume Looks Perfect?
When AI polish makes every resume look perfect, screen on facts someone outside the candidate recorded and dated: commit histories, published bylines, license and registration numbers, filings. Polish can't touch those. Prose quality and achievement bullets never proved much and now prove nothing. Specifics you can falsify, like a quota a reference can confirm, move to a ten-minute call. Whole fields leave no public trail, so consulting, product, marketing and operations roles need a work sample earlier. Don't screen for what sounds AI-written. Nobody can tell.
The takeThe resume was never as good as hiring treated it. It was a writing test in disguise, and it rewarded whoever could produce the tidier account of the same career. That unevenness is what quietly did the sorting. Models didn't break the screen; they removed the last thing hiding what it was. So the current panic looks to me like a good outcome arriving at a bad moment: hiring now has to pay for evidence it used to get free. I'd expect the teams resenting this most are the ones whose screen was doing the least work.
Where Olive fits
Open a role and see what the work shows
A resume can only report how someone works with AI. Olive puts the work in front of you instead: a 40-to-60-minute assignment in the candidate's own occupation, done with an AI assistant, returned as six findings a reviewer wrote by hand with the moment behind each one, and the candidate is granted the same report.
Rank your shortlistWhat still counts as evidence on a resume?
Three things: records a third party holds and has dated, specifics that are falsifiable in a phone call, and nothing else. A commit history, a published byline, a license number, a court docket, a filing: those existed before the candidate wrote the resume and can be opened without asking permission. Everything written in the candidate's own voice is now a draft a model tightened.
That splits an inbox into three tiers, and the tiers are worth writing down before you read another application.
- Externally held and dated. A repository with two years of small commits and review comments. Bylines with publication dates. A state license or registration number. A patent, a filed prospectus, a docket, a published paper, a package other people depend on. You can check these yourself, in a browser, in under a minute each.
- Falsifiable in a call. Team size, the name of the system, the quarter, the vendor, the number and where it came from, who signed off. A model writes these as easily as anything else, so they are not evidence on the page. They are specific enough to be wrong, though, which makes them the material for a screen call.
- Unfalsifiable. "Led the workstream." "Drove a 30% lift." "Partnered cross-functionally to align stakeholders." These were never evidence. What changed is that they used to be unevenly written, and the unevenness was quietly doing the sorting. Now every applicant clears that bar, so weighting it is weighting noise.
Tier three is where this gets awkward, because it is most of a resume and most of what an experienced candidate has to offer. A director of operations at a private company leaves no public trail at all. That is not a reason to distrust them. It is a reason to stop pretending the screen was measuring them, and to move the measurement somewhere it can actually happen.
What survives AI polish in your field?
It depends on whether your occupation leaves a public trail. Software engineering, journalism, academic research, licensed professions and anything filed with a regulator leave records you can open in a browser. Consulting, product management, marketing, operations and most general management work leave almost none. The deliverables sit behind an NDA, and the resume bullet was always the only account of them.
Concretely, per field:
- Software engineering. The repository is not the evidence; the history is. A model will produce a clean project in an afternoon, so read commit cadence over months, issue threads, review comments the candidate left on someone else's code, and whether anything they built has users who are not them. Who actually built the portfolio project is a separate question from whether it runs.
- Journalism, research, design. Bylines, papers, credits and case studies carry dates and editors. Check that the dates cluster the way a career does rather than the way a portfolio assembled last month does.
- Licensed and regulated roles. Nursing, accounting, law, insurance, real estate, medicine. The license lookup is free, public and definitive, and so is a disciplinary record. It is the cheapest verification in hiring, and most inboxes leave it until offer stage.
- Finance and analysis. Registrations, published notes, named transactions, and the model itself. Ask for a workbook and read the formulas rather than the summary tab.
- Consulting, product, marketing, operations, sales. Nothing survives, because nothing was ever public. Sales is the partial exception: quota, territory, ramp and attainment are specific enough that a reference either confirms them or hesitates.
For that last group the honest conclusion is that the resume screen has no evidence left to work with, and the process has to compensate earlier rather than read harder. Screening for judgment in a non-technical role starts from that constraint instead of arguing with it.
Should you screen out resumes that look AI-written?
No, and you can't anyway. Across six experiments and 4,600 participants, people identified whether a professional self-presentation was written by a person or by a language model with 50 to 52% accuracy, which is chance, and it stayed at chance after training and after payment for getting it right 1. A hiring manager who says a resume sounds like AI is reporting a hunch that has been measured and found empty.
Automated detectors do not rescue it. Seven detectors flagged an average of 61.3% of TOEFL essays written by non-native English speakers as AI-generated, and 97.8% of those essays were flagged by at least one of them, while essays by US-born students were classified accurately 2. That is not a tuning problem. Simpler vocabulary and shorter sentences read as machine-like to these tools, so the error lands on the same people every time.
Ethics is one edge of this; compliance is the other. A rule you apply to decide who advances is a selection procedure, and if it disproportionately excludes people on a protected basis you have to show it is job-related and consistent with business necessity, an obligation you carry even when a vendor supplied the tool 3. "It read like ChatGPT to me" is not that showing. Whether detectors work in hiring at all has a short answer, and it is this one.
How do you verify a resume claim in ten minutes?
Pick the two bullets the hire actually turns on and ask for the facts around them: what the constraint was, who disagreed, what got cut, and what happened afterward. A person who did the work answers in specifics and volunteers the parts that went badly. Someone reciting a generated bullet gives you a second version of the bullet.
The call runs ten minutes and has three moves.
1. Name the claim. "Your resume says close time went from twelve days to five. Walk me through the first week." 2. Ask for the part that went wrong. Real projects have a fight in them: a vendor that missed, a number that would not reconcile, a stakeholder who blocked it. Someone who was there names one without being prompted. 3. Ask what they would keep and what they would not. This is the closest a phone call gets to judgment, and it is the move a rehearsed answer handles worst.
Don't ask whether they used AI to write the resume. Assume they did, the way you assume they used spellcheck, and ask about the work instead. When the claim in question is AI fluency itself, the same call works with different questions, and verifying "uses AI daily" in ten minutes is that version of it.
Reference calls carry the weight they always did, which is less than people expect. Their remaining use is narrow and real: a reference who cannot confirm the scope of what the candidate claims is telling you something.
What replaces the resume screen?
A shorter filter and an earlier work sample. Cut the resume screen back to hard requirements you can check (the license, the work authorization, the years in a named system), and stop reading the rest for signal it no longer carries. Then spend the time you saved on a work sample, where the tasks resemble the job's tasks closely enough that performance on one tracks performance on the other 4.
That trade is usually affordable. An hour spent reading fifty resumes for tone is the same hour a structured sample takes on the eight candidates who cleared a checkable filter. The sample also behaves better on fairness: work samples generally show little or no performance difference between men and women or across racial groups, and candidates tend to accept them as fair because the connection to the job is visible 4.
Two cautions. A take-home a model finishes in thirty seconds measures nothing, so the task has to allow AI openly and be built so that using it well is the thing being watched. And length costs you people: a two-hour task loses fewer candidates than a full day, and a full day loses the ones who are employed and busy, which is most of the ones you want.
If your inbox got worse because a keyword filter stopped separating anyone, that filter was doing the same job as the tone screen and broke for the same reason, which is that every resume now matches. The replacement is not a better filter. It is a smaller filter and a real task. Where volume allows, some teams drop the resume screen entirely and assess everyone who meets the hard requirements, which is cheaper than it sounds once the reading stops.
Common questions
Is a GitHub profile or a portfolio still worth checking?
Yes, but read the history rather than the artifact. A polished project proves an afternoon with a model. Two years of small commits, issues answered, and reviews left on other people's code prove a working habit. Check whether anyone besides the candidate depends on the thing. For designers and writers, the equivalent is dated, credited, published work rather than a case study assembled last week.
Should you still ask for a cover letter?
Only if you replace it with something checkable. A free-text letter costs the candidate nothing now and tells you nothing. A short structured question works better: one paragraph on a specific decision they made, with the constraint and the tradeoff named. It is harder to fake usefully and easier to compare across applicants, because everyone is answering the same prompt rather than performing enthusiasm.
Is 'this sounds like AI' a defensible reason to reject someone?
No. People perform at chance when asked to tell AI-written professional text from human-written text 1, and detectors misfire heavily on writing by non-native English speakers 2. Whatever rule you use to decide who advances is a selection procedure: if it screens out people on a protected basis at a higher rate, you have to show it is job-related and consistent with business necessity 3. A hunch about tone cannot carry that.
What do you screen on when the role leaves no public record?
The checkable facts (scope, systems, dates, who they reported to), and then move the real evaluation forward. Consulting, product, marketing and operations candidates have nothing you can open in a browser, so the screen's job shrinks to separating people who meet hard requirements from people who don't. The judgment gets measured in a work sample or a structured call, not in the inbox.
Do years-of-experience filters still work?
Keep them only where the number stands for something specific: a regulated qualification, or a system nobody learns in a month. As a general proxy it was always weak, and it now selects for whoever read the posting most carefully and wrote to it. If the requirement matters, name the thing it is a proxy for and check that instead.
References
- 1. Human heuristics for AI-generated language are flawed ✓ pmc.ncbi.nlm.nih.gov Across six experiments with 4,600 participants, people identified the source of professional self-presentations with 50 to 52% accuracy.
- 2. GPT detectors are biased against non-native English writers ✓ pmc.ncbi.nlm.nih.gov Detectors flagged 61.3% of non-native TOEFL essays as AI-generated on average; 97.8% were flagged by at least one detector.
- 3. Employment Tests and Selection Procedures ✓ eeoc.gov A selection procedure with disparate impact must be job-related and consistent with business necessity, and the employer carries that obligation for vendor-supplied tools.
- 4. Assessment and Selection: Work Samples and Simulations ✓ opm.gov Work samples carry high content and criterion-related validity, and generally show little or no performance difference by sex or race.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.