Every résumé perfectly matches the role. The interviews pass with flying colors. The task of choosing seems impossible. A month later, the hire grades as a bust.
Upon closer inspection, the cover letter, the take-home, and the portfolio also read like AI. Every artifact a candidate hands over may tell you nothing useful.
What’s missing in today’s evaluation process? Assume everyone uses AI, and the vetting changes from skill to judgment. What they question, check, and take on faith.
At Olive, we evaluate candidates on how they apply judgment to guide AI.
I
The BROKEN AI SKILL requirement
ZipRecruiter’s economists survey more than 1,000 verified US hiring managers and talent-acquisition professionals. 74% call AI skills a strong advantage or a flat requirement.
The demand is not new. In a Microsoft and LinkedIn survey from February and March 2024, 66% of leaders already say they would not hire someone without AI skills, and 71% said they take a less experienced candidate with AI over a more experienced one without. Two years ago before recent findings emerge (on the next screen).
The problem: employers fail to check for the skills they want. The disconnect results in bad hires, failed AI projects, and wasted investment. The gap is observable.
TestGorilla’s survey of 2,000 senior hiring leaders in the US and UK. Each square is one in a hundred — touch one.
95 of every 100 say AI fluency is now a formal hiring requirement, written into the hiring process.
Just 26 of those 100 make a candidate demonstrate AI use and verify the result. 19 leave the whole judgment to whichever hiring manager happens to be in the room.
An independent survey of 1,000+ US hiring managers, June 2026, counts lower: 74 call it an advantage or a requirement, 13 require it for every role.
Different surveys, same shape: wanting displaces checking.
ZipRecruiter’s own report says employers are changing what they screen for and automating the postings. The results still say nothing about how to verify the skill ask before an offer goes out.
II
The FLAWED listing
Olive measures every active posting in its own corpus — 2,151,213 of them. 34,743 carry an AI-usage requirement specific enough for the skill extractor to name: only 1.62% show enough detail to define a skill.
Indeed’s economists run a version of the same test on the wider market. When they put a full year of AI-mentioning postings through a topic model, one in four gives little indication of how AI may surface in the role. Specific language about the work remains rare.
Uniquely, Olive counts the same corpus two ways — one reading retrieves what the extractor found, the other reading directly pulls the raw text. A mistake in one cannot hide inside the other. See below.
2,151,213 active postings on August 11, 2026. Six months, February to July, counts the same day, two way.
The floor identifies an AI-usage requirement. The ceiling matches phrases in the raw text directly, so the floor avoids a drift with the extractor’s own mistakes.
3.04× on the floor, 2.15× on the ceiling — five to six months, February to July.
Between 1.40% and 8.12% of postings now carry the requirement.
Hiring managers begin to quickly catch on to the updated AI listing needs.
Two confident facts point in the same direction. Postings in human resources that mention AI double across 2025, climbing from 4.4% to 8.8%. Nationally, postings asking for AI skills grew 144% in a year while postings overall grew just 7%. Olive finds that the listings predict the drift before the process confirms.
III
The other side of the table
Candidates remain ahead by doing the same thing employers ask about. In Gartner’s survey of 3,290 applicants, from late 2024, 39% say they use AI somewhere in the application process. Capterra’s survey of 2,997 candidates across twelve countries, run separately, puts AI use somewhere in the search at 58%. And 59% of job seekers, by Resume Genius in March 2026, use AI to write their résumé. Today, AI saturation seems obvious and inevitable.
Hiring managers describe what they see when the same kind of résumé lands on the desk again and again. 69% say résumés read perfectly but more generic than they did five years ago. 77% say many now look AI-generated, and 39% name perfect grammar with no variation as the tell. 86% say AI makes it too easy to embellish, 42% call the trend a serious hiring risk, and 34% report résumés often or always fail to match the skills behind them. Résumés carry less and less value for information.
IV
The CURRENT PROCESS that does not survive
The obvious fix — grade what a candidate hands in, and let the artifact speak for itself. The next record chronicles when a vendor tries exactly such a plan, on its own AI-run interviews, at scale.
Fabric’s analysis of 19,368 AI-led interviews. Each mark is one interview.
38.5% of candidates flag for AI assistance — 48% in technical roles, 12% in sales.
61.1% of the flagged might pass anyway.
The point goes beyond cheating. The test measures a candidate’s artifact after the fact, and once AI produces it too, the artifact stops telling anyone who did the work.
The statistical version finds cheating useless. Among 5,179 customer-support agents, AI raises output 34% for the least experienced and barely moved the most experienced — the gap between them narrows. Among 453 professionals in a study published in Science, AI raises output quality 18% and narrows the spread between workers. A test that grades just the artifact loses its power to tell people apart.
On a task squarely inside what a model handles well, the gain can be enormous: one Copilot trial finishes a task 55.8% faster. The honest word here becomes jagged, not better or worse.
The fundamental question becomes not whether employees will use AI at work.
Can the hiring process find candidates best capable of leveraging AI for productivity?
“Professionals who had a negative performance when using AI tended to blindly adopt its output and interrogate it less.”— Dell’Acqua et al., working paper, 758 consultants
What separates the helped from the harmed was not prompting. The advantage surfaces when the person interrogates what came back — inquires of the LLM on a fine point, checks a figure, or refuses an answer that merely sounded right. The secret identifies a behavior, not a skill with the tool. And Olive excels at recording behavior.
V
What Olive ASSESSES instead
What Olive reads is the judgment a person applies while working with AI, not how well they prompt. The assessment rests on the candidate’s own reasoning, spoken (typed as an accommodation), as it happens; a task specific to one of twelve occupations; dimensions aim at how a claim gets checked rather than how an answer gets phrased; and evidence excerpts in place of a single number.
Others try pieces of our techniques. Canva’s engineers write that candidates unable to tell when a model was wrong struggle in their interviews. Zapier judges how a candidate pushes back on an output. CodeSignal names a skill it calls GenAI Limitation Awareness. Bryq names one such dimension in its assessment. Each names a different piece of the same judgment.
Eleven of twelve roles sit in the Bureau of Labor Statistics’ top exposure tier — 206 of 831 occupations. Software engineering, financial analysis, management consulting, data & analytics, product management, marketing, journalism, healthcare revenue cycle, legal operations, accounting & audit, underwriting & claims, supply chain.
97 cases sit across twelve roles with a high productivity potential with AI.
In 91 of them, we plant challenges or “occasions” — a real-world fact made wrong on purpose, the truth recorded right beside.
An occasion: problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification. 91 of each — 546 in total, laid down in sequence as the work goes on.
Not a quiz about AI. A job, with real places to check.
These places, within an hour with an AI colleague, receive inquiries during the work rather than placed afterwards. The next section shows what that hour looks like, from the opening document to the written debrief.
VI
How OLIVE works
The session begins in a browser, one real task from the role’s case set, a packet of documents, and a provided (or bring your own) AI assistant in the same workspace. The candidate thinks aloud throughout. The clock runs 55 minutes for a loan case, 60 for an engineering ticket, with a hard cap of 75 either way. Olive, an AI colleague, joins partway through with scripted questions about approach and workflow, spoken and shown in text at once. A written debrief follows submission.
Olive and the voice exist to get the thinking out while it happens, not after the fact.
“Speaking and reasoning aloud is a legible record for assessment.”
Six dimensions:
- Problem framing — Did the first move go after understanding, or straight for output?
- Evidence sourcing — Did they demand evidence for the claim that mattered?
- Delegation boundary — What did they keep, and what did they hand over?
- Working structure — Did anything exist between the brief and the answer?
- Output rejection — Was anything the assistant produced refused, and on what grounds?
- Verification — Was anything tested against the world, and did the result change something?
The assessment records the commentary, against the case’s own answer key.
Separately and represented for context: prose quality, prompt syntax, tool trivia, completion speed, tone of voice, accent, pace or hesitation, and AI use. Think-aloud audio becomes available for a reviewer. Never captured: a camera, the screen outside the task, keystroke rhythm, a face, a room, an identity document.
VII
FINANCE CASE: A loan, on a Thursday
You are a credit analyst at a private credit fund. It has been asked to hold $110.0M of a senior term loan, with a second $20.0M line — the revolver — unused at close. 55 minutes, 75 at the hard cap, for a 600–900-word pre-read memo to the committee.
Seven documents sit in the packet: the tasking email, the audited accounts, two extracts from the seller’s pitch document — the CIM — the term sheet, the rate sheet, and a sector file. Six were built to carry one checkable wrong fact; the audited accounts plant nothing, since the other six get checked against them.
The seller’s pitch document puts Adjusted EBITDA — the yearly profit figure the loan is sized against — at $30.0M.
The audited notes count in thousands: $1,400 there is $1.4M here. Note 8 — a footnote to the accounts — says that sum is already inside the $2.3M line the pitch document adds back again. Counted twice.
Take the $1.4M out once and senior leverage — debt divided by that profit — moves from 3.7x to 3.85x. The fund’s own policy tests 4.00x. Counting just the $110.0M actually drawn, both readings clear it.
The answer didn’t change. Whether it was checked did.
From the same role: a distributor is offered a competitor for $95M. The seller’s deck says $11M a year in savings pays for it in under four years. In a second case, a different colleague has a different number.
A second case, same role: “Friday, and a number.” Nothing planted. Here the colleague also offers remarks — one useful, one wrong, neither labelled.
“For the write-up, the deck’s payback line holds — eleven a year against the ninety-five clears in just under four, so that part at least can be taken as read.”
$95M over $11M is 8.6 years.
Whether that division ever happens becomes visible in the record, and the written view for the CFO concurrently changes.
A confident number marks a spot. Someone checks, or nobody did.
A concentration claim cites Note 14. The accounts’ notes run 1 through 11; it does not exist. Problem framing: what would make the answer wrong.
Day-count — the rule for counting interest days — is misstated as 30/360; this loan runs on actual days. Delegation boundary: the assistant, asked, repeats it.
Two Term SOFR readings exist — the seller’s 3.72%, already worked into a $9.9M interest line, and the desk’s own sheet at 4.23%. Working structure: something between the brief and the answer.
But the policy never says whether the $20.0M revolver counts — 3.85x if not, 4.55x if so. Output rejection: refusing an answer the packet leaves assumed.
The sector file’s benchmark leverage describes a distributor’s economics, not this manufacturer’s — a different capital intensity entirely. Verification: tested against the world, and the result changed something.
VIII
A SOFTWARE ENGINEERING CASE
You serve as a software engineer at a parcel carrier. A tool prints each depot’s end-of-day delivery sheet. The ticket adds failed deliveries, the re-attempt fee — the fee for trying a delivery again — and refunds. The engineer who owns this left recently. 60 minutes, 75 at the hard cap: working code, plus a written note to the operations lead.
Six items sit in the packet: the ticket, a departed engineer’s handover notes, two emails, a pay schedule, and the code itself. Every one of them carries a planted claim. Nothing here checks against a clean copy — in code, any file can be wrong.
The code bills $6.00 to try a delivery again, wired in from a January bulletin. The July pay schedule sets the fee at $4.50 and says it overrides anything in the code.
Weekend re-tries add a quarter of the fee on top of that. A quarter of $4.50 is 112.5 cents — a number with half a cent still to resolve.
A colleague’s rounding function — correct, already tested — turns 112.5 into 112. The pay schedule specifies half up: 113. One cent, and the whole point rides on it.
Correct code. Stale numbers. A single cent, and no error for any of it.
A second, separate case from the same role. An internal service fails four times a quarter, and each failure costs a finance person a working day — a finance-day. A vendor quotes $84K a year to replace it.
A second case, same role: “The service nobody owns.” Nothing planted. Here the colleague’s napkin also offers remarks — one useful, one wrong, neither labelled.
“Napkin version — four failures a quarter at a finance day each is real money, and on the recovered time alone the … quote pays for itself.”
16 finance-days a year — $10K to $25K.
One multiplication of the packet’s own numbers refutes the recovered-time argument. The vendor’s case has to rest on failure risk, engineering time, or roadmap cost instead — and whether that gets said is in the record.
Same failure. Same shape. In code this time.
The handover claims the system breaks above 2^31, ~$21M in cents; it is exact to 2^53−1 — too far apart for one axis, so no bar. Evidence sourcing: one line settles it.
The operations lead recommends a ready-made rounding package by name. It does not exist. Problem framing: a recommendation is itself a claim.
A date helper reads a bare date string as midnight UTC, so a depot’s Saturday reads as Friday. Delegation boundary: code the assistant reproduces, untested.
Both examples share a shape: a fluent number, a document or colleague behind it, one check that happens in the record or not. What comes next is what that record looks like.
IX
What the employer reads
Six dimensions, each with what a reviewer finds and the excerpts from the reading.
Each finding below comes from the record. Each excerpt quotes the moment and the source. The candidate reads the same page an employer does, at the same time.
Sample — authored. No session ran.
Financial analysis
“Six dimensions, each with what a reviewer found and the excerpts it was read from. There is no overall score.”
Problem framing
“Did the first move go after understanding, or straight for output?”
- Finding
- Sample finding — the citation to Note 14 was checked against the notes index, 1 through 11, and named unverified.
- 04:12 · capture
- “Sample excerpt — Note 14 isn’t in the notes. I can’t confirm that number.”
Evidence sourcing
“Did they demand evidence for the claim that mattered?”
- Finding
- Sample finding — the pitch document’s settlement line was traced to Note 8.
- 14:30 · AI panel
- “Sample excerpt — This adds back, landing EBITDA at 30.0.”
- 16:05 · submitted work
- “Sample excerpt — The settlement was counted twice.”
Delegation boundary
“What did they keep, and what did they hand over?”
- Finding
- Sample finding — the terms summary’s day-count line entered the memo as written. No check of it appears in this record.
Working structure
“Did anything exist between the brief and the answer?”
- Finding
- Nothing in this session spoke to this dimension.
Output rejection
“Was anything the assistant produced refused, and on what grounds?”
- Finding
- Sample finding — the assistant’s leverage figure was refused; the memo carries both revolver readings.
- 38:47 · spoken answer
- “Sample excerpt — The policy doesn’t say if the revolver counts, so I ran both.”
Verification
“Was anything tested against the world, and did the result change something?”
- Finding
- Sample finding — the sector benchmark was tested against the audited statements, changing the recommendation’s limit.
- 51:20 · debrief
- “Sample excerpt — That benchmark is for distributors, not this manufacturer.”
Candidates see this report too.
The single comparison this page allows is a property of the procedure. What an employer reads before anything else is a finding — the shape a reviewer wrote about a session. Reading a section and stopping tells you something about that session.
“A single figure would be quoted and compared across people never measured on the same work. Olive and and the six dimensions are not weighted against one another to produce it.”
Appendix
The receipts
Every figure on the page above, with the instrument behind it, what it does and does not license, and who profits from telling it.
The gap cited
TestGorilla’s survey of about 2,000 senior hiring leaders (US and UK, 56% senior decision-makers) sells AI skills assessments and publishes no fieldwork dates; its 95%, 26% and 59% figures all come from that one instrument. ZipRecruiter’s economists surveyed 1,000+ verified US hiring managers and talent-acquisition professionals, fieldwork 2026-06-11 to 2026-06-18, and sell no assessment. 19% leave assessment to hiring-manager discretion.
The specificity finding cited
Indeed’s economists, in their own words, found that “roughly a quarter” of a year’s AI-mentioning postings gave little indication of how AI would actually be used in the role — the page states this as one in four.
The corpus snapshot measured
2,151,213 active postings measured 2026-08-11; 34,743 carrying an AI-usage requirement the skill extractor could name. 34,743 ÷ 2,151,213 = 1.615%; /benchmarks and measured.js publish 1.62, which is used throughout.
The climb measured
Six monthly cohorts, February–July 2026, two instruments: extracted 0.46% → 1.40% (3.04×); raw text 3.78% → 8.12% (2.15×). August (45,683 postings) is excluded as incomplete against a ~998,166-posting monthly norm. 300 postings per arm remain to be hand-adjudicated before any single point estimate publishes.
Indeed’s series cited
AI_posting.csv, Indeed Hiring Lab, 2019-01-01 to 2026-07-31, 24,922 rows, 9 countries, re-downloaded and every statistic independently recomputed. US series: all-time high 6.285% (2026-07-31); fastest 91-day growth ever 1.264×; most recent 91 days 1.149×. The US series has never reached 8%.
Candidates cited
39% (Gartner, 4Q24, n=3,290, sells no hiring tool) · 58% (Capterra, n=2,997, 12 countries, July 2024) · 59% on the résumé specifically (Resume Genius, n=1,000 US, March 2026). Two different 59%s exist in this evidence: this one is job seekers using AI on the résumé; the other, in Act III, is organizations reporting a bad AI hire. The noun rides on the same line as the figure everywhere on this page. Resume Genius, the publisher behind that résumé figure, also supplies Act III’s “every résumé looks perfect” finding (C-9, its own 2026 Hiring Insights Report) and sells résumé-writing tools, profiting from both sides of the story it tells; Act III’s companion figures on that same beat (86%, 34%) are Express Employment Professionals/The Harris Poll’s (E-11), a staffing firm, not Resume Genius’s. 83% of hiring managers report some kind of hiring regret in the past year, AI or not (Resume Genius, n=1,500 US hiring managers, E-12) — the base rate against which the 59% bad-AI-hire figure in Act III must be read.
The interview wall cited
Fabric, which sells cheating detection, ran 19,368 AI-led interviews on its own platform through its own detector. 38.5% flagged, 61.1% of those flagged still scoring above the bar. No false-positive rate has been published for the detector; the population had already agreed to an AI interview; the 48%/12% technical/sales split is the detector’s own count. The lit shares on the canvas are drawn at these percentages, not separately counted.
Compression and mechanism cited
J-1: working paper (HBS/MIT, n=758 BCG consultants), pre-registered RCT, not peer-reviewed; the peer-reviewed version (Organization Science) was unreachable and no wording is attributed to it. J-2: METR, n=16 developers, an existence proof, not a population estimate. J-4: NBER working paper, published in the Quarterly Journal of Economics; n=5,179. J-5: Science 381(6654), n=453; cite the published figures (453/40%/18%), not the earlier SSRN draft’s (444/0.8SD/0.4SD). J-6: GitHub/Microsoft-run study of GitHub’s own product, n unstated, single greenfield task — vendor-run, and not evidence about judgment generally. J-7: PNAS, ~1,000 Turkish high-school students, field experiment. J-3: CHI 2025, Microsoft Research, n=319 knowledge workers, survey; its Evaluation coefficient is confirmed at full text and not printed. J-9: Psychological Bulletin 2011, 94 studies, ~3,500 participants; its think-aloud r is verified at the abstract level alone and not printed.
The typed path cited
V-11 is an allegation — a discrimination complaint filed with the Colorado Civil Rights Division and the EEOC, 2025-03-19 — not an adjudicated finding. V-12 measured ASR word-error rates of roughly 78% for deaf speech against roughly 18% for hearing speech, on 45 samples with a 2017-era API; no figure from this study is printed on the page. Thinking aloud is not free: V-1’s own experiment measured 125.1s versus 117.7s (p=0.008) for the think-aloud condition against silent work — a real cost the page’s clock absorbs without printing.
Capture measured
grep -rnE "video:[[:space:]]*true" src e2e tools content in judgment-gate-app on 2026-08-29 returns exactly one line across 1,516 TypeScript files: the comment at src/lib/capture.ts:14 asserting there are none. getUserMedia is called for audio alone (src/components/PushToTalk.tsx); getDisplayMedia is scoped to the current tab.
The roles cited
US Bureau of Labor Statistics, ai-exposure-categories.xlsx, 2026-08-27, 831 occupations in four tiers (Low 213 / Moderate 206 / High 206 / Very high 206). Eleven of the twelve roles sit in “Very high”; legal operations (paralegals and legal assistants) sits one tier down, in “High.” Product management is mapped to SOC 11-3021, Computer and Information Systems Managers — the closest available code, not a dedicated product-management classification. Instance depth across the twelve roles is uneven (12/12/12/12/12/12/10/3/3/3/3/3) because the five newest roles were added on liability and credential filters, not on exposure ranking. No one — Olive included — has published evidence that a judgment assessment like this one predicts job performance; Sackett, Zhang, Berry & Lievens’ corrected validity estimates belong to structured interviews and work samples, built over decades this product does not have. Bryq’s own named dimension (J-16) is not yet a verified row in this ledger; its row is pending.
The census measured
97 instance files under content/assessments/, all status: "live", bank version 2026Q3-002. 91 carry the seeded six-probe structure (546 probes, 91 per dimension); the other 6 are open-brief ground-truth cases with no probes array.
The adverse-impact audit cited
No four-fifths ratio has been published for this product. The audit has not been run because there is not yet enough completed-assessment volume across any protected-class subgroup for the ratio to be statistically meaningful.
The two instance files cited
judgment-gate-app/content/assessments/financial-analysis/instance-001.json and …/software-engineering/instance-001.json, both status: "live". Every run card and definition-list entry above cites one probe’s planted line from these files; the pass condition that decides whether a candidate’s response is graded correctly is never rendered on this page.
The two ground-truth cases measured
SIM (“Friday, and a number”) and SWE_SIM (“The service nobody owns”) are restated from judgment-gate-app/content/ground-truth/financial-analysis.json and …/software-engineering.json. No real candidate session is shown on chart D or chart E; the conversational lines quoted are the ground-truth case’s own scripted content, not a transcript. Chart E card 2 elides the vendor’s name.
The sample report sample
“A sample is a shape to read, not a field to rank” — an engineer’s source comment in src/lib/server/sample-data.ts:217-218, not a product string. Authored. No candidate record exists for fin-001. A released report on this case would also carry five collaboration traits beneath the six dimensions shown here; the sample renders just the six. Excerpts here are shortened for the page; a shipped report renders them whole. Verbatim strings and their sources: “Six dimensions, each with what a reviewer found and the excerpts it was read from. There is no overall score.” (reportShapeDek(false), traits.ts) · “Nothing in this session spoke to this dimension.” (DimensionSections.tsx:101) · “Candidates see this report too.” (ProfileTab.tsx:58).
The cost line cited
Shipped lowercase mid-sentence on olive.is (“ten attempts a month, no card”); the coda capitalises it and adds the period.