Pipeline

Where Did the Time Go After AI Entered Every Hiring Stage?

Time-to-hire went up in the year AI entered every stage, and the days are not in the tools. AI made producing an application nearly free and left judging one at its old cost, so screens that used to thin the pool now pass most people through. In high-volume roles the days pile up as queue before the first human contact. In scarce roles they sit in verification rounds added after a specific miss. Find the stage whose pass rate jumped, then rebuild or cut what sits below it.

The takeThe part that never makes the deck: those screens were never doing the work you credited them with. Effort was. For years an application cost a candidate an evening, and that toll did the sorting while the keyword filter took the applause. The toll is gone, and no product is bringing it back. So the reflex now, longer forms, tighter filters, a two-day window, is the expensive mistake of the next two years. Nobody has run that experiment cleanly, and the shape of it is plain: a little volume back, and the candidates with options gone first. That is the opposite of the trade you think you are making.

Where Olive fits

Open a role and see what the work shows

If a stage comes out of the loop, whatever takes its place has to earn those days back in evidence. Olive is priced per attempt rather than per seat and returns six evidenced findings on one candidate (an input to your decision, never a filter), with ten attempts a month free, so a replacement round can run beside your current one and be timed against it.

Rank your shortlist

Where did the extra days actually go?

Into the volume upstream and the checking downstream, not into the tools themselves. Generating an application now costs a candidate minutes, and reading one costs a reviewer what it always did. So every screen that used to thin the pool passes more people through, the stage below it inherits the load, and the days show up one step away from whatever broke.

The volume side is measured. LinkedIn's Economic Graph defines labor market tightness as the number of jobs members apply to divided by the number of applicants, and reports that job seekers now submit roughly twice as many applications as before the pandemic while the job-to-seeker ratio has returned near pre-pandemic levels 1. Nothing in that is candidates behaving badly. Someone applying to forty roles instead of eight is responding correctly to a market where each reply got less likely.

And the assistance works, which is why it spread. In a field experiment across nearly half a million jobseekers in an online labor market, jobseekers who received algorithmic writing assistance on their resumes were hired 8% more often 2. A tool that raises your odds gets used by everyone, and the pool it produces is not worse. It is more uniformly competent, which is a different and more expensive problem for whoever reads it.

Run the arithmetic on one stage. A resume screen that used to advance 12 of every 100 applicants and now advances 40, because the bottom of the pool reads like the middle, hands the round below it more than three times the interviews on the same calendar and the same headcount. That round got busier at the same speed, and busy shows up as wait time rather than as work. This is the same mechanism behind applications per opening tripling. The funnel stopped narrowing at the top.

Which stage absorbed the days in your kind of role?

It splits cleanly by how scarce the candidates are. In high-volume roles the days sit at the top, between the application arriving and the first human contact, as queue rather than as work. In scarce roles the days sit at the bottom, in rounds that were added to check work nobody is sure of any more. Same total, opposite ends, and the fix differs accordingly.

High-volume roles: campus, support, operations, retail management. The application-to-first-contact gap is the whole story. Your recruiters are not slower; they are reading four times the pile with the same week, and the pile no longer sorts itself by effort because effort is no longer what it costs. If your keyword filter now matches nearly everyone, that stage has stopped being a stage: what replaces ATS keyword screening is a separate decision and it should be made before you buy anything else.

Software engineering. The days moved into review. In Stack Overflow's 2025 developer survey the most-reported frustration with AI tools was "AI solutions that are almost right, but not quite," at 66%, and 45.2% said debugging AI-generated code is more time-consuming 3. A reviewer opening a take-home has exactly that problem, on a deadline, for a stranger. So a round that took forty minutes now takes ninety, and it takes ninety twice as often.

Legal operations, compliance, audit. The added days are verification rounds, and they are rational. A preregistered evaluation of the major legal research tools found Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time, against vendor marketing that had claimed hallucination-free citations 4. When the tool a candidate used carries a rate like that, checking one submission is real work, and someone decided to make it a stage.

Financial analysis and consulting. The reconciliation round. A memo that traces cleanly to sources takes an analyst twenty minutes to confirm and an afternoon to disprove, and the second case is now common enough that teams stopped skipping it.

Senior and scarce roles generally. An extra reference call, a second panel, a founder conversation that used to be a formality. Each of those was added by a person who could name a specific miss, which is why none of them come off easily.

Why does the calendar disagree with everyone's stopwatch?

Because per-task speedups are self-reported and calendar days are not. Everyone in the loop can truthfully say their part got faster while the loop gets longer, since nobody's stopwatch covers the queue between their part and the next one. Before you act on anyone's estimate (a recruiter's, a vendor's, your own), get timestamps.

The size of that gap is measurable. In a randomized trial by METR, 16 experienced open-source developers working 246 real issues in repositories they maintain took 19% longer to complete issues when AI tools were allowed. They had forecast a 24% speedup beforehand, and afterwards still believed AI had sped them up by about 20% 5. That is a roughly forty-point error in self-assessment among people measuring their own core craft. A recruiter's sense that screening takes half as long now deserves the same discount.

So instrument the loop, and keep two numbers per stage rather than one. Touch time is how long the work takes when someone is doing it. Wait time is how long the candidate sits between stages. Almost all of the increase lands in wait, and total time-to-hire hides that completely, which is why the metric that alarmed you is also the metric that cannot tell you anything.

Add a third number and the diagnosis finishes itself: pass rate per stage, this year against two years ago. A screen that used to advance 12% and now advances 40% has not become generous. It has stopped discriminating, and it is now a scheduling step wearing a filter's name. Rank every stage by that delta and the answer to "where is the time going" turns into a list with names on it.

Cut the stages that no longer separate anyone

Take the stage-by-stage pass rates and cut the ones that advance nearly everybody, then cut the ones two reviewers cannot score the same way. Adding a round is the reflex because it feels like more rigor; it is more calendar. Subtraction is the only move that returns days, and the stages that stopped discriminating are the ones it costs nothing to lose.

Start with the unstructured conversation. The revised meta-analytic estimates put structured interviews at .42 operational validity and unstructured interviews at .19, and the authors note that the predictors at the top of the list are the ones specific to individual jobs rather than general measures of a candidate's attributes 6. The culture chat that every candidate passes is costing you a week of scheduling to produce a number nobody records.

Cutting is also a compliance act, not only a speed one. Under EEOC guidance, a selection procedure that screens applicants for hire must be job-related and consistent with business necessity where it produces disparate impact, and the employer remains responsible for that validity even when a vendor supplied the tool and its documentation 7. A stage that advances 95% of the people who reach it still makes decisions about the other 5%, and you would have to defend those on the same terms as a stage that does real work.

Then spend the reclaimed days once rather than three times. One occupational task, done the way the job is actually done, produces more usable evidence than three conversations about it, and swapping a stage instead of adding an hour is the version of this that survives contact with a hiring manager who likes their round. Where two finalists now hand back the same quality of artifact, the separating evidence is in the process rather than the output, which is its own tiebreak problem. See how Olive measures this.

What does cutting a stage actually cost?

The bill comes in three parts, and none is the one executives usually fear. You lose your baseline, you lose a political ally, and you take on more variance in the first two months. What you almost never lose is quality, because a stage passing 95% was contributing almost nothing to the decision it appeared to be part of.

The baseline problem is the real one. Change two stages at once and you can no longer attribute the difference, so the next argument about the loop gets settled by whoever speaks with most confidence. Change one stage per quarter, keep the old rubric running in shadow on the same candidates for one cycle, and compare the same three numbers you started with: touch time, wait time, pass rate.

The political cost is worth naming out loud. Every round belongs to someone who can name a candidate it caught, and that story is real even when the aggregate says the round catches nobody. Bring the pass rate rather than the argument. "Your round advanced 47 of the last 49" ends the conversation more cleanly than any opinion about rigor, and it lets that person keep the part of their round that was doing the work.

And be honest about what a shorter loop cannot do. Cutting stages returns days; it does not tell you which candidate to hire, and a loop that got fast without getting more evidential has traded one problem for another. The test of the redesign is not the calendar. It is whether the round you kept produces something you could show a skeptical person six months later when the hire is either working out or not.

See a sample report

Common questions

Does adding AI to hiring actually make the process slower?

Not directly, but the second-order effect usually wins. AI cut the cost of producing an application far more than it cut the cost of judging one, so volume per opening climbs while review capacity stays flat. Screens that used to thin the pool pass most people through, the stage below inherits the load as wait time, and teams add verification rounds to compensate. Each of those decisions is locally correct and the sum is a longer loop. The tools are not the problem; the untouched stage structure around them is.

Which hiring stage should you cut first?

The one with the highest pass rate. A stage advancing more than about 90% of the people who reach it is a scheduling step, not a filter, and it costs calendar days in queue for a decision it is not making. After that, cut whatever two reviewers cannot score the same way twice. In most loops that is the unstructured culture conversation: revised meta-analytic estimates put structured interviews at .42 operational validity against .19 for unstructured ones. Pull the pass rates before the meeting so the discussion is about numbers rather than about rigor.

How do you find where the days are going in your own funnel?

Export timestamps per stage per candidate for the last two years, then split each stage into touch time and wait time. Touch time is the work; wait time is the candidate sitting between stages. Nearly all of the increase lands in wait, and total time-to-hire hides it. Add pass rate per stage, this year against two years ago. The stage whose pass rate jumped is the one that stopped discriminating, and the days will be sitting immediately below it.

Why does everyone say the tools are faster while the calendar says otherwise?

Because self-reported speedups are unreliable and nobody's stopwatch covers the queue. In a randomized trial, 16 experienced developers working 246 real issues in their own repositories took 19% longer with AI tools allowed, having forecast a 24% speedup and still believing afterwards they had been sped up about 20%. That is roughly a forty-point error among people measuring their own craft. Treat every stage-level speed claim as a hypothesis and settle it with timestamps.

Should you add a verification round instead of cutting one?

Only if it replaces a stage rather than joining the queue. Verification is genuinely where the work moved: 45.2% of developers in Stack Overflow's 2025 survey said debugging AI-generated code is more time-consuming, and the major legal research tools were measured hallucinating between 17% and 33% of the time. So the checking deserves a stage. It does not deserve an extra one. Put it in the slot vacated by the round whose output a model now produces in seconds, and keep the total number of stages flat or lower.

Does a shorter loop mean more bad hires?

Not from removing stages that pass almost everyone, since those were contributing little to the decision. The risk lives in what replaces them. A loop that got fast without getting more evidential has traded a calendar problem for a quality one, and you will not see it for two quarters. Guard against that by changing one stage per quarter, running the old rubric in shadow on the same candidates for one cycle, and comparing touch time, wait time and pass rate before and after.

References

  1. 1. Labor Market Tightness: LinkedIn's Measure of Job Competition LinkedIn Economic Graph, 2025. economicgraph.linkedin.com Tightness is measured as the number of jobs members apply to divided by the number of applicants; job seekers now submit roughly twice as many applications as before the pandemic while job-to-seeker ratios have returned near pre-pandemic levels. Page text verified 2026-08-25.
  2. 2. Generative AI at Work? Evidence from an Online Labor Market (NBER Working Paper 30886) National Bureau of Economic Research, 2023. nber.org Field experiment in an online labor market with nearly half a million jobseekers: treated jobseekers received algorithmic writing assistance on their resumes and were hired 8% more often. Abstract verified 2026-08-25.
  3. 3. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most-reported frustration with AI tools is "AI solutions that are almost right, but not quite" at 66%, and 45.2% of respondents report that debugging AI-generated code is more time-consuming. Page text verified 2026-08-25.
  4. 4. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Stanford RegLab and Institute for Human-Centered AI (arXiv 2405.20362), 2024. arxiv.org Preregistered evaluation finding Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinate between 17% and 33% of the time, against vendor claims of hallucination-free citations. Abstract verified 2026-08-25.
  5. 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized trial with 16 experienced open-source developers across 246 real issues in their own repositories: 19% longer to complete issues when AI tools were allowed, against a forecast 24% speedup and a post-study belief of a 20% speedup. Page text verified 2026-08-25.
  6. 6. Revisiting the Design of Selection Systems in Light of New Findings Regarding the Validity of Widely Used Predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. cambridge.org Revised operational validity estimates place structured interviews at .42 and unstructured interviews at .19, and the authors note the predictors at the top of the list are those specific to individual jobs rather than general measures. Page text verified 2026-08-25.
  7. 7. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Tools used to screen applicants for hire must be job-related and consistent with business necessity where they produce disparate impact, and while a vendor's validity documentation may help, the employer is still responsible for ensuring its tests are valid under UGESP. Guidance dated December 2007; page text verified 2026-08-25.

7 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.