Pipeline

Proving a Hiring Process Change Actually Worked

To prove a hiring process change improved anything, keep running the old process on part of the pipeline. Assign whole requisitions, or whole weeks, and never split two candidates competing for the same role. Hold the role family constant, and name the outcome before you start rather than after you see the numbers. If the outcome is quality of hire, you will wait two or three quarters for it, so declare an interim measure now and label it a proxy out loud, before it quietly becomes the headline.

The takeMost hiring process changes are never tested, and the honest reason is that testing one risks finding out. A control costs a few reqs running a version somebody has already decided is worse, which feels wasteful right up until the measurement comes back flat and saves a year of building on it. Firms in other functions run randomized trials on their own staff as a matter of course. Recruiting treats the same design as exotic, then argues about charts instead.

Where Olive fits

Open a role and see what the work shows

Every report Olive releases exports with its rubric, scorer and bank versions attached, which is what lets a later comparison say which version produced which finding. Ten attempts a month are free, so one arm of a split can run it while the other keeps the existing loop.

Rank your shortlist

Why a Before-and-After Chart Proves Nothing

Because everything else moved too. Between the two quarters on that chart the labor market changed, the req mix changed, the sourcing channels changed, two interviewers left, and the panel got better at the new questions simply by asking them fifty times. Any one of those can produce the improvement you are crediting to the change, and no amount of analysis afterwards can separate them.

Measuring a comparable effect properly takes a scale no hiring pilot reaches. Three company-run randomized trials at Microsoft, Accenture and an anonymous Fortune 100 firm had to be pooled across 4,867 developers before they produced a readable result, a 26.08% increase in completed tasks with a standard error of 10.3%, because each experiment on its own was too noisy to carry a finding 1. That is a randomized design with thousands of participants, measuring something as countable as tasks finished per week, and it still could not stand alone.

Set that beside a recruiting change evaluated on forty hires across two quarters, with no control and an outcome nobody defined in advance. The comparison is not close. It does not mean the change was worthless; it means the chart cannot tell you either way, and a chart that cannot tell you either way tends to get read as confirmation.

Run the Old Process on Some of the Reqs

Assign whole requisitions, never individual candidates. Two candidates for the same role getting different loops is a fairness problem, and the first question to put to counsel. Two reqs in the same family running different loops is ordinary variation in how a company works. Alternate by req as they open, or by week, and write the assignment rule down before the first req lands so nobody can choose sides case by case.

Hold constant what you can. Same role family, same level band, same sourcing channels, same recruiters where possible. If the new process is only running on engineering reqs and the old one on everything else, you have measured engineering. Where the volume is too thin to split within one family, alternate in short blocks across time, which at least mixes the seasonality through both arms.

Size it honestly before you start. Count how many reqs the family opens in a quarter, halve it, and ask whether the resulting number of hires could distinguish anything. Often it cannot, and knowing that in advance changes the plan in a useful way: extend the window, widen the families, or pick a nearer outcome. Discovering it afterwards produces an argument instead.

The pilot version of this is smaller and worth doing first. Piloting an assessment before making it a hiring gate runs the new stage alongside the existing loop without letting it decide anything, which gives you a clean read on what it would have changed.

Name the Outcome Before You Start

Write the outcome, the window and the decision rule in one paragraph, dated, and share it before any data exists. This is the single cheapest step in the whole method and the one most often skipped, because naming the outcome forecloses the move everyone makes later, which is to search the dashboard for whichever line went up and present that one.

Pre-registration is standard practice in the studies people quote at each other in these arguments. The Boston Consulting Group experiment on AI and consultant output was pre-registered across 758 consultants on 18 tasks, and its published result includes both directions: inside the tool's capability, output and quality rose sharply, and on one task chosen to sit outside it, the same group did worse 2. A study free to report only the tasks that worked would have published half of that.

Your paragraph needs four things. The primary outcome, stated as a number somebody else could compute from your systems. The window, in weeks, with a start date. The decision rule, saying what result would make you keep the change, revert it, or extend the test. And the exclusions, naming which reqs will not count and why. Anything you add to the analysis afterwards is exploratory, which is fine as long as it is labelled that way and does not become the headline.

One rule for reading the result: a difference you would not have predicted, in a measure you did not name, on a subgroup you did not plan to split, is a hypothesis rather than a finding.

Which Interim Measures Are Honest Proxies?

The ones that measure the mechanism you changed. If the change was a scored work sample, the honest interim measures are agreement between reviewers, how complete the evidence is on each scorecard, and candidate drop-off inside the new stage. Each of those tells you the stage is working as designed. None of them tells you the hires are better.

You need interim measures because the real outcome is slow. Median time-to-fill for nonexecutive roles ran 44 days in SHRM's 2021 benchmarking of its member organizations, with the slowest quarter at 73 days or more, counted from opening the req to an accepted offer 3. Add a start date, then a ramp, then a retention window matched to the role, and the first honest read on hire quality sits two or three quarters out.

So publish the proxy with its label attached. Write "interim proxy, not quality of hire" in the same sentence as the number, every time, including in the slide that goes to leadership. Proxies drift into headlines by repetition, one slide at a time, and the label is what stops it.

Two proxies are worth more than one because they fail differently. Reviewer agreement rises when a rubric is clearer and also when it is blander, so pair it with something that moves the other way under blandness, like the share of scorecards carrying a specific piece of evidence. Writing a rubric two reviewers score the same way is the work behind the first of those. And before you read any stage-level rate as a result, check that the denominator held still: which funnel metrics still mean anything is the same trap one level down. Once the window closes, measuring quality of hire on outcomes the manager does not own is what the interim numbers were standing in for.

See a sample report

Common questions

Is running two hiring processes at once legally risky?

Ask counsel about your jurisdiction, and the design detail that matters is the unit of assignment. Two candidates competing for the same role should face the same process. Two different requisitions running different processes is closer to how most companies already operate, since loops differ by team and by level anyway. Document the assignment rule, apply it mechanically, keep the records, and avoid any split that correlates with a protected characteristic.

What if the volume is too low to split the pipeline?

Alternate across time in short blocks rather than one long before-and-after, so seasonality and market shifts land in both arms. Widen the comparison from a single title to the whole role family. Or accept that what you have is a pilot, and set the goal accordingly: learning whether the stage is operable, how long it takes, and what candidates say about it are all real answers a small sample can give.

How long should the test run?

Long enough to accumulate the number of hires your outcome needs, which is usually longer than anyone wants. Work backwards: pick the outcome, estimate how many hires it takes to see a difference worth acting on, divide by the family's hiring rate. If the answer is more than a year, change the outcome to something nearer rather than shortening the window and reporting the near-miss as a result.

Can candidate feedback count as evidence the change worked?

As one outcome, named in advance, collected the same way from both arms. Candidate survey scores are quick and genuinely informative about the experience, and they are also sensitive to who bothers to answer, so report the response rate beside the score. What they cannot do is stand in for hire quality. Treat them as a measure of the process, which is a legitimate thing to try to improve on its own.

What if the test says the change made no difference?

That is a result, and usually a cheap one relative to the alternative. A flat measurement means the change is not paying for the time it costs, which frees that time for something else and stops a year of building on top of it. Record the verdict in writing with the design attached, so nobody rebuilds the same stage in eighteen months. Reverting a change you tested is a stronger position than keeping one you did not.

References

  1. 1. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers MIT Department of Economics (working paper; later Management Science), 2025. economics.mit.edu Supports the claim that even randomized workplace experiments are too noisy alone: three firm-run trials pooled across 4,867 developers for a 26.08% increase with a standard error of 10.3%.
  2. 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the pre-registration argument: 758 consultants on 18 tasks, with results reported in both directions including the task where the treatment group did worse.
  3. 3. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the claim that the real outcome is slow to arrive: median time-to-fill 44 days for nonexecutive roles, 75th percentile 73 days, req open to offer accepted, on data collected April to November 2021.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.