Pipeline
A Procedure Is Biased When Decisions Move on Job-Irrelevant Facts
A hiring process is biased when its decisions move with a fact about the candidate that is not related to the job, whether or not anyone meant them to. That makes bias a property you can test rather than a motive you have to allege. The unit is the selection procedure, meaning any measure used as a basis for an employment decision, and unlike a state of mind it has three inspectable parts: the inputs it reads, the rule it applies, and the record it leaves.
The takeThe person-level definition survives because it is comfortable. It puts the fault inside a head, where nobody can inspect it six months later, and it prescribes a remedy that leaves the procedure exactly as it was. Ask what the stage wrote down instead. A team that cannot reconstruct why ten people were rejected has not got a bias problem it can argue about, it has a stage that cannot be examined at all, and that is the cheaper of the two to fix.
Where Olive fits
Open a role and see what the work shows
A number standing for a person cannot be interrogated, which is what makes a rejection hard to explain and a stage hard to inspect. Olive returns six findings written by a human reviewer, each carrying the timestamped excerpt from the session it rests on, and no composite figure at all.
Rank your shortlistWhat makes a procedure biased rather than a person?
A procedure is biased when its decisions move with a fact about the candidate that is not related to the job. Intent is not part of that definition. The Supreme Court's 1971 decision in Griggs held that Title VII reaches "not only overt discrimination but also practices that are fair in form, but discriminatory in operation," and that the touchstone is business necessity 2.
Federal selection law gives the object a name. A selection procedure is "any measure, combination of measures, or procedure used as a basis for any employment decision," and the 1978 Uniform Guidelines draw that range widely enough to include "informal or casual interviews and unscored application forms" alongside tests and work samples 1. The coffee chat with the hiring manager is a selection procedure. So is the glance somebody gives a resume before setting it down.
That matters because a procedure, unlike a mental state, has parts you can put on a table:
- Inputs. Every field the stage actually consumes: what is on the screen, what was said in the room, what the tool was fed.
- The rule. What has to be true for a candidate to advance, and who applies it.
- The record. What the stage wrote down, and whether a stranger could rebuild the decision from it.
The legal test runs on those same parts. Title VII, as amended by the Civil Rights Act of 1991, says that once a complaining party shows a particular practice causes a disparate impact, the employer has to demonstrate the practice is "job related for the position in question and consistent with business necessity" 3. That is a claim about a rule, and somebody has to produce the rule to make it. This is public legal fact rather than advice, and what applies to your process is a question for counsel.
Why is the list of named biases a dead end?
Because a mental state is not discoverable at month six and a rule is. Affinity, halo, anchoring and confirmation describe real ways a person goes wrong, and every one of them is invisible in the only artifact that survives a hiring decision. When a rejected candidate asks why, nobody can produce an interviewer's state of mind.
The panel's rule and its notes are producible, which is why the remedy that follows from the person-level definition so often misses: training changes no input, no rule and no record. Meanwhile the stage practitioners trust most is the one that can carry the least. Highhouse's review reports that the interrater reliability of the traditional unstructured interview is so low that even with a perfectly reliable and valid criterion, interview-based judgments could never account for more than 10% of the variance in job performance, while a 1990s survey of 201 HR executives rated that same unstructured interview, in their own perception, more effective than any paper-and-pencil procedure 5. Both numbers are his summary of other people's work. The survey measured belief about effectiveness. The 10% is a ceiling that follows from how little two interviewers agree, so read it as the most such judgments could explain.
The instability turns up in ordinary scorecard data too. Across interviews with more than one interviewer in a large applicant-tracking dataset, around 38% of scorecard pairs carried at least a one-point difference, and nearly half of those one-point gaps fell between 2 and 3, crossing the yes-or-no line on a four-point scale 4. The dataset records disagreement and nothing more: it holds no outcome measure and no demographic variable, so it says nothing about who was right or about any protected group. What it does say is that the number is least stable exactly where it decides.
Which is the practical case for evidence over ratings. Nobody can interrogate a 3. A sentence saying the candidate could not name what she would check before sending the figure out is either accurate or it is not, and the person who wrote it can be asked. The same failure runs the other way when hiring managers reject a candidate for sounding like AI: the input doing the work is unstated, so nothing about it can be checked.
Run the reconstruction test on ten past decisions
Take ten decisions from one stage in the last quarter and hand the written record to somebody who was not in the room. If they can say why each candidate advanced or did not, using only what the process wrote down, the stage is inspectable and you can start arguing about whether those reasons were job-related. If they cannot, you have found something worse than bias.
A stage whose only record is a number, a thumbs-down or a blank is unfalsifiable in both directions. Nobody can show it is biased and nobody can show it is clean, so the question never closes and the same argument returns next quarter with new numbers and no new evidence. That condition is more common than bias and much cheaper to repair, because repairing it costs a form field rather than a culture.
Here is the version that fits in a Monday:
1. Pick the stage that rejects the most people. In most funnels this is application review, which rarely appears on the funnel diagram because nobody schedules it. 2. List every field that stage consumes per candidate. Work from what is actually on the screen, including what a reviewer reads and no form captures. The documented list is usually shorter. 3. Write a job-related reason beside each field, in one sentence, or leave it blank. The blanks are the finding. 4. Cut the blanks, or write down why you kept them. A field with no stated reason is a field the decision rides on for reasons nobody has to defend. 5. Add one required sentence to the stage's record: what this candidate did or did not show, against which requirement.
Step five is the one that gets dropped and the one that makes every later question answerable. It is also what stands behind a rejection when somebody asks for a reason, which is where what a candidate is owed when an AI screen turned them down begins: with whether a reason exists in writing at all.
Which stage should you open first?
The stage that rejects the most people and writes down the least. Those two properties multiply: volume sets the cost of a bad rule, and thinness of record sets how long the rule can stay wrong without anyone noticing. Reach for the debrief last, even though the debrief is where the argument usually starts, because a debrief inherits its candidates from every decision upstream of it.
That inheritance is why a gap you can see is often not a gap that stage made. Which stage a gap actually opens in is a separate calculation, and getting it wrong is expensive in a specific way: teams retrain interviewers for something created earlier in the funnel, then measure the training against a number it was never able to move. A rate difference on its own does not settle the attribution either, which is what an outcome gap does and does not license.
A last caution about the word. Calling a procedure biased is a design claim you own and can act on this week. Calling it unlawful is a different claim with its own elements, jurisdictions and dates, and it belongs in front of counsel. The two get run together constantly, usually by someone who wants the first claim to carry the second one's weight. Keep them apart and the design work can proceed at its own speed while the legal question takes as long as it takes.
Common questions
Can a process be biased if everyone in it means well?
Yes. Intent is not part of what makes a procedure biased, and US adverse-impact analysis has not required it since Griggs in 1971. A rule can move decisions on something unrelated to the job while every person applying it acts in good faith, which is the case the phrase about practices fair in form but discriminatory in operation was written for. That is good news operationally: a fault in a rule is something a team can find and change this quarter, and finding it does not require anybody to be accused of anything.
Is an unstructured interview a selection procedure?
Yes, and it surprises people. The 1978 federal Uniform Guidelines define a selection procedure as any measure used as a basis for an employment decision, and say the range runs from paper-and-pencil tests through informal or casual interviews and unscored application forms. Swapping a scored assessment for a conversation does not move the step outside federal selection law. It moves the step outside your own record, which is a different problem and usually a worse one.
Does unconscious bias training fix a biased procedure?
Judge it by whether anything about the procedure is different afterwards. The inputs the stage reads, the rule it applies and the record it leaves are all the same the week after the session as the week before, so none of the three parts a biased procedure is made of has moved. Training may still be worth running for other reasons. What it bought here was a feeling of progress rather than progress.
What if a stage keeps no record at all?
Then that is the finding, and it outranks the bias question. A stage with no record cannot be shown biased or shown clean, so every argument about it is unresolvable and returns next quarter. Add one required sentence per candidate naming what they showed against which requirement. It costs a form field and it converts a permanent argument into a question with an answer.
How is a biased procedure different from unlawful discrimination?
Biased is a design claim: decisions move on something not related to the job. Unlawful is a legal conclusion with elements, jurisdictions and dates attached, and under Title VII, as amended by the Civil Rights Act of 1991, it turns on whether the employer can show the practice is job related for the position in question and consistent with business necessity. A procedure can be badly designed without being unlawful, and a lawful one can still be worth fixing. Fix the design on your own authority and take the legal question to counsel.
References
- 1. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) govinfo.gov Supports the definition of a selection procedure as any measure used as a basis for an employment decision, expressly including informal or casual interviews and unscored application forms.
- 2. Griggs v. Duke Power Co., 401 U.S. 424 (1971) law.cornell.edu Supports the claim that intent is not required and that a hiring device has to be tied to job performance rather than to the person in the abstract.
- 3. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases uscode.house.gov Supports the statutory test quoted here: job related for the position in question and consistent with business necessity.
- 4. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the scorecard-pair disagreement figure used to argue that a rating is least stable at the point where it decides.
- 5. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the reliability ceiling on unstructured interviewing and the survey finding that practitioners rated it the most effective method.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.