Screening
Report the Self-ID Rate Beside Every Ratio or the Ratio Is Unreadable
No response rate on the demographic questions makes a ratio trustworthy on its own, so bound the result instead of hunting for a threshold. Recompute the impact ratio twice: once counting every unknown applicant into one group, once into the other. If the verdict flips between those two runs, what you have is a collection problem. If it holds, report it with the response rate printed beside it, and treat that as a floor, because splitting the unknowns between groups opens the range wider still.
The takeA diversity tile that prints a percentage with no response rate on the same surface is not a measurement, and the software will not volunteer the missing number. The default is to compute over whoever answered, which quietly assumes the people who declined resemble the people who did not, and that assumption carries more weight than the analysis resting on it. Ask what the response rate was before you ask what the ratio was. If nobody can produce it, you already have your first finding.
Where Olive fits
Open a role and see what the work shows
A response rate is a measure of what candidates are willing to hand over, which is why Olive asks for one thing and gives one thing back: an occupational assignment done with an AI assistant, and six findings, each anchored to a timestamped excerpt from the session. The candidate is granted the identical report, free, on every tier, so there is no version of it they cannot read.
Rank your shortlistWhat response rate is high enough?
No threshold has been published, and any page that hands you one invented it. It depends on the size of the gap and on which way the unknowns could plausibly fall, and no fixed bar sees either. The federal Uniform Guidelines on Employee Selection Procedures (1978) permit impact records on a sample basis only where the sample is appropriate in terms of the applicant population and adequate in size 2. The people who chose to answer are not that sample.
The question you actually need answered is narrower than it looks. You are not trying to establish the true composition of your applicant pool. You are trying to work out whether one specific conclusion, that a gate treats groups differently, survives the uncertainty in the data. How much missing data a conclusion can absorb depends on how wide the gap is, which is why a fixed bar was never the instrument. The next section tests the conclusion directly.
Test the verdict against both bounds
Run the ratio three times: once on the answered subset, once with every unknown applicant counted into one group, and once with every unknown counted into the other. You already know what happened to those applicants. The only thing missing is which group they belong to, which is what makes the bound computable.
A gate takes in 1,000 candidates and passes 250. Demographics are known for 600 of them, and the answered subset looks spotless:
| Assignment of the 400 unknowns | Group A rate | Group B rate | Ratio |
|---|---|---|---|
| Answered subset only | 120/400 (30%) | 60/200 (30%) | 1.00 |
| All unknowns to group B | 120/400 (30%) | 130/600 (21.7%) | 0.72 |
| All unknowns to group A | 190/800 (23.8%) | 60/200 (30%) | 0.79 |
The answered subset reports perfect parity. Neither bound clears 0.8. Nothing about this gate is settled, and the finding to write down is the response rate, not the ratio. Note the direction of the second and third rows: they fail on opposite sides, because the unknowns passed at a lower rate than either known group, and whichever group absorbs them takes the hit.
Moving the block wholesale is the version you can compute in a minute and defend in a meeting. The range goes wider than that. Split the same 400 instead, the 70 who passed into group A and the 330 who did not into group B, and the ratio falls to 0.28. Treat a verdict that survives the two wholesale rows as a floor under the uncertainty, and say in the write-up which assignment produced the number you are reporting.
Nobody had to invent this. Researchers who read the bias audits published under New York City's Local Law 144 ran a cruder version of it, because a published audit does not say what became of the applicants whose demographics are missing: for each group they assumed the whole missing block belonged to it, then took the best and the worst selection outcome for that block. Of the 116 audits they collected between the law taking effect in July 2023 and November 2024, 53% report at least one impact ratio below 0.8. At the level of individual impact ratios, 70% have a possible lower bound below 0.8 under that extreme assignment, against 28% when the missing applicants are assumed to look like the known ones 1. The authors call the extreme version demonstrative and print the gentler one beside it. Read the pair as a range the published point estimates never showed, and the same range worth demanding when a vendor holds the data behind an audit you did not run.
Why aren't the missing answers random?
Because declining is a decision, and nothing guarantees that decision is independent of the demographics being asked about. A candidate weighs what they think happens to the answer, and that weighing has no reason to be uniform across the people asked or across the points at which you ask. Treating the unknown category as random is an assumption, and it is the assumption that hands the missing rows the average behavior of the people who did answer.
You can test the second half of that in your own data this afternoon. Pull the response rate at each stage separately: at application, at the assessment, at offer. If the rate climbs through the process, the same people are answering a question they refused earlier, which puts the refusal down to where it was asked. If it is flat, that is worth knowing too, and it makes one bound tighter.
What to refuse outright is filling the gap by inference. A survey experiment on the names used in correspondence audits found the congruent perception rate for Black first names was 75.0 percent with no last name, 82.5 percent with a Black last name and 66.5 percent with a white last name, and that names common among highly educated Black mothers were markedly less likely to be read as Black than names common among less educated Black mothers 3. That study tested how survey respondents perceive a name. It did not measure what employers did, and no callback gap follows from it. What does follow is that a name is a noisy indicator of the category you want, and the noise is patterned by class and education rather than scattered at random. Imputing from names or addresses replaces a known unknown with an unknown error that runs in a direction you cannot see, and it puts the employer on record assigning a protected characteristic to somebody who declined to state one. The cheaper route is to improve the ask, which is where the demographic question has to sit relative to the decision. What an analysis may touch, and what it may be used for once it exists, is a question for counsel on your own process.
Report the rate per stage, then work on raising it
Print the response rate next to every ratio, at every stage, or the number is unreadable by anyone who knows what to look for. One rate for the whole funnel is not enough: the ask sits at one point and the population changes at every gate after it, so the share of unknowns at the offer stage can look nothing like the share at application, and each stage ratio inherits its own.
The response rate cannot be broken out by group, because the group of the non-responders is the thing you do not know. What can be published is the overall rate per stage and per requisition, the composition of the answered subset, and the bounds. Anyone who hands you a response rate by race has computed something else and mislabelled it.
Four ways to raise the rate, in rough order of effect:
1. Move the ask off the application form to its own step after submission, so it is not competing with a candidate's incentive to look employable. 2. State the purpose in one sentence naming what the data is used for and who cannot see it. Longer is worse. 3. Make declining an explicit option rather than a blank, so a considered refusal is recorded as one. 4. Ask once, later, in a second place, such as the offer stage, and keep the two rates as separate measurements.
The thing never to do is make the question mandatory. It converts a voluntary instrument into a compelled one, it produces answers people did not want to give, and it buys a denominator by damaging the numerator. When the counts themselves are too small for any ratio to carry meaning, the response rate is not the first thing in your way: start with what to measure when there are too few hires.
Monday's first query is the response rate, before anything else gets pulled. It takes a minute, and it decides whether the rest of the afternoon is measurement or arithmetic performed on a guess.
Common questions
Is a 50% response rate usable?
Sometimes, and a response rate on its own cannot tell you which, at 50% or at any other level. Compute the ratio with all the unknowns counted into one group, then into the other, and see whether the verdict holds both times. How much missing data a verdict can absorb depends on how wide the gap is, which is why no single percentage answers the question. Publishing the rate lets a reader make that judgment themselves, which is the real reason it belongs on the page. A number that only holds under one assignment of the unknowns is a hypothesis, not a finding.
Should you drop candidates who did not answer?
Dropping the non-responders is a choice with consequences, so make it explicitly rather than by default. Excluding the unknowns assumes they behave like the people who answered, and that assumption is the one most likely to be wrong here. Keep them as their own category, report how many there are at each stage, and use them to build the bounds. If a conclusion only exists once the unknowns are removed, the honest description of that conclusion is that it depends on an assumption nobody tested.
Can you report the response rate by group?
No. The people whose group you would need for the denominator are the ones who did not tell you. You can report the composition of the answered subset, and you can compare that composition against a stated external reference such as an occupational or regional benchmark, with the mismatch between the two populations noted. What you cannot do is present either of those as a response rate per group without inventing the missing half.
Does a low response rate mean the tool being audited is fine?
No. A low rate widens the range of results the data is compatible with, which usually means it cannot rule out a problem either. Bounding often shows that a ratio reported as clean has a plausible lower bound well under 0.8. Read a low response rate as a loss of resolution rather than as good news, and be suspicious of any report that presents missing data as a reason for confidence in either direction.
Where should the demographic question sit to get more answers?
On its own step, after the application is submitted, with a one-sentence purpose and a visible decline option. The worst placement is inside the application form beside the work history, where it competes with a candidate's reasonable instinct to present themselves as employable and where it risks landing on the screen a recruiter reads. A second ask later in the process, kept as its own measurement rather than merged into the first, usually collects more without pressuring anyone.
References
- 1. Auditing the Audits: Lessons for Algorithmic Accountability from Local Law 144's Bias Audits facctconference.org Supports the bounding precedent and keeps its two figures apart, since they are not two points on one series: 53% of the published audits include at least one impact ratio below 0.8, while 70% of individual impact ratios have a possible lower bound below 0.8 under the extreme assignment of the missing data, against 28% under the distribution-matching assumption.
- 2. 29 CFR 1607.4 - Information on impact (Uniform Guidelines on Employee Selection Procedures, 1978) law.cornell.edu Supports the claim in 4(A) that impact records may be kept on a sample basis only where the sample is appropriate in terms of the applicant population and adequate in size.
- 3. How Black Are Lakisha and Jamal? Racial Perceptions from Names Used in Correspondence Audit Studies sociologicalscience.com Supports the claim that a name is a noisy and patterned indicator of race: congruent perception rates of 75.0 percent with no surname, 82.5 percent with a Black surname and 66.5 percent with a white surname, with names common among highly educated Black mothers less likely to be read as Black than names common among less educated Black mothers.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.