Screening
Proxies Survive Deletion Because the Record Rebuilds Them
A proxy is any field your hiring screen's decision rides on that carries information about a protected characteristic and has no job-related reason to be in the rule. The working test is load rather than correlation: hold the job-relevant evidence constant, change the field, and see whether the decision moves. That test also explains why deletion under-delivers. The rest of the record rebuilds what you removed, and the fields carrying real weight in a screen today include process artifacts that appear on no standard list.
The takeThe circulating proxy lists are five demographic fields and a deletion instruction, which answers for a screen most teams stopped running two years ago. Nobody lists response latency, a paid model subscription, bandwidth, or a quiet room for a timed exercise, and those are fields a modern screen genuinely consumes. Run the load test over what your stage reads today rather than over somebody's list. If a field survives with a written reason attached, keep it and say so out loud.
Where Olive fits
Open a role and see what the work shows
No screen can establish who wrote a document, so Olive assesses the person rather than the artifact: a 40-to-60-minute occupational assignment done with an AI assistant, returned as six findings with the timestamp behind each one. The candidate is granted the identical report on every tier.
Rank your shortlistWhat counts as a proxy?
Any field the decision rides on that carries information about a protected characteristic without a job-related reason for being in the rule. Correlation alone is not the test: plenty of genuinely job-related evidence correlates with something protected, and where it does the field stays and the reason gets written down.
The load test is three lines long. Take a candidate the stage rejected. Change the suspect field and nothing else, holding the job-relevant evidence exactly as it was. If the decision flips, the decision was riding on that field, and the only question left is whether there is a reason for it. Doing this on ten real records produces a shorter, truer list than any published inventory.
One proxy is now named in statute. Illinois Public Act 103-0804 amended the Illinois Human Rights Act, effective January 2026, to make it a civil rights violation for an employer to use artificial intelligence that "has the effect of" discriminating on the basis of a protected class, or to use zip codes as a proxy for a protected class 5. Two things about that are worth carrying. The Illinois test turns on effect, while Texas wrote an intent standard into its own AI act effective the same month and said expressly that a disparate impact alone does not establish intent 6, so being compliant with AI hiring law is not one coherent claim across jurisdictions. The second thing is that a statute can name one field out of a large family. A screen has to cover the family. Jurisdiction, timing and application are questions for counsel.
Why doesn't deleting the field remove the information?
Because the rest of the record rebuilds it. A resume that no longer names a college still carries the club somebody captained, the city they worked in, the years they were out of the workforce, and the way they write. Any rule with access to the remainder can reconstruct most of what the deleted field carried, and a statistical rule will do it without being asked to.
The best-known example is usually told wrong. Reuters reported that Amazon built an experimental recruiting engine from 2014, trained on patterns in a decade of resumes submitted to the company, and that by 2015 it was penalising resumes containing the word "women's" and downgrading graduates of two all-women's colleges; the team was disbanded by the start of 2017 1. Reuters reports no candidate outcome at all, records that recruiters looked at the tool's output but never relied solely on it, and rests the whole account on five people speaking anonymously to one reporter. The durable sentence is the one usually cut from the retelling: Amazon edited the programs to be neutral to the flagged terms, and that was no guarantee the machines would not devise other ways of sorting candidates that could prove discriminatory.
Deletion as a remedy has been tested under randomisation, and it went the other way. When the French public employment service randomly assigned about 600 participating firms to receive anonymised or name-bearing resumes, anonymisation widened the interview gap: the interview rate for minority candidates fell and the rate for majority candidates rose 2. The authors' reading is not that anonymisation is bad in principle. Firms that volunteered for the program were already interviewing relatively more minority candidates, and stripping the name also strips the context that let a recruiter discount a negative signal such as an employment gap for someone they wanted to place. Removing an identifier is an intervention with its own effects rather than a neutral safety measure, which is the whole of why blind screening under-delivers against its reputation.
Thinner records are not safer records either. In a study of embedding-based resume retrieval, cutting the document down to a name and a job title moved the measured differences the wrong way: significant race differences turned up in more of the bias tests than full-length resumes produced, and significant gender differences in a good deal more 3. The paper's own denominators do not reconcile cleanly, so the finding here is a direction. The gender increase also came entirely from more preferences for resumes carrying female names. The mechanism is mundane. Strip the content and the name becomes a larger share of what the system has to work with.
Which fields are actually doing the work in your screen?
The ones a person configured, mostly. The most consequential filters in a modern screen are not learned by a model; they are rules a recruiter or an administrator typed into a system, and they name no protected class, which is precisely why they survive review year after year. That puts the fix in the criteria list, where a person can read it.
One of them has been measured. In a survey of 2,275 executives across the US, UK and Germany, 48% of employers said they filtered middle-skills candidates out on an employment gap of more than six months, and more than 90% of the employers who run a recruitment management system used it to filter or rank candidates at initial screening 4. Both figures count employers. The 48% says nothing about how many resumes the rule threw out, the 90% is a share of the roughly two-thirds of respondents who had such a system, and the survey was fielded in early 2020, so it describes how screens get configured and no current rate follows from it. An employment-gap rule mentions nobody's protected characteristic. Its weight falls on caregivers, veterans, people who were ill and people leaving incarceration.
Then there are the fields nobody lists, the ones that arrived with the tooling:
- Response latency. A screen that advantages whoever replied within the hour is measuring availability during the working day.
- Access to a paid frontier model. An exercise that assumes a subscription measures who holds one before it measures anything about the work.
- Bandwidth and a quiet room. A timed, live or video exercise reads household circumstances as composure.
- Employment continuity. Priced at the screen whether or not anyone configured a filter for it.
- Writing register. Sentence rhythm that reads as unusual to a reviewer often marks a second-language writer or somebody using assistive tooling.
Each is a field the decision rides on, and each fails the load test unless somebody can write the job-related reason beside it. The last two are also how a screen that lost its old filter quietly acquires a new one: when an ATS keyword filter stops working because every resume matches, the reviewer's own sense of what reads well takes over, and it is unwritten. The same question, whether a familiar field still predicts anything, is what GPA for entry-level roles turns on.
Write the reason next to every field, then cut the blanks
The deliverable is one page and it takes an afternoon. List every field your screen actually consumes, including what a human reviewer reads and no form records. Write one sentence beside each naming the job-related reason it is in the rule. Cut the ones where the sentence will not come, and keep the ones where it will, with the sentence attached to the field where a successor can find it.
Three refinements stop the exercise turning into theatre.
Write it for the stage you run, not the stage you designed. The list has to come from watching a screen happen, because the gap between the documented criteria and what a reviewer's eye actually stops on is where most proxies live.
Do not correct for a field you have decided to keep. Reweighting a field to cancel its group effect leaves the field in the decision and adds a second rule nobody can explain a year later. If it has a reason, keep it plainly. If it does not, remove it.
Where the reason is real and the load is uneven, go looking for the alternative that does the same job. A requirement that predicts what you need while carrying less unrelated information is better on both counts, and that search is what a less discriminatory alternative looks like in practice.
The last one bites hardest on the fields that arrived with generative tools, because a screen reading a document whose author cannot be established from the document is reading tooling access. What is left to screen on when every new-grad resume is AI-built is that case at its sharpest. The answer there and here is the same: move the decision onto evidence you asked for and can trace, and stop asking a document to prove something it never carried.
Common questions
Is a field a proxy just because it correlates with a protected characteristic?
No, and treating correlation as the test produces a list nobody can act on. Plenty of genuinely job-related evidence correlates with something protected. The question is whether the decision moves with that field once the job-relevant evidence is held constant, and whether anyone can write down a job-related reason for it being in the rule. A field with a written reason stays. A field with a correlation and no reason is the one to cut.
If we delete names and schools, is the screen fixed?
No. The rest of the record rebuilds most of what those fields carried, and when anonymised resumes were put through a randomised trial the interview gap widened. Deletion is an intervention with its own effects, including removing context a reviewer used to discount a negative signal. Removing a field is worth doing when nobody can justify it, but understand it as cutting an unjustified input rather than as neutralising the information.
What are the proxies people miss in an AI-era screen?
The ones that arrived with the tooling: response latency, access to a paid model subscription, bandwidth and a quiet room for a timed or live exercise, employment continuity, and writing register. None of them appears on the standard proxy lists, all of them can move a decision, and each fails the load test unless somebody writes a job-related reason for it. A screen that reads sentence rhythm as a quality signal is reading language background and assistive tooling along with it.
Is using a ZIP code illegal?
In Illinois, using zip codes as a proxy for a protected class became a civil rights violation in January 2026, alongside a broader prohibition on AI that has the effect of discriminating. That is one state with an effects standard, and Texas wrote the opposite into its own AI act effective the same month, requiring intent and saying a disparate impact alone does not establish it. The design answer holds everywhere: a postal code in a screen almost never has a job-related reason attached, which is enough to cut it without waiting for a statute. Application is a question for counsel.
What do we do with a field that has a real reason and still hits one group harder?
Keep it, write the reason, and go looking for something that does the same job with less unrelated information attached. That search is the practical form of the less-discriminatory-alternative question, and it is more productive than reweighting, which leaves the field in the decision and adds a rule nobody can explain later. Record what you considered and why you kept what you kept, because that record is what makes the choice reviewable.
References
- 1. Amazon scraps secret AI recruiting tool that showed bias against women web.archive.org Supports the account of the abandoned Amazon experiment and the line that editing out the flagged terms was no guarantee against other proxies.
- 2. Unintended Effects of Anonymous Resumes docs.iza.org Supports the randomised finding that anonymising resumes widened the interview gap, and the two mechanisms the authors identify.
- 3. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval arxiv.org Supports the claim that a thinner record produced more measured group differences in an embedding-based screen, quoted as a direction rather than as counts.
- 4. Hidden Workers: Untapped Talent hbs.edu Supports the employment-gap filter figure and the point that the most consequential screening rules are configured by a person.
- 5. HB3773 Enrolled (Public Act 103-0804), amending the Illinois Human Rights Act ilga.gov Supports the statement that one state prohibits AI with a discriminatory effect and names zip codes as a proxy for a protected class.
- 6. Texas H.B. 149 (89R), Texas Responsible Artificial Intelligence Governance Act, enrolled text capitol.texas.gov Supports the contrast with Illinois: Texas requires intent and states that a disparate impact alone does not demonstrate it, which is why compliance is not one claim across jurisdictions.
6 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.