Screening
Write the Screening Criteria Before the First Application Opens
A screening rubric needs three or four criteria, written before the posting goes live, and each one has to name the evidence that satisfies it and where that evidence comes from. Anything you cannot state that way, including impact, polish and communication read off a page, comes out of the screen and moves to a stage where the candidate produces something. Date the file and keep it, because a bar set before anyone saw an applicant is the only version that can be argued with later.
The takeMost screening rubrics fail on their timing. The content advice inside them is mostly fine. Written after the first fifty applications, they encode the pool instead of the job: a criterion appears because three strong-looking people happened to have it, and nobody records that the bar moved. A dated file written before the pool arrived is the only artifact that lets a colleague, an auditor or a rejected candidate ask a question you can answer without reconstructing your own memory of a week of reading.
Where Olive fits
Open a role and see what the work shows
Olive publishes its six dimensions and the evidence standard for each one before a candidate begins, which is the same discipline as writing criteria before the pool arrives. A human reviewer writes every finding, and the candidate is given the identical report free on every tier.
Rank your shortlistWhat belongs in a screening rubric, and what does not
A criterion belongs if an application can settle it and two readers would settle it the same way. That admits far less than the standard templates do. Employment dates and continuity, a licence or registration issued by a body you could call, a stated and job-related minimum, a named system or client attached to a named place: those are checkable. Impact, achievement depth, communication and polish are not, because they are scored from prose.
The prose criteria were reasonable proxies when the applicant wrote the document. They measured effort and care, indirectly, and that was better than nothing. Scored today they partly measure which drafting tool the applicant opened, and the ranking templates in circulation still weight impact and communication as though nothing had changed. Call that an argument, because nobody has measured it. The fix it points at is a test you apply to every criterion before it goes in the file.
Write each one in three parts and the test applies itself:
- The claim. "Has closed books for a multi-entity company."
- The evidence that satisfies it. "An employer and a date range where that was the role, or a named system used to do it."
- Where that evidence comes from. "The employment history section, or a link the candidate provided."
If the third part is empty, the criterion cannot be screened and has to move downstream. That is not a loss. A criterion nobody can settle from an application was never being applied consistently; it was being felt, differently, by whoever was reading that afternoon. The situation where every document already looks strong is worked through in what to screen on when every resume looks perfect.
When should the rubric be written?
Before the posting goes live, and certainly before anybody opens an application. The order matters more than most teams expect. A rubric written after the first fifty applications is tuned by those fifty: a criterion gets added because several impressive-looking people had it, a threshold gets softened because the pile was thin, and none of that is recorded as a change to the bar.
The cost of getting the timing wrong lands on the people cut earliest. Applications one through fifty were screened against a different standard from applications fifty-one through four hundred, and nobody can say which. That is invisible in an outcome report, unfixable afterwards, and it is the honest reason to date the file. Nobody has to trust your memory of what the bar was in week one if the file says.
There is a second timing rule that gets skipped. Two people screen the first twenty applications independently, against the written criteria, before the rest of the pile is touched. Then compare. Where you disagree, the criterion is ambiguous and gets rewritten before it has been applied four hundred times. This costs an hour and it is the only cheap opportunity you will get, because after the pile is read the disagreement is buried in decisions nobody will revisit.
Expect the first comparison to be worse than you think. Even at interview stage, where scoring is deliberate and the scale is written down, around 38% of scorecard pairs in one large benchmark set differ by at least a point, and nearly half of those one-point gaps cross the yes-or-no threshold on a four-point scale 1. That is disagreement rather than error, measured with no outcome variable attached and no demographic variable either, so it supports nothing about accuracy or bias. What it supports is the modest claim that two careful readers routinely diverge at the boundary, which is precisely where a screen operates.
Write anchors, and use the rule you wrote
An anchor is a written description of what each rating means, and it is what stops a scale drifting between readers and between Tuesdays. "3 = has held the responsibility at a named employer for at least a year" is an anchor. "3 = good" is a number wearing a label. Anchoring costs an afternoon and it is the part of a rubric that most teams skip, because it is the boring part.
Anchors buy something real, and the size is worth stating precisely. Levashina and colleagues report Taylor and Small's meta-analysis of 19 past-behaviour interview studies, in which interviews using anchored rating scales showed higher criterion-related validity, .35 against .26, and higher interrater reliability, .77 against .73 2. The reliability gain is small; the validity gain is the larger of the two. The studies cover past-behaviour interview questions, and none of them is a screen. The comparison also runs across studies, so nothing here tests the same interview with and without anchors, and the review itself is blunt that anchored scales became popular for their logical appeal ahead of the evidence. Take it as a reason to write anchors. It does not forecast what yours will buy. The same review reports a separate finding that matters more in practice: anchoring every point of a five-point scale produced ratings more resilient to disability bias than anchoring only the endpoints 2, so a half-anchored scorecard is not what was studied.
Then use the rule you wrote instead of forming an impression of the whole person. In a meta-analysis of selection and admissions studies, combining applicant data mechanically by a formula correlated .44 with job performance against .28 when the same kinds of data were combined holistically by expert judgment 3. Mechanical here can be as plain as adding up unit-weighted scores on a written scorecard, and the paper says so. The job-performance comparison rests on nine studies and other criteria show much smaller gaps. The finding is about how evidence gets combined at the end, and it licenses no number standing in for a person.
What that means at a screen is narrow and practical. Decide in advance which criteria are disqualifying and which are additive, write it down, and apply it. The impression is still yours to have. It just stops deciding, and the argument for rating consistency between two people is carried further in writing a rubric two reviewers score the same way.
How do you know the criteria are wrong?
Read what they rejected. Criteria that are too narrow look identical to criteria that are exactly right when you only inspect the survivors, and the difference only shows up in the pile you cut. Pull twenty rejects at random from a closed requisition, hand them to whoever knows the work, and count how many they would have wanted to meet.
Employers have been agreeing about this outcome for years without acting on it. In the Harvard Business School and Accenture survey of 2,275 executives, 88% agreed that qualified high-skills candidates are vetted out because they do not match the exact criteria in the job description, rising to 94% for middle-skills roles 4. That is agreement with a statement about their own process, not a measured exclusion rate, and it was fielded in early 2020, so it describes criterion-based filtering that employers configured themselves. It predates any argument about AI, which is why the reject sample is standing practice in a normal quarter.
Four signals that your criteria need rewriting:
1. The reject sample keeps surfacing people the manager wanted. More than a couple out of twenty and something in the file is selecting for the wrong thing. 2. Two readers disagree on the same application. The criterion is ambiguous, and the repair belongs in its wording. 3. The pass rate moves without the pool changing. Usually a criterion being reinterpreted over time. 4. Nobody can say why a specific person was cut. The decision was made on something that is not in the file.
The qualification worth stating out loud, because it is the real trade: fixed criteria will let through people who look ordinary on paper and cut people who would have interviewed well. You are buying a screen that means the same thing on application four hundred as on application four, and you are paying for it in a certain kind of upside. Whether that trade is worth making is the question behind knowing whether your screen is throwing away the wrong people, and the reject sample answers it from your own pile.
Common questions
How many criteria should a screening rubric have?
Three or four, and rarely more than five. Every additional criterion multiplies the ways two readers can diverge, and past a handful nobody applies them all in the time a screen actually takes, so the extras become decorative. There is a second reason to keep the list short: a long list of requirements is read by candidates as a set of literal eligibility rules, and it suppresses applications from people who could do the job. If your list runs to nine, most of those items belong in the interview plan rather than in the screen.
What is the difference between a screening rubric and an interview scorecard?
The evidence available to each. A screening rubric can only use what an application can settle, which is checkable facts and specific claims attached to named places. An interview scorecard can use behaviour observed in a conversation, so it can carry things a document never could. Teams get into trouble by copying scorecard criteria upward into the screen, because impact and communication read fine as interview dimensions and become unreadable as screening criteria. Write them separately and let the screen be the narrower document.
Should the criteria be shared with candidates?
Sharing the checkable ones is usually safe and often useful. If a criterion is a fact about employment, a licence or a stated minimum, publishing it in the posting reduces out-of-scope applications and shortens the pile you have to read. The criteria worth keeping internal are thresholds that would invite gaming, and there is a diagnostic in that: if publishing a criterion would let people fake it, the criterion is probably not checkable enough to be in a screen at all. That is a signal about the criterion rather than about candidates.
Can we change the criteria mid-requisition?
Yes, and the one condition is that you record it. Write the new version, date it, note what changed and why, and keep the old one. What causes trouble is not the change; it is a change nobody wrote down, because the pool then gets screened against two different bars with no way to tell which applied to whom. If the change is substantial and the pool is large, consider rescreening the applications that were cut under the old criteria, which is cheap when the criteria are checkable facts.
Do we still need a rubric if one person does all the screening?
Yes, because the second reader you are protecting against is yourself next week. One person screening four hundred applications over ten days is not one consistent reader, and nothing about a single screener makes the standard stable across sessions, moods or the composition of the pile that day. A written file also means somebody can answer a question about a decision months later without reconstructing it. Single-screener processes are the ones where criteria drift hardest, precisely because nobody ever has to state the rule out loud.
References
- 1. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that two careful readers routinely diverge at the decision boundary: around 38% of scorecard pairs differ by at least a point, nearly half of those crossing the yes-or-no threshold on a 1 to 4 scale.
- 2. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the payoff from written anchors (.35 against .26 validity, .77 against .73 interrater reliability across 19 past-behaviour interview studies) and the separate finding that anchoring every point of a five-point scale produced ratings more resilient to disability bias than anchoring only the endpoints.
- 3. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports applying a written combination rule rather than forming a holistic impression: .44 for mechanical combination against .28 for clinical, on nine job-performance studies.
- 4. Hidden Workers: Untapped Talent hbs.edu Supports the reject-sample practice by showing employers already agreed their own criteria vet out qualified candidates: 88% high-skills, 94% middle-skills, 2,275 executives surveyed in early 2020.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.