Assessment design
A Rank Order Is a Weaker Claim Than a Bar
Set a bar. Write the criterion from tasks the job actually contains, judge every candidate against it in writing with the evidence attached, and order only the people who already cleared it. A rank is a claim about who else applied this week, which changes weekly and cannot be defended twice. A bar is a claim about the role, and a rejected candidate can be shown what it was.
The takeRanked output survives because it sorts cleanly and looks like more information than it is. The distance between the fourth and fifth candidate on one afternoon of work usually sits inside the measurement error of whatever produced it, and nothing about a ranked list tells you where that error starts. A bar is the weaker-looking claim and the stronger one: it names what the job needs, it holds still while the pool changes, and it is the only one of the two you can say out loud to the person you turned down.
Where Olive fits
Open a role and see what the work shows
If you are building the bar in-house, the two costly parts are the answer key and the evidence behind each judgment. Olive returns six findings on one candidate, each carrying the timestamped excerpt it rests on, with outcomes written as demonstrated, partly demonstrated or not demonstrated rather than as a number.
Rank your shortlistWhat a rank claims, and what a bar claims
Two different claims. A rank says candidate four placed above candidate five among the people who applied this week. A bar says this candidate can do something the job requires, judged against a standard written before anyone applied. Only the second claim survives a changed applicant pool, a rejected candidate asking why, and the same question asked again next quarter.
Defending the decision is where it first shows up. Under Title VII, the federal statute amended in 1991, the test for a practice that produces a disparate impact is whether it is job related for the position in question and consistent with business necessity 1. That is a description of a bar. A percentile has no answer to it, because a percentile is a fact about the other applicants, and the other applicants are not the position in question.
It shows up again in what can be said to a person. A bar can be shown to a candidate: here is the task, here is what clearing it looked like, here is where the submission stopped. An ordering cannot be shown to anyone, because saying it out loud means naming the strangers a candidate was measured against, and the substitute most teams reach for, that the field was strong, explains nothing.
It shows up a third time next quarter. A criterion written from the work still means the same thing in March and can be applied to a candidate who arrives alone. An ordering has to be regenerated from whoever is in the pool, so two people of identical ability get different outcomes depending on the week they applied. If the standard is genuinely hard to write down, that is the real problem: set a defensible bar for the role.
Why the gap between fourth and fifth is usually noise
Because the instrument has a wide band around it and the ordering never shows the band. Structured interviews sit top of the corrected validity table at .42, with an 80% credibility interval running from .18 to .66 2. That is the spread across settings for the best evidenced single method there is, and a ranked list reports none of it.
Two consequences follow. The first is that a small difference between two candidates on one sample is not a finding, it is the instrument breathing. The second is that combining two or three different kinds of evidence does more for a decision than sharpening the ordering inside one of them: a mechanically combined composite reaches about .61, and dropping cognitive ability out of it costs .05 2.
The 2022 re-analysis underneath those numbers puts work sample tests at .33 and warns against reading that against the .42 for structured interviews as one method beating another, because the credibility intervals overlap heavily 3. If the authors decline to separate two methods on a difference of nine hundredths, a hiring team should be slower still about separating two people on an afternoon of work.
Nothing here says stop comparing candidates. The question is which comparison the decision is allowed to rest on. Two finalists whose submissions look equally good is a real situation with a real answer, and the answer is more evidence rather than a finer sort: work out who actually did the thinking.
Write the bar before the first application arrives
Write it from the work, in the week the requisition opens. Take three or four things the role does in its first quarter, state what an acceptable performance of each looks like in a sentence a stranger could apply, and decide in advance what evidence counts as clearing it. A bar written after the first strong candidate is a description of that candidate.
A usable criterion has four properties. It names a task from the job rather than a trait. It says what evidence settles it, so two people reading the same submission reach the same call. It exists in writing before anyone applies. And it sits at the level the job requires, not the level the best applicant happened to reach.
- Weak: strong analytical skills.
- Better: reconciles a month of transactions and explains the two largest variances.
- Weak: comfortable with AI tools.
- Better: takes a confident model output on a task with a checkable answer, finds the part that is wrong, and says how they knew.
Three or four criteria is the working number. Twelve produces a checklist nobody applies the same way twice, and it quietly reintroduces the ordering, because a reader facing twelve marks starts totalling them. Keep the set small enough that a rejection can be explained in one sentence naming which criterion and what was missing.
Decide two policy questions at the same time. What happens to a candidate who clears every criterion but one, which is a rule about the role rather than a judgment about a person. And who may move the bar mid-process, which should be nobody, because a bar that drops when the pool looks weak is an ordering wearing different clothes.
When is ranking the right tool?
When more than one person has cleared the bar and there is one seat. Ordering inside a qualified group is a resource decision rather than a claim about competence, and it is the only place a comparison carries its own weight. Make that explicit in the debrief: the ordering starts after the criterion, it applies to this seat, and it says nothing about whether the second person can do the job.
Keeping the two steps apart pays off later. Candidates who cleared the bar and lost the seat are a real shortlist for the next requisition, because the record says what they demonstrated rather than who else was in the room that month. Candidates who did not clear it are a different set, and the record says which criterion and why.
The debrief is where this usually goes wrong, since a room asked to compare people will compare people whether or not a criterion exists. Collect the judgments against the criterion first, in writing, and let the ordering be the last five minutes rather than the first thirty: collect the scores before the debrief.
Common questions
Is a pass bar the same as a cut score on a test?
Close, but a cut score is one number on one instrument and a bar is usually several criteria, each with its own evidence. The practical difference is what can be said afterwards. A cut score of 62 invites the question of why not 60, and the honest answer often does not exist. A criterion stated as a task and an acceptable performance of it can be defended by pointing at the job. If you do use a numeric cut, write down how the number was chosen before you use it, because that reasoning is the part anyone will ask for.
What if only one candidate clears the bar?
Then you have one candidate, which is information rather than a failure. Either the bar is right and the pool was thin, or the bar describes something the job does not actually require. Decide which before the next requisition, not during this one. Moving a criterion because a specific person fell short of it converts a claim about the job into a claim about that person, and it is the most common way a defensible process stops being one.
Does setting a bar make hiring slower?
Not usually, because the expensive part of hiring is the re-judging rather than the judging. A written criterion resolves most submissions in one pass and sends only the genuinely ambiguous ones to a second reader. What does take time is writing the criterion, once, before the role opens. Teams that skip it pay the hours back in debriefs where four people argue about what good looks like, having never written it down.
How does a bar interact with adverse impact analysis?
A bar gives you a selection rate, which is the quantity every impact analysis is built from: how many people from each group cleared the criterion, out of how many who tried. An ordering with no cut point gives you nothing to compute until someone draws a line, and the line tends to get drawn afterwards, wherever it makes the shortlist come out the size somebody wanted. The analysis itself is a conversation with counsel; the point here is only that a bar produces the number and an ordering does not.
Can candidates be told where the bar was?
Yes, and the process improves when they are. Publishing what the assessment asks for costs little, because the criterion is about the job, so a candidate preparing against it is preparing to do the work. An ordering cannot be published at all, which is why teams that sort candidates end up sending rejection reasons that mean nothing to the person receiving them. Put the criteria in the invitation and the rejection writes itself.
References
- 1. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases uscode.house.gov Supports the claim that the statutory test asks whether a practice is job related for the position in question and consistent with business necessity, which is the shape of a bar rather than an ordering.
- 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors cambridge.org Supports the .42 estimate for structured interviews with its 80% credibility interval of .18 to .66, and the composite figures behind the claim that different kinds of evidence beat a finer ordering.
- 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the .33 work sample estimate and the authors' own warning that it should not be read against .42 as one method beating another.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.