Assessment design
Replace the Five-Point Scale With a Call and Its Reason
When interviewers score a round, ask for a call, hire or no hire, plus the evidence that produced it, and treat a rating out of five as a column your applicant tracking system may require rather than the point of the exercise. A binary forces the interviewer to commit, which is the useful part, but a bare hire or no-hire is only a shorter number. Keep a scale if something downstream consumes it by a stated rule, and never average one across rounds that observed different things.
The takeNumbers are not the enemy here, and the research most often quoted against them says the opposite. A rule that combines the same fields the same way every time genuinely beats a room full of impressions, and a written checklist adding unit-weighted values is such a rule. What fails is the middle case nearly everyone ships: a value produced by judgment, combined by conversation, and then defended because it is a value. Commit to the rule or drop the column.
Where Olive fits
Open a role and see what the work shows
If you are building this in-house, the expensive parts are the answer key and the evidence trail behind each finding. Olive ships authored cases grounded in one occupation and returns six separately evidenced findings, each reported as demonstrated, partly demonstrated or not demonstrated rather than as a value.
Rank your shortlistWhich field should be required?
The evidence box and the call, in that order, with the rating left optional. Those two are the fields a debrief can question. A number is not, which is why nobody ever argues with a 4 and can only answer it with a 3. An interviewer who writes a decision plus the two specific things that produced it has done the whole job.
A usable output block fits in four lines and takes about ninety seconds to fill in:
- Call. Hire or no hire, and the round it rests on.
- Because. Two claims, each with the thing that was said or done sitting next to it.
- Against. The strongest reason not to, written by the same person.
- Unseen. What this round could not reach.
The Against line does more work than the rest of the form. It costs thirty seconds, and it turns a one-sided card into something the room can test, because the person who most wants to hire has already written down the best argument on the other side. It also makes a specific and common failure visible: an interviewer whose Against line is empty on every candidate has stopped weighing and started agreeing.
The fields that go above this block, and the order they belong in, are the subject of what actually goes in a hiring scorecard. The output field is the last decision you make about that form.
Does a number help or hurt?
It helps when a rule consumes it and hurts when a conversation does. The evidence on combining candidate data is unusually clear, and people reach for it to defend their scores when what it tested was the combination rule. What beats an impression is applying the same procedure to the same fields for every candidate, and a written checklist does that as well as an algorithm does.
The meta-analysis behind that claim compared studies in which applicant data were combined mechanically by a formula against studies in which the same kinds of data were combined holistically by expert judgment. Average correlations with job performance were .44 for the mechanical route and .28 for the holistic one, which the authors describe as a population-level improvement in prediction of more than 50% 1. Read the limits before spending it: the job-performance comparison rests on 9 studies, other criteria show much smaller gaps, and the improvement is in a correlation rather than in hires made. The mechanism is what travels, and the authors are explicit that a mechanical method can be as plain as adding up unit-weighted values on a written scorecard.
So nothing here indicts numbers. A number with no rule waiting for it is decoration, and decoration on a hiring form gets read as measurement. If nothing downstream adds that value up, compares it to a threshold, or weights it against another one, delete the column and keep the call. If something does, publish the rule to everyone who fills in the form, before the loop opens.
Keep the scale narrow and anchor every point
If you keep a scale, keep it short and write an anchor for every point on it, including the middle. The payoff from anchors is real and modest, and it is larger for validity than for agreement between raters. Their main work is forcing the team to say in advance what a middle rating would have to look like in this role.
The standard review of structured interviewing reports a meta-analysis of 19 past-behaviour interview studies in which interviews using anchored rating scales showed higher criterion-related validity (.35 against .26) and slightly higher interrater reliability (.77 against .73) than interviews without them 2. Past-behaviour questions only, and a comparison across studies rather than the same interview run both ways, so anchored scales travel with whatever else careful teams already do. The review is blunt that anchored scales became popular for their logical appeal ahead of the evidence, which is a fair description of most scorecard design.
The count of points matters less than the anchors, which is why four-against-five is the least productive part of this argument. Removing the middle option forces a lean without making it informative. If nobody on the team can write a sentence describing what a middle rating looks like for this role, that inability is the finding, and no amount of scale design will repair it.
Never average a tool's value with a human rating
A value produced by a tool goes in its own row, labelled with what produced it, and it never enters the same column as a rating a person wrote after watching somebody work. Averaging the two yields a number whose meaning nobody in the room can state. The two are not measuring comparable things, so anyone who wants them in one column should have to state the rule that combines them first.
The evidence on machine-assessed interviews is thin enough to make that the only supportable position. A peer-reviewed study of automated video interview assessments, whose authors say they are unaware of any previous literature examining such scores against organisational criteria, reports an uncorrected correlation of .24 with job performance across five organisational samples totalling 1,124 people, against .32 for human-rated structured interviews stated in the same uncorrected terms 3. Four of the five authors worked for the vendor whose algorithms were evaluated, five samples stand against more than a hundred effect sizes on the human side, and the authors themselves call the research area in its infancy. Nothing in it licenses treating a machine value and a human rating as the same kind of quantity.
Whether any assessment value still tracks performance once candidates work with an assistant is a live question in its own right, taken up in whether assessment scores still predict performance. The same separation applies to submitted work: a polished deliverable and an observed process belong in different rows, which is the practical problem in grading take-homes that all come back polished.
On Monday, change the required field. Make the form ask for a call and the evidence that produced it, demote the rating to an optional column, and give any tool-produced value its own labelled row. Then sit through the next two debriefs and count how often somebody quotes a number instead of quoting the candidate. That count is the measurement, and it moves before the meeting gets any shorter.
Common questions
Does a hire or no-hire call make interviewers harsher?
It makes them commit, and a commitment reads harsher than a 3 does. The more common effect runs the other way: a binary makes a soft no visible, where a 3 lets an interviewer avoid deciding and hand the problem to the debrief. If the team drifts hard in either direction after the change, that is a calibration issue rather than a field-type issue, and it shows up as one person's calls diverging from everyone else's across a run of candidates.
What about a four-point strong hire to strong no-hire scale?
It is a scale with four points and a friendlier vocabulary, and it inherits everything a scale does. It can be averaged, quoted without evidence, and argued about by adjective. Nothing is wrong with using it, provided the required fields underneath are the evidence and the reason. If those are present, the wording of the top option is a matter of taste. If they are absent, four labels do no better than five numbers.
Can interviewer ratings be averaged across a loop?
Only across rounds that observed the same claim, and almost no loop is built that way. Averaging a technical round against a values round produces a number about nothing in particular, and it buries the case that matters most, which is one strong signal sitting beside one clear concern. If a single value has to exist for a system, derive it from a stated rule everyone has seen, and keep the underlying calls visible next to it rather than behind it.
Should candidates see the rating scale?
The anchors can be shared and often should be. Telling a candidate what a round assesses and what a strong answer contains changes how they prepare, and preparation is not the thing being measured. The numbers themselves travel badly, because a value lifted out of the form reads as a verdict on a person rather than on one round. If you share anything, share what the round is looking for.
What if the applicant tracking system requires a numeric field?
Fill it in and stop treating it as the output. Most systems accept a value derived from the call rather than the reverse: no hire maps to the bottom of the range, hire to the top, and the content lives in the evidence field above it. What matters is which field the team argues about in the debrief. If that field is the number, the tool has quietly redesigned the process and nobody has voted on it.
References
- 1. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the claim that combining candidate data by a stated rule outperformed combining it by expert judgment (.44 against .28), and the authors' point that a written unit-weighted scorecard counts as such a rule.
- 2. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the size of what anchored rating scales buy: validity .35 against .26 and interrater reliability .77 against .73 across 19 past-behaviour interview studies, plus the review's own caution that anchors spread ahead of the evidence.
- 3. Psychometric Properties of Automated Video Interview Competency Assessments hirevue.com Supports the .24 uncorrected job-performance correlation for machine-assessed video interviews across five samples of 1,124 people, against .32 for human-rated structured interviews, with the vendor authorship stated.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.