Assessment design
Every Number on the Scale Needs an Observable Behavior
Each point on an interview rating scale should name a behavior an interviewer could have watched, in the words of work this role actually contains, with one real example beside it. Adverb labels like meets or exceeds describe the rater's feeling, so identical performance draws different numbers. Three anchored points beat five vague ones, every point gets an anchor including the middle, and a separate no-evidence box keeps the scale from absorbing rounds where the dimension never came up.
The takeThe adverb scale is not a neutral default; it is the reason two raters can both be honest and disagree. A template that ships poor-to-outstanding has quietly decided the panel will rate impressions, and it costs nothing to print, which is why it keeps arriving. If a scale point cannot be written as something a person did, that point is not measuring anything. Delete it before the panel spends an hour arguing about what it means.
Where Olive fits
Open a role and see what the work shows
Olive reports each of six dimensions as demonstrated, partly demonstrated or not demonstrated, with the timestamped excerpt that produced the finding printed beside it, so the judgment and the evidence for it arrive together. A human reviewer writes every word, and the candidate is given the same report.
Rank your shortlistWhat should each number actually mean?
A point earns its place when it names something a candidate did or said that a second interviewer could have watched and recognized. Not a level of quality, not an adverb: a behavior, drawn from work this role actually contains. That is the difference between a scale two raters apply the same way and a scale that records how each of them felt about the person in front of them.
Here is one dimension, checks a claim before relying on it, written as three points:
- 1. Carried a figure from an assistant's answer straight into the recommendation, and when asked where it came from, restated the answer.
- 2. Named which figures would need checking and why, but checked none of them during the session.
- 3. Opened the source, found the figure covered a different period, and changed the recommendation.
Beside each point, paste one real sentence from a past submission that landed there, stripped of anything that identifies who wrote it. The quote is what stops a panel arguing about the wording of the anchor. Two raters can spend an hour disagreeing about what checked means and agree in a minute once they are looking at the same sentence.
Narrowing discretion is the second thing anchors do, and it is where the fairness argument sits. A 2026 preprint auditing 36,880 applications to 9,220 postings for new US college graduates found callback gaps widest in roles combining high analytical and interpersonal demands with low routine content, with callbacks 28 to 43 percent lower for Black men, Black women, White women and Hispanic men than for otherwise identical White men in management occupations 3. Those are callbacks at the screening stage, the paper is a preprint under review, and the range covers four different groups. The authors propose discretion as the mechanism; the audit did not manipulate it. It still points at the same lever: the less a judgment rests on written criteria, the more room there is for something else to move it.
Write the anchors before you see a candidate
Draw the anchors from the job rather than from the pool. Anchors have to exist before the first candidate is seen, or they turn into a description of the first two people you met, and every candidate after that gets measured against those two. The sources are the work itself, the last few submissions you kept, and twenty minutes with whoever currently does the job.
That person is the one who can tell you what a wrong answer looks like in this material, which is most of what the top anchor has to say. A recruiter working from a template cannot supply it, and a template is where an adverb scale comes from in the first place.
Three fully anchored points beat five where only the ends are written. A five-point scale with anchors on 1 and 5 leaves the panel to invent 2, 3 and 4 privately, which is the condition the exercise was meant to remove. If the applicant tracking system forces five, write all five.
The payoff is real and worth sizing honestly. Levashina and colleagues' review reports a meta-analysis of 19 past-behavior interview studies in which interviews using anchored rating scales showed higher criterion-related validity (.35 against .26) and higher agreement between raters (.77 against .73) than interviews without them 1. The agreement gain is small; the validity gain is the larger of the two. The studies were sorted by whether they used anchors, never run twice with and without them, so part of the gain belongs to whatever else careful teams do.
Anchors also do not transfer between roles, which is the part teams resist because it makes the work recur. A scale written for a support role and reused for an analyst role keeps the numbers and loses the behaviors. The same constraint runs through a rubric two reviewers score the same way: agreement comes from the examples, and examples are role-specific.
Why does the middle of the scale need its own anchor?
Because a rater with nothing to go on picks the middle, and an unanchored middle turns that into a mediocre rating. The round that never reached the dimension and the candidate who was genuinely middling produce the same 3, and nobody downstream can tell them apart. So write the middle anchor as carefully as the top one, and put a separate box beside the scale reading no evidence gathered.
That box does two jobs at once. It keeps absences out of the numbers, and it tells you which dimension the loop failed to cover, which is a finding about your process rather than about the candidate. When the same dimension comes back empty three candidates running, a round is not doing the job it was added for.
Anchoring every point appears to matter for who gets rated fairly, too. The same review reports a study in which anchoring all five points of a five-point scale produced ratings more resilient to disability bias than anchoring only the top and bottom 1. One study, and it looked at ratings on their own. Treat it as a reason to write the middle anchor. It proves nothing about who your funnel hires, and it costs nothing to act on.
When the argument is about where the bar sits, that is separate work and it comes first. Setting a defensible bar for good enough in a specific role has to be settled before the anchors, because a scale can only be as clear as the standard underneath it.
Test the scale on two raters and one old submission
Take the dimension your panel argues about most, rewrite its points as three observable behaviors with one real quote each, and hand it plus one past submission to two interviewers who rate independently. If they land on the same point for the same reason, the anchor holds. If they land on the same point for different reasons, it is loose, and the quotes beside it need replacing before the next candidate sees it.
Ask each rater to write the sentence that produced the rating next to the number, naming the capability it demonstrates. There is some evidence that naming the capability carries signal: text-mining post-interview notes on 7,650 people hired at one large Chinese technology company found that the number of job-relevant capabilities an interviewer named in the notes tracked later performance and promotions and ran the other way against turnover, at roughly 2 percent of a performance measure per standard deviation 2. Small, correlational, one firm, and only observable for people who were hired, since nobody sees how the notes would have predicted for a rejected candidate. Still, the notes carried something the rating on its own did not.
Then leave the rest of the scorecard alone until the next requisition opens. A panel can absorb one rewritten dimension mid-loop. A wholesale rewrite between candidates means the first three were rated on a different instrument, and the comparison you were running has quietly ended.
If the numbers still refuse to mean anything after all of this, the honest next move is asking for a call and its reason instead of a rating. A scale exists to make ratings comparable, and a panel that cannot make its ratings comparable is better served by a decision with the evidence attached.
Common questions
How many points should the scale have?
Three, unless something forces more. Three points can each carry a written behavior and a real example, and most panels can hold three distinctions in their heads while listening. Five points are usually three points plus two the team never defined, and the two undefined ones absorb the cases nobody wants to argue about. If the applicant tracking system requires a five-point field, keep five and write all five anchors, including the middle. What matters is not the count but whether every point on it names something observable.
Can we reuse anchors across roles?
The dimension can travel; the anchors cannot. Checks a claim before relying on it is a sensible dimension for an analyst, a journalist and a nurse manager, and the behavior that counts as a 3 is different in all three. Reusing the text keeps the appearance of a rubric and drops the part that made two raters agree. Copy the dimension names, then spend twenty minutes with someone doing the new job to rewrite the examples underneath them.
What if the panel cannot agree on the top anchor?
That disagreement is worth more than the scale is, and it should be settled before a candidate is booked. A panel that cannot describe what excellent looks like in this role is not disagreeing about wording, it is disagreeing about the job. Ask each person to bring one real example of work at that level, then write the anchor from what the examples have in common. If nobody can produce an example, the dimension is aspirational and does not belong on the scorecard yet.
Should candidates see the anchors?
Send them. It costs little and removes a class of unfairness that has nothing to do with ability. A candidate who has been told what the top anchor requires can aim at it; one who has not is guessing at a standard the panel wrote down and kept. The usual objection is that publishing the anchors invites rehearsed answers, which is a real risk for questions with one correct shape and a small one for anchors describing what a person actually did.
What do we do with a dimension nobody could observe?
Mark the no-evidence box and leave the rating empty, then decide whether to run another round or to proceed without it. Never average an empty cell into a total, and never let a middle rating stand in for one. The dimension that keeps coming back empty is telling you something about the loop rather than about the people in it: either no round is designed to surface it, or the question meant to surface it is not working.
References
- 1. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the anchored-rating-scale figures (validity .35 against .26, interrater reliability .77 against .73, across 19 past-behavior interview studies) and the review's further report that anchoring every point of a five-point scale produced ratings more resilient to disability bias than anchoring only the endpoints.
- 2. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company frontiersin.org Supports the claim that naming job-relevant capabilities in interview notes carries predictive signal a rating alone does not, at roughly 2 percent of a performance measure per standard deviation across 7,650 hires.
- 3. Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit arxiv.org Supports the claim that measured callback gaps are widest where evaluation is least routine, cited as an unreviewed preprint about the screening stage rather than about interview ratings.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.