Assessment design
A Fixed Rule Beats the Same Manager's Holistic Read
A simple scoring rule does beat an experienced hiring manager's judgment on the direct test: hand an expert who knows the job and a fixed formula the same applicant evidence, and the formula predicts job performance better, .44 against .28 across the selection and admissions studies where both were run on the same predictors. In many of those studies the expert held more information than the formula did. The weights have to be set before anyone meets a candidate, and unlimited overrides put you back where you started.
The takeThe hedge that quietly kills structure is use the scorecard as an input, then apply judgment, because that is the holistic step with arithmetic in front of it. Judgment has real work here: deciding what to test, writing the questions, reading the answer, noticing what the rubric failed to ask. What it should not do is the addition at the end. Executives who accept that division usually find the argument was never about the quality of their judgment but about where in the process it gets applied.
Where Olive fits
Open a role and see what the work shows
Olive reports six dimensions as demonstrated, partly demonstrated or not demonstrated, each printed beside the timestamped excerpt a human reviewer wrote it from, and produces no total. A candidate who disagrees has a specific moment to point at, and the candidate is given the same report the employer reads.
Rank your shortlistDoes a scoring rule beat an experienced manager?
Where the comparison has actually been run, yes. In a meta-analysis of employee selection and academic admissions studies, the average correlation with job performance was .44 when applicant data were combined mechanically by a formula and .28 when the same kinds of data were combined holistically by expert judgment, which the authors describe as a population-level improvement in prediction of more than 50 percent 1.
The comparison is worth taking seriously precisely because it is narrow. A study qualified only if it ran both methods on the same data from the same predictors, and the authors add that in many cases the expert had access to more information about the applicant than the formula did, so this is not a claim that tests beat interviews or that data beats experience. It is a claim about the last step, where several pieces of evidence become one decision, and it says a formula does that step better than a judgment call.
The limits belong in the same breath. The job-performance comparison rests on 9 studies, with 1,392 people behind the mechanical estimate and 1,156 behind the judgment one. Other criteria show much smaller gaps, the authors flag a possible file-drawer problem, and under the most stringent new-sample shrinkage estimates the mechanical advantage was eliminated, though not reversed, for two criteria.
One more thing the summaries drop: the paper's own definition of mechanical is a formula applied the same way for each decision, and the first example it gives is aggregating scores with simple unit weights. A written scorecard added up identically every time qualifies. Nothing in this finding argues for a model choosing people, and nothing in it licenses a single number standing in for a candidate.
What has to be true for the rule to work?
The weights have to be fixed before anyone meets a candidate. A rule written after the finalists are known can be tuned until it selects the person the room already prefers, and it will then be defended as arithmetic, which is worse than no rule because it is harder to argue with. Name the dimensions, set the weights, date the document, and leave it alone for the requisition.
The second condition is that the rule actually runs at the end. The same meta-analysis classifies group consensus meetings as holistic combination, sitting beside individual expert judgment, because what makes a method holistic is that evidence is combined by judgment rather than by a procedure applied identically each time 1. A debrief that ends in a shared feeling is that method, whatever the scorecards said on the way in. The meeting is still where evidence gets surfaced and corrected, which is a separate and genuinely useful job, worked through in how a debrief converts opinions into a decision.
The combination step is rarer in practice than the vocabulary suggests. A content analysis of 104 interviews reported in 103 articles published between 1997 and 2010 found statistical prediction present in 11 percent of them and detailed notes in 19 percent 4, and that measures what researchers built into published studies rather than what employers do. Practice is unlikely to be more structured than the literature it is drawn from.
So the honest description of most structured processes is that they structure the gathering and leave the combining alone. That is the half that the evidence says carries the gain.
Write the override protocol, then count the overrides
Allow overrides and make them expensive to file. A written override names the dimension, the evidence that outweighs the rule, and the person accountable for the call. Then count them, because the override rate is the only measure of whether the rule is running at all. A process that overrides itself half the time is a holistic process with a document attached to it.
The override rate has been measured against outcomes at least once. Comparing managers at the same location across 15 firms hiring low-skill service workers, a one standard deviation higher rate of hiring against a job test's recommendation went with 6 to 7 percent shorter job durations, in a setting where introducing that test had raised completed job tenures by 0.23 log points, just over 25 percent 2. Not a randomized experiment, tenure rather than performance as the quality measure, and a high-turnover setting where the median completed spell ran about three months. It says the average override was worse, never that the test was right about any individual person, and overrides in that data were common rather than exceptional.
The older result in this literature is starker and worth keeping in view. Sarbin's 1943 admissions study found high school rank plus a college aptitude test correlated .45 with academic achievement, while the same two predictors plus counselors' intuitive judgment correlated .35 3. Adding the human judgment on top made prediction worse. That is an admissions study from a different era, and it is famous for exactly this reason.
The same review notes that although it is commonly accepted that some interviewers are better than others, research on variance in interviewer validity suggests the differences are attributable to sampling error 3. Your best interviewer is not an evidenced category, which matters most when that person is the one asking for the override.
Why is a total harder to contest than a finding?
Because there is nothing inside it to point at. A candidate told they came out at 74, or a manager told the rule placed someone third, can only disagree with the whole number, and an objection at that resolution reads as complaint whatever is behind it. A finding that carries the sentence it rests on can be contested in one specific place: that is not what the transcript says, or that moment was the assignment's fault rather than the candidate's.
This is where a fixed rule and a composite score get confused, and the difference is the whole argument. A rule is a written procedure for combining evidence that stays visible after the decision: anyone can see the dimensions, the weights and the ratings that fed them. A composite is the output with the inputs discarded, and it travels well precisely because it has thrown away everything anyone could interrogate.
Keep the dimension ratings and their evidence attached to whatever the rule produces, and resist publishing a total as if it measured a person. That discipline costs nothing and it is what makes an unfavorable decision explainable, to a rejected candidate, to a manager who disagrees, and to a regulator who asks why. A decision nobody can reconstruct is not defensible just because it was consistent.
The rule also does not tell you whether the evidence feeding it was any good. That is a separate question, taken up in whether assessment scores still predict performance once AI is in the workflow and in the recurring case of a candidate who interviewed brilliantly and struggled in the first quarter. A rule that combines weak evidence consistently produces consistent mistakes, which is an improvement only in that you can find them.
Common questions
Does this mean the hiring manager should not have the final call?
It means the final call should be made against a rule the manager helped write, before the candidates arrived. Accountability for the hire stays where it was. What changes is the order: the weights are argued about in the abstract, when nobody has a favorite yet, and applied afterward. A manager who wants to depart from the result can still do so, in writing, naming the evidence that outweighs the rule.
What may the rule not contain?
Anything you could not defend as job-related, and anything that stands in for a characteristic you may not consider. Culture fit, school prestige, communication style, and years of experience used as a proxy for skill all deserve scrutiny before they get a weight, because a weight makes them consistent rather than making them valid. If a dimension cannot be described in terms of the work, it does not belong in the rule. Whether a particular dimension is defensible for a particular role belongs in front of counsel.
How many dimensions should the rule have?
Few enough that each one gets real evidence in the loop. Four to six is a workable range for most roles, because every dimension needs a round designed to surface it and a rater who can describe what good looks like. Ten dimensions on a five-round loop guarantees that several will be rated from impression, and an impression given a weight is more dangerous than an impression given none.
What is a healthy override rate?
There is no published benchmark to quote, so the useful comparison is your own rate over time rather than an external number. What the rate tells you is whether the rule exists: near zero suggests either an unusually well-specified rule or a team quietly reverse-engineering ratings to reach the answer they want, and a high rate means the rule is decorative. Read the reasons, not just the count, and rewrite the rule when the same reason appears three times.
Doesn't a rule discard information an experienced interviewer has?
Some of it, deliberately. The selection literature holds that people are effective at collecting information and less effective at combining several sources of it into a final decision, which is the division of labor the rule proposes rather than a dismissal of experience. Information an interviewer genuinely has can enter as evidence on a dimension, where it is visible and can be challenged. What the rule refuses is the version that arrives only as a conclusion at the end, with nothing behind it anyone else can inspect.
References
- 1. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the .44 against .28 job-performance comparison and its sample limits, the description of the gain as a population-level improvement of more than 50 percent, the definition of mechanical as a method applied the same way for each decision, and the classification of group consensus meetings as holistic combination.
- 2. Discretion in Hiring nber.org Supports the 0.23 log point (just over 25 percent) tenure improvement from introducing a job test, and the association between a one standard deviation higher exception rate and 6 to 7 percent shorter job durations among managers at the same location.
- 3. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the Sarbin comparison (.45 for two mechanical predictors against .35 once counselors' intuitive judgment was added) and the claim that variance in interviewer validity is attributed to sampling error.
- 4. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the content-analysis finding that statistical prediction appeared in 11 percent and detailed notes in 19 percent of 104 interviews reported across 103 articles published between 1997 and 2010.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.