Interviewing

A Scorecard Holds Evidence, a Call, and Nothing Else

A hiring scorecard needs five things: the job-relevant claims an interview round was meant to test, what the candidate said or did against each one, a line naming where that observation came from, a row for the claims the round never reached, and a single call. Everything else on a standard template, the weights, the five-point scales, the optional notes box, is arrangement. A card is worth filling in when somebody who was not in the room can reconstruct the call from what is written above it.

The takeThe rating is the least useful field on the card and it is the only one most templates make mandatory. Keep it if the system demands a value in that column, then treat it as a label on the evidence rather than a summary of a person. If a card cannot survive having its number deleted, the number was carrying weight that nothing underneath it earned, and the debrief is about to spend forty minutes arguing with a digit.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing how they would check a confident claim, and a scorecard then records the description. Olive puts the checking itself in front of them as work, a 40-to-60-minute occupational assignment run with an AI assistant, and returns six findings with a timestamped excerpt behind each one.

Rank your shortlist

What belongs on the card, in order?

Five fields, in this order: the claims the round was assigned to test, the evidence for each, where it came from, the claims the round could not reach, and one call with its reason. Two of those, the provenance line and the coverage-gap row, are missing from the standard template, and they are what make the card readable to anyone who was not there. The rating, if you keep one, goes last, because everything above it produces it.

  • Claims, not competencies. "Can scope an ambiguous request before starting" beats "Problem Solving". Four to six of them, written before the round, taken from the job the person will actually do.
  • Evidence, quoted. The sentence they said, the step they took, the thing they refused to do. Close paraphrase is fine. "Strong communicator" is not evidence and never was.
  • Provenance. One short line per claim: watched live, read in a submitted file, relayed by another interviewer, produced by a tool.
  • Coverage gaps. The claims this round could not reach, named, so the loop knows which ones still need a round.
  • The call. Hire or no hire, plus the one or two claims it rests on.

What the notes name turns out to matter. Text-mining the post-interview notes written about 7,650 candidates who were hired at a large Chinese internet technology company, researchers found that the number of job-related capabilities an interviewer named in the notes tracked the candidate's later performance and promotions and ran against turnover, with a one standard deviation rise in the match between notes and job analysis corresponding to roughly a 2 percent rise in performance 1. Small, correlational, one firm, and visible only among people who were hired. Note-taking on its own was never the variable there, which is what makes the narrow version worth having: what an interviewer writes down about the capabilities the job called for carries something a rating does not.

Why does the evidence box have to come first?

Because position on the form decides what gets written. A field placed last and marked optional collects whatever attention is left after the required ones are done, and on the templates in circulation that field is the notes box. The card that leaves the room is then a row of ratings with nothing underneath them, and a debrief cannot argue with a rating. It can only answer a 4 with a 3.

This is not only a template problem. A content analysis of 104 interviews reported in studies published between 1997 and 2010 found that they used an average of 5.74 of the 15 structure components in Campion, Palmer and Campion's framework, and that the evaluation-side components were the thin ones: detailed notes were described in 19% of them 2. Each component was coded from what the article itself described, so part of that 19% is incomplete reporting. It also counts what researchers built into published studies rather than what employers do, so read it as a ceiling. Practice is unlikely to be more structured than the literature it is drawn from.

Making the evidence field required has a second effect worth having. An interviewer who has to write down what was said has to have listened for something specific, which pushes the work back into the round design where it belongs. It is the same discipline that decides whether two people reading the same answer land in the same place, which is worked through in writing a rubric that two reviewers score the same way.

Add a line naming where you saw it

Add the one field standard templates omit: how each observation reached you. A claim you watched a candidate make and a claim a summary reported to you are not the same evidence, and once both are typed into the same box they stop being distinguishable. Four values cover nearly everything. Watched. Read. Relayed. Tool-produced. It costs an interviewer one word a claim.

The field earns its space because loops now mix rounds carrying very different amounts of human observation. An async video round, a vendor-assessed exercise and an automatic summary all emit something that looks like a finding, and by the time three of them are in one applicant record they read as equivalent. The authors of the 2022 revision of the standard selection-validity estimates, whose own table ranks structured interviews first, note that as administration and scoring move toward automation the interview becomes more amenable to running at volume, and that validity needs to be evaluated under those changed circumstances 3. That is an open question flagged by the people best placed to flag it, evidence for nothing in either direction, and precisely the kind of thing a provenance line keeps visible instead of averaging away.

Provenance also hands the debrief a cheap diagnostic. When two people rate the same claim differently, the first question stops being who is right and becomes whether they were looking at the same material. That conversation is shorter and it usually ends somewhere useful. Two finalists whose work came back at the same quality is this problem with the evidence already gathered and the provenance still missing.

Write down what the round could not see

Give the card a row for coverage gaps and require it. Every round runs out of time, every panel skips a claim, and the standard template has no way to say so, which leaves a blank cell that reads as an oversight and a rated cell that reads as an observation. Neither is true. A named gap is the only version of this a later reader can act on.

The reason it matters is arithmetic rather than diligence. Six claims on the card, four reached in the round, and two ratings appear anyway because the field was required and the interviewer was reasonable about it. Those two now travel with the same weight as the four, and no reader downstream can tell them apart. Marking them not observed costs nothing and it exposes the pattern that a hiring team most needs to see: the same claim going unreached in three consecutive rounds, usually the awkward one nobody wants to ask about.

On Monday, the change is small and takes one edit. Open the template. Move the evidence field above the competency list, mark it required, add a provenance line, add a coverage-gap row. Keep the rating if the system insists on a value in that column, and stop treating it as the artifact. Then pull one filled-in card from last month and ask whether the call is reconstructable from what is written above it. If it is not, the template is still the problem, and the next card will be no better than the last one.

See how it works

Common questions

How many claims should a scorecard cover?

Four to six per round, and fewer if the round is short. The constraint is not attention span. It is that every claim on the card needs a question or a task behind it capable of producing evidence. A round listing eight competencies and asking four questions will come back with eight ratings, and four of them will be invented. Write the claims first, check that the agenda reaches each one, then delete whatever it cannot reach and hand that claim to another round.

Should competencies be weighted?

Weighting is harmless and it is rarely the binding constraint. It matters when several people fill in the same card and something downstream adds the values up, because then the weights are the rule and they deserve one argument in advance. It matters not at all when the card ends in a conversation, which is the usual case. If the debrief is where the combining happens, spend the argument on what counts as evidence instead.

Should the candidate see the scorecard?

Most teams do not share the card itself, and the more useful test is whether it could be shared without embarrassment. A card written as specific observations survives being read by the person it describes. A card that says 'not a culture fit' or 'felt junior' does not, and that discomfort is a signal about the card rather than about the candidate. Write every line as though it will be read aloud, because in a dispute it may be.

What goes on the card when the interviewer never observed an assigned claim?

The words not observed, and nothing else. A blank cell reads as an oversight and a rated cell manufactures evidence that will sit beside real evidence later with no way to tell them apart. Add the reason beside it if there is one, short: ran out of time, the candidate took the question somewhere else, the exercise never reached it. That line is what tells whoever runs the next round whether to pick the claim up or drop it from the loop.

Where does a tool-produced value or summary belong on the card?

In its own row, labelled with what produced it, never inside the interviewer's evidence field. The point of the card is that a reader can tell which observations a person made. A value from a screening tool or a summary from a notetaker mixed into that box destroys exactly that, and it is unrecoverable afterwards. Keep it visible, keep it separate, and say in the debrief which of the two you are quoting.

References

  1. 1. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company Frontiers in Psychology, Volume 11, Sec. Organizational Psychology (Shanshi Liu, Yuanzheng Chang, Jianwu Jiang, Haigang Ma and Huaikang Zhou), 2021. frontiersin.org Supports the claim that notes naming the capabilities the job called for track later performance, promotions and retention, including the 7,650-candidate sample and the roughly 2 percent figure.
  2. 2. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature Personnel Psychology, 67(1), 241-293 (Levashina, Hartwell, Morgeson and Campion); author copy opened at morgeson.com, 2014. doi.org Supports the claim that the evaluation-side components of interview structure are the rarest, with detailed notes described in 19% of the 104 interviews content-analysed and an average of 5.74 of the 15 components in use.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the point that the authors who ranked structured interviews first flag automated administration and scoring as a setting where validity still needs to be evaluated, which is why provenance belongs on the card.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.