Assessment design

Two Graders on the First Ten, One After That

An assessment needs two graders on its first ten submissions, marked blind and separately. Then compare the two mark sheets criterion by criterion and rewrite the rubric everywhere they diverged. That repair is what buys reliability; the second pass on submission eleven does not. Once two people apply the rubric the same way, grade single with a periodic audit, and run the double-grade again whenever the assignment, the rubric or the tooling changes.

The takeAdding reviewers is the most expensive available way to avoid rewriting a rubric. Panels of three exist because disagreement is uncomfortable. Nobody measured what the third opinion bought. A rubric two trained people cannot apply the same way is not strict, it is unfinished, and averaging three readings of an unfinished rubric produces a figure with a decimal point and no shared judgment inside it. Spend the hours once, on the wording, and the reliability keeps. Every reviewer you add after that is rent.

Where Olive fits

Open a role and see what the work shows

The costly part of building this is not the marking, it is producing evidence a second reader can check. Olive returns six findings on a candidate, each written by a human reviewer and anchored to a timestamped excerpt from the session, so a disagreement about a finding has something specific to be about.

Rank your shortlist

How many graders does reliability actually take?

Two, for the first ten submissions, then one. Reliability is a property of the rubric: the degree to which two people applying it to the same work reach the same call. You find that out by testing it, and once it holds, a second reader on the eleventh submission buys reassurance and no extra accuracy.

The advice this replaces, "use multiple reviewers for fairness," spends hours without ever asking whether the reviewers agree. Three readings averaged into one mark look more careful than one reading. If the three disagree about what the criteria mean, the averaging hides that behind a decimal point.

Grading also carries consequences outside the team. The Uniform Guidelines, the 1978 federal regulation on employee selection, define a selection procedure as any measure, combination of measures, or procedure used as a basis for an employment decision, and say the definition reaches everything from performance tests through informal or casual interviews and unscored application forms 1. The scope is definitional. Validation obligations attach only where adverse impact appears. What it forecloses is the escape hatch: swapping a graded exercise for a conversation moves nothing outside the category.

Two people who disagree about a submission are almost never disagreeing about the candidate. They are disagreeing about a word in the rubric: thorough, appropriate use, senior-level judgment. That is fixable, and fixing it costs less than a permanent second reviewer (write a rubric two reviewers apply the same way).

Run the first ten blind, then compare mark by mark

Give both graders the same ten submissions with names and any disclosure of AI use stripped out, have them mark independently with no sight of each other's sheet, and collect both sheets before either grader sees the other. Then compare the sheets one criterion at a time, because two graders often reach the same overall verdict for different reasons, and the total conceals it.

Record four things on each of the ten:

  • The mark on every criterion from both graders, not a single overall verdict.
  • Every case where the two agreed on the outcome but cited different evidence.
  • The exact sentence in the rubric each grader was applying where they diverged.
  • How long each submission took to mark, which tells you what the rubric costs to run at volume.

Ten is a working minimum, and the reason to name a number at all is that teams otherwise calibrate on two submissions and declare the matter settled. Ten gives enough different failure shapes to catch ambiguous wording. If the first ten are all clean passes, keep going until several genuinely difficult ones have been through, because a rubric is only tested where it is hard.

Do this before the assessment goes live, using pilot submissions or work from people already in the role. Calibrating on real candidates means the first ten applicants through the process were graded by a rubric you later decided was wrong, and there is no way to give them that back.

What disagreement is telling you about the rubric

That a criterion has more than one reasonable reading. Every place two trained graders diverge is a place the wording did not decide the question, so the divergence sheet is a to-do list for the rubric rather than a report card on the graders. Rewrite each one until the criterion names an observable thing in the submission and states what settles it.

The federal guide to structured interviews already names this as one of three properties: the same questions in the same order, a common rating scale, and interviewers agreeing in advance on what an acceptable answer looks like 2. That guide is written for federal hiring and binds no private employer, and the third property is the one that gets skipped wherever it is applied. It is also the only one that makes the first two mean anything.

The hardest divergences cluster in one place, which is submissions that look right and are not. In a field experiment with consultants, on one task deliberately chosen to sit outside the model's capability, people using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right against 60% and 70% in the two AI conditions 3. That is one task and one sample, and the capability line moves with each model release. What it illustrates is the grading problem exactly: confident, well-presented, wrong.

So a rubric that rewards presentation will produce agreement quickly and measure nothing. Anchor each criterion to something a grader can point at: the claim that got checked against a source, the instruction the submission declined to follow, the number that does not reconcile. Where every submission arrives polished, the divergence usually sits in what graders infer from a clean surface, which the rubric should forbid them to infer at all (grade take-homes when everything comes back polished).

When to double-grade again

Whenever the thing being graded changes. A new assignment, a rewritten criterion, a new grader, a change to which tools candidates may use, or a model release that changes what a submission looks like. Each of those is a reason to run another ten. None of them announces itself as one, so put a standing date in the calendar as well.

A rolling audit costs less than it sounds. Pull one in ten graded submissions, have a second reader mark it blind, and record whether the two agree. If agreement holds, nothing happens. If it slips, the next ten go double-graded and the rubric gets another pass. That is a small permanent overhead, and what it buys is a number you have checked.

Two failure modes are worth watching for. Graders drift toward each other over time, converging on each other's habits until the rubric is no longer what either one is applying. That is why the audit reads blind. And a rubric can be reliable and still wrong: two people applying a criterion identically to something the job does not require will agree perfectly about nothing. Reliability is a floor under validity, not a substitute for it (tell a validated assessment from a bias-audited one).

Where the assessment feeds a hire-or-not call, this work and the standard are one job. A criterion nobody can apply twice cannot support a bar either, and a process that cannot state its bar falls back on comparing candidates to each other (decide whether to rank candidates or set a bar).

See what gets scored

Common questions

What counts as good enough agreement between two graders?

High enough that the disagreements left are about genuinely hard submissions rather than about wording. If you want a statistic, agreement corrected for chance is the family to look at, and any of the standard coefficients works as long as the same one is used each time. The practical test is harder to fool: pull the cases where the two graders diverged and ask whether a third person can tell from the rubric alone who was right. If not, the rubric is the problem whatever the coefficient says.

Should graders discuss a submission before or after marking it?

After, and only on the ones they marked differently. Discussion beforehand replaces two independent readings with one, because the second grader is now rating the first grader's argument. Mark blind, then hold a short conversation about the divergences, and record what changed and why. If a mark moves, it should move because someone quoted a part of the submission the other reader had not weighed, not because one of them is more senior.

Does a second grader reduce bias?

Only where the two bring different blind spots and the rubric forces both to point at evidence. Two people trained the same way, reading the same polished submissions, tend to be wrong together. What actually narrows the room for bias is stripping identifying material before marking and requiring every mark to cite a specific part of the submission. A second reader on top of that is a check. A second reader instead of it is a second opinion about the same impression.

How long does calibration take in practice?

Budget two sittings. The first is the blind marking of ten submissions, which costs each grader whatever ten submissions cost, usually a few hours. The second is the comparison and the rubric rewrite, an hour or two with the divergence sheet on the table. The rewrite is where the value sits, so protect that time specifically. Teams that run the marking and skip the rewrite have measured their disagreement and then kept it.

Can one experienced person grade everything?

They can, and nobody will know whether the results mean anything, including them. A single grader is internally consistent by definition and can be consistently applying a standard nobody else would recognise. The double-grade is what turns one person's judgment into something the organisation owns: afterwards the criteria are written well enough that a second person could take over, which is what happens anyway when the experienced grader leaves.

Do candidates ever see the rubric?

They should see the criteria, if not every anchor. A criterion drawn from the work is not a secret worth keeping, and a candidate who prepares against it is preparing to do the job. Publishing it also disciplines the writing: a criterion you would be embarrassed to show a candidate is usually one your graders cannot apply either. Keep back the specific answer key, publish what is being looked for.

References

  1. 1. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that a graded exercise and a casual conversation are the same category of thing under federal selection law, stated as definitional scope rather than an audit duty.
  2. 2. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the claim that agreeing in advance on what an acceptable answer looks like is part of the standard definition, not an optional refinement. Federal HR guidance, binding on nobody in the private sector.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the submissions hardest to grade are the confident, well-presented wrong ones, using the single out-of-frontier task where AI-assisted consultants did worse.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.