Assessment design
How Do You Write an AI-Use Rubric Two Reviewers Score the Same Way?
An AI-use rubric that two reviewers score the same way names acts, not quality. Give each dimension three levels and write each level as something a reviewer can point at in the role's own material: the source behind a headline statistic opened, a failing test written before the fix. Rate one dimension at a time across candidates, score independently before anyone discusses, calibrate on three real sessions. Then measure kappa per dimension, never averaged; under about thirty double-marked sessions, call it a calibration check rather than a reliability estimate.
The takeReliability is the easiest thing in an assessment to improve and the easiest to mistake for validity. Two reviewers can be trained into near-perfect agreement about a finished deck, and the figure that comes out will look like rigor while describing a document showing none of the acts the rubric names. My read is that most teams who reach a respectable kappa got there by tightening the wording rather than by capturing more of the session, because wording is free and capture is not. Agreement is cheap. Evidence is the part you have to go get.
Where Olive fits
Open a role and see what the work shows
If you are building this yourself, the expensive parts are the answer key and the evidence trail behind each rating. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer and anchored to a timestamped moment in the session, with the candidate granted the identical document.
Rank your shortlistWhy Two Reviewers Disagree About the Same Session
Because the scale asks for a judgment nobody defined. 'Uses AI effectively, 1 to 5' is an empty container each reviewer fills from their own practice, so a session that looks disciplined to someone who works that way looks slow to someone who does not. Reliability, in assessment terms, is exactly this: rating consistency among the people doing the rating 1.
The federal structured-interview guidance is blunt about what raters need instead: a concrete behavior written out for each proficiency level to refer to rather than an adjective, and a note documenting the reasoning behind every rating 1. A one-line dimension name is neither. It is a prompt for an opinion, and two competent people will hold different ones.
Three things go wrong, and each has a different fix.
- The scale rates the artifact. A finished deck, a clean pull request and a fluent memo cost almost nothing to produce now, so a rubric that rewards a good deliverable is measuring the assistant. Reviewers who weight the artifact differently will disagree forever, because the artifact was never the thing in dispute.
- Volume reads as effort. Forty prompts look like engagement to one reviewer and thrashing to another, and both readings are defensible under a scale that never said which. How much AI someone used is not a level of anything. A candidate who decided the model was the wrong instrument for a step and did it by hand has done the thing you are trying to see.
- Halo travels down the grid. Read one candidate end to end and the first strong rating colors the next three. This is the cheapest of the three to fix and almost nobody does it: rate every candidate's first dimension, then every candidate's second, and never open a session already knowing what you gave it last time.
Getting a whole panel to converge is this same problem one layer out, worked through in How to Get an Interview Panel to Judge AI Use the Same Way.
Write the Anchors in the Role's Own Evidence
An anchor is a sentence describing what the act looks like in this occupation, with the artifact that proves it happened. 'Demanded evidence' is unscorable. 'Opened the source behind the statistic the recommendation rests on, found the sample was three years old, and changed the claim' is scorable, because two reviewers can point at the same moment and agree it is or is not there.
The method for writing them is old and dull and it works. Federal structured-interview guidance has subject-matter experts sit down per competency and write, for each proficiency level, how an actual employee at that level would answer. The stated purpose is to give raters concrete behavior to refer to rather than an adjective, and so a common frame for reading different candidates 1. Do that, one dimension at a time, with people who hold the job.
The anchors cannot be generic, because the act is not generic. Take one dimension (tested a claim against something outside the conversation) across three roles.
| Role | Not demonstrated | Partly demonstrated | Demonstrated |
|---|---|---|---|
| Marketing | The headline statistic goes into the brief as the assistant supplied it | The source is named in a footnote but never opened | The source is opened, the sample turns out to be a 2019 consumer panel, and the claim in the brief changes |
| Financial analysis | The growth rate the model produced is carried into the memo | One input is spot-checked against the filing | The figure is rebuilt from the filing, comes out lower, and the recommendation moves with it |
| Software engineering | The generated patch ships because it compiles and the suite is green | The existing tests are run against the change | A test reproducing the reported bug is written first, watched to fail, then made to pass |
That engineering anchor is written where the defect actually lives. In Stack Overflow's 2025 developer survey the most common frustration with AI tools, at 66%, was output that is almost right but not quite, and 46% of developers said they distrust the accuracy of what the tools produce against 33% who trust it 6. Almost-right compiles. It passes the happy path. It fails the case nobody wrote a test for, which is why the anchor names the test and not the diff.
Two rules keep an anchor scorable. Every level describes an act with a before and an after, so a reviewer can point at where it happened or concede it did not. And no level mentions how much the assistant was used, in either direction. Anchoring outside engineering takes more work rather than less, and How to Screen for AI Judgment in Finance and Marketing Roles works through where the evidence hides in those roles.
What Goes on the Scale: Four Dimensions, Three Levels
Four dimensions, three levels, and no average across them. Three levels (not demonstrated, partly demonstrated, demonstrated), because a five-point scale invites the middle and forces distinctions the record cannot carry. Four dimensions because each has to be independently visible in the session, and the fifth candidate dimension is almost always a restatement of one already on the sheet.
Four that hold up across roles, each with the question it asks and the evidence that settles it.
- Framing. Did the first move go after understanding the problem, or straight at the deliverable? Settled by the opening exchange, before any output exists.
- Evidence. Was a source demanded for the specific claim the answer rests on, and opened? Settled by what was retrieved, not by what was cited.
- Boundary. What was kept, and what was handed over? Settled by the acts in the record the candidate did themselves.
- Refusal and verification. Was anything the assistant produced rejected on substance, and was any claim tested against something outside the conversation? Settled by a change you can trace: a section cut, a number moved, a limit added.
Rate refusal as a rate rather than a count, refusals per answer actually taken up, or three well-aimed prompts lose to thirty scattered ones for a reason nobody would defend out loud.
Do not average the four into one figure. The average is the part reviewers disagree about least and the part that tells you least: two candidates with the same mean can be opposite people, and the disagreement between your reviewers, which is the information you were collecting, vanishes into it. There is a paperwork cost too. Where a selection procedure assigns weights to its parts, the documentation standards expect the weights and the validity of the weighted composite to be reported 5. Scoring a single AI-assisted answer runs on the same rules, and How to Score an Interview Answer Produced With AI covers that case.
Run the Calibration Session: A 90-Minute Script
Both reviewers score the same three sessions alone, then meet for ninety minutes and work only the disagreements. Pick sessions spanning the range you expect (one weak, one middling, one strong), and require a note citing the moment behind every rating, because a rating with no cited moment cannot be argued about, only defended.
The note is not paperwork. The federal guidance asks for notes of sufficient quality and quantity to document the reasoning behind each rating on each competency, and treats them as the record supporting the employment decision 1. In a calibration session they are the whole substrate. Without them the meeting is two people restating impressions at each other.
- 0-10 minutes: compare grids only. Both scoring sheets side by side with the reasons covered. Count exact matches per dimension. That number is the honest starting point and it is usually worse than either reviewer expected.
- 10-45: work the splits, one dimension at a time. Take each disagreement and have both reviewers read out the moment they rated. One rule makes this productive: whoever changes their mind has to say which sentence of the anchor moved them. If neither can point at a sentence, the anchor is at fault rather than the reviewer.
- 45-70: rewrite what split. Edit the anchor in the room, and write the disputed session's actual behavior into it as the worked example. An anchor that survived a real disagreement is worth ten written in the abstract.
- 70-85: re-score session three under the new wording. If agreement does not move, the dimension is doing two jobs at once. Split it or cut it.
- 85-90: version the rubric and log what changed. Every rating from here carries the anchor version it was made under. A rubric that changed mid-round cannot be compared across the candidates who took it before and after.
Then keep the order. Each reviewer observes, records and rates independently, and only once both grids exist do they discuss and settle a rating together 1. Discussion before scoring does not produce agreement. It produces one reviewer's opinion held by two people, and it will read as agreement in every measurement you take afterwards.
Recalibrate when a reviewer joins, when the task changes, and quarterly regardless. Run the whole thing over historical sessions before it gates anybody: How to Pilot an Assessment Before Making It a Hiring Gate covers the shape of that trial.
Measure Agreement Before You Trust the Rubric
Count agreement with Cohen's kappa, per dimension, never averaged across them. Two reviewers who both mark 'demonstrated' most of the time will match on raw percentage while agreeing about almost nothing, which is the reason kappa exists: it subtracts the agreement that chance alone would have produced 3.
The benchmarks worth quoting in a review: 0.60 to 0.79 is moderate agreement, 0.80 to 0.90 is strong, and any kappa below 0.60 indicates inadequate agreement, with little confidence to be placed in what rests on it 3. Per dimension. An average across four will happily hide one sitting at 0.2, and that dimension is where every disputed decision comes from.
The sample size is where most teams find out they cannot make the claim yet. The same paper puts the heuristic floor at about 30 comparisons 3. Thirty double-marked sessions per dimension is more than a first round produces, so if you have eight, report exact agreement and the list of what split, and call it a calibration check rather than a reliability estimate. A kappa computed on eight sessions carries no information.
The regulatory expectation runs the same direction and is modest. Where a procedure rests on content validity, its reliability should be a matter of concern to the user, and appropriate statistical estimates of it should be made whenever that is feasible 2. The documentation standards expect records showing the ratings given by each rater to be kept and produced on request 5. And a scored rubric that decides who advances is a selection procedure, which puts the burden of showing it is job-related and consistent with business necessity on the employer 4.
One limit sits underneath all of it, and it decides what your agreement is worth. A rubric can only rate what the record contains. If what reaches your reviewers is a finished artifact plus the candidate's own account of how it was made, two reviewers can reach a kappa of 0.85 about a document showing none of the four acts the rubric names. Agreement on the wrong material is still agreement, and it is the failure that survives every calibration session, because calibration measures whether two people read the same evidence the same way, never whether the evidence was there.
Common questions
How many reviewers should score each session?
Two, independently, for as long as you can afford it. One reviewer is a rubric with no reliability evidence at all. Once agreement holds across a round, drop to single-marking with a sample double-marked: every fifth session, plus every session that leads to a rejection. Keep both grids. The second reader is not there to make a better call on that candidate; it is the only way to know whether the first reader's ratings mean what the rubric says they mean.
Should reviewers discuss a session before scoring it?
No. Score independently, then discuss. Talking first produces a single opinion held by two people, and every agreement figure computed afterwards measures the conversation rather than the rubric. The order in the standard guidance is deliberate: each rater observes, records and evaluates alone, and consensus comes after both sets of ratings exist. Discussion is where disagreement gets resolved, not where it gets prevented.
What kappa is high enough to stop calibrating?
0.60 per dimension is the usual floor, and 0.80 is what you would want on a dimension that gates a hiring decision. Compute it for each dimension separately, because an average across four hides the one that is broken. Below 0.60 the honest reading is that the two reviewers are measuring different things, and the fix is in the anchors rather than in more training. Report the sample size beside the figure; under about 30 double-marked sessions, the number is not yet an estimate.
Can I use a 1-5 scale instead of three levels?
You can, and agreement usually drops. Five levels ask reviewers to separate a 3 from a 4 on evidence that rarely supports the distinction, and the middle absorbs everything ambiguous. Three levels force the reviewer to decide whether the act is in the record or not. If a role genuinely needs finer resolution, add a dimension rather than more points on the same one.
Does a homegrown rubric need to be validated?
If it decides who advances, it is a selection procedure, and the employer carries the burden of showing it is job-related and consistent with business necessity. That does not mean a criterion study on day one. It does mean a job analysis you can show, anchors traceable to duties in that job, records of who rated what, and reliability evidence once the volume exists to compute any.
What if two reviewers keep splitting on the same dimension?
The dimension is doing two jobs. 'Used the assistant well' bundles framing, evidence and refusal, and reviewers split on which one they weighted. Read the disputed notes and check whether the two are describing different moments in the session. If they are, split the dimension in two and re-anchor both halves. If they are describing the same moment and reading it differently, the anchor's wording is the problem and one of them should rewrite it before the next round.
References
- 1. Structured Interviews: A Practical Guide ✓ opm.gov Defines reliability as rating consistency among interviewers; has subject-matter experts write example behaviors for each proficiency level so raters have concrete behavior to refer to rather than adjectives; requires notes of sufficient quality and quantity to document the reasoning for each rating on each competency; and has panel members individually observe, record and evaluate before discussing and reaching consensus.
- 2. 29 CFR 1607.14 - Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) ✓ ecfr.gov Section 1607.14(C)(5): the reliability of a selection procedure justified on content validity should be a matter of concern to the user, and appropriate statistical estimates of reliability should be made whenever feasible.
- 3. Interrater reliability: the kappa statistic ✓ pmc.ncbi.nlm.nih.gov Kappa corrects raw percent agreement for the agreement two raters would reach by chance; the interpretation table puts 0.60-0.79 at moderate and 0.80-0.90 at strong, states that any kappa below 0.60 indicates inadequate agreement among raters, and gives a heuristic floor of about 30 comparisons.
- 4. Employment Tests and Selection Procedures ✓ eeoc.gov A test or other selection procedure used in an employment decision must be job-related and consistent with business necessity, and the employer carries that burden.
- 5. 29 CFR 1607.15 - Documentation of impact and validity evidence (Uniform Guidelines on Employee Selection Procedures) ✓ ecfr.gov Requires records showing the ratings given to each sample member by each rater, and where weights are assigned to different parts of a selection procedure, the weights and the validity of the weighted composite are to be reported.
- 6. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.