Assessment design

Send the Criteria, Withhold the Weights and the Key

Send candidates the criteria from your rubric with the exercise brief; keep the weights and the worked answer back. A candidate should know what the exercise is for, what a reviewer will look at, and roughly how long it should take, because that is what makes submissions comparable, and none of it is a specification to build to. The weight on each criterion and the answer you consider correct are the two things that get optimized to instead of met, so those stay with the reviewer.

The takeThe advice to publish the rubric comes from university teaching centres, and it was right for the world it was written in. A student could not optimize perfectly to a published rubric inside an exam week, so publishing raised the floor without flattening the top. That constraint is gone. Hand a model a rubric and a brief and it returns a submission satisfying every line, which means a published rubric now sorts candidates mostly by whether they thought to paste it in.

Where Olive fits

Open a role and see what the work shows

The split between what candidates are told and what stays with the reviewer is the hard part of building this in-house. Olive names the six dimensions its reviewers write against, and the candidate is granted the identical report the employer receives, free, on every tier.

Rank your shortlist

Should you send the rubric with the brief?

Send part of it. The criteria, yes: naming what a reviewer will look at costs you nothing and improves the comparability of what comes back. The weights and the worked answer, no. Those two together are a specification, and a specification handed to a candidate working with a model produces a submission built to satisfy it rather than one that shows how the person thinks.

Search this question today and teaching-centre pages fill the results. None of them was written about hiring, and none was written for a candidate who can hand the rubric to a model and get back a submission that answers every line of it. A course carries what a hiring exercise does not: an instructor who already knows the student, a term of prior work standing behind the submission, and a curve. A brief has one artifact from one stranger, so the published half of the rubric carries far more of the weight.

The anchored half of a rubric does buy something, and what it buys goes to the people doing the judging. In the meta-analysis reported in Levashina and colleagues' review, past-behaviour interviews using anchored rating scales showed higher criterion-related validity than those without, .35 against .26, and slightly higher interrater reliability, .77 against .73 1. Nineteen studies, one question type, and a comparison across studies rather than a controlled test, so hold it loosely.

Anchors pay off where two reviewers have to agree about one submission, and that agreement is settled before the brief ever goes out. So they belong in the scoring pack.

What belongs in the brief and what stays back?

The brief carries everything a candidate needs to aim. The scoring pack carries everything a reviewer needs to judge. That split is clean enough to apply line by line, and applying it to a rubric you already have takes about twenty minutes.

In the brief:

  • what the exercise is for, and where it sits in the process
  • the deliverable, its format, and any hard bound on length
  • the intended time, and a sentence saying extra polish earns nothing
  • the criteria, in plain sentences a person could repeat back
  • whether an assistant is permitted, and whether to describe how it was used
  • who reads the submission, and roughly when

In the scoring pack:

  • the weight on each criterion
  • anchored descriptors for each level, written from real past submissions
  • the reference answer, and the outcome the team actually reached
  • the rule for two reviewers who disagree
  • the bar, and what happens on either side of it

The second list is the one that has to be reliable, because two reviewers reading the same submission should arrive at the same finding. That is a harder problem than it looks, and it is worth solving before the first candidate sees the brief, which is the whole of writing a rubric two reviewers score the same way.

Why does a published rubric stop sorting anyone?

Because it turns a judgment task into a checklist, and a checklist gets completed rather than reasoned about. In a study where engineering students used ChatGPT on take-home open-book exams and submitted their transcripts, some of the strongest evidence of reasoning appeared where students evaluated the model's incorrect or incomplete answers, and the author concluded that correctness of the final answer alone may no longer be sufficient evidence of comprehension 2.

That is one qualitative study of a single course, with no control group and no grades comparison, and transcripts bring problems in hiring that a classroom never faced: a log can be curated afterwards, and demanding one is a disclosure requirement a candidate may reasonably refuse. When the answer is cheap, the judging is what remains worth reading.

The tell that you shared too much is convergence. Submissions start arriving with the same section headings in the same order, each criterion addressed under its own subheading, every box ticked and nothing argued. At that point the exercise measures who thought to paste the rubric in.

Two repairs, in order of cost. Ask for something a rubric cannot enumerate: the option that was rejected and why, the claim that was checked and what it was checked against, the part the candidate would not hand to an assistant. And stop scoring completeness, which is the criterion a published rubric optimizes hardest. Scoring an AI-assisted answer is mostly a question of what you refuse to give credit for.

Give the rubric to the candidate after the decision

Parity belongs on the other side of the decision, and it is a different act from publishing the key in advance. A rejected candidate should be able to see what the reviewer saw: the criteria, the finding on each one, and the passage in their own submission each finding rests on. The exercise is over by then, so there is nothing left to optimize to and nothing left to protect.

It also changes how the outcome lands. In a vignette experiment with 921 working-age Austrians, a rejection by AI with no explanation scored lowest on every outcome measured, while an AI rejection that came with an explanation drew the same fairness ratings as a human rejection that did not 3. Hypothetical rejections imagined by an online panel, every condition below the scale midpoint, differences of fractions of a point: this is a ranking among unhappy outcomes rather than a route to a happy one.

What it supports is narrow and worth acting on. The explanation does more work than the identity of whoever decided, and an unexplained automated rejection is the worst version measured. Three sentences of evidence beats a bare decision, and the material already exists in the scoring pack you kept back. Explaining a rejection that involved AI is the same content with the weights removed. Whether a candidate-facing report lowers your legal exposure or raises it is a separate question, and it is worth putting to counsel before the first one goes out.

So split the rubric into the half you would happily print in the brief and the half you would not. Send the first half with the exercise. Send the second half, rewritten as findings against your own criteria, to the person you turned down.

See what gets scored

Common questions

Do candidates do better work when they see the criteria?

They produce work that is easier to compare, which is the benefit worth having. Naming the criteria removes the guesswork about what the exercise is for, so submissions differ on judgment rather than on how each person imagined the assignment. What it does not do is raise the ceiling. The strongest submissions were already aimed at the right thing, and publishing the criteria mostly pulls the weaker ones onto the same target.

What do I say if a candidate asks for the full rubric?

Say what you share and why, in one sentence: the criteria are in the brief, the weights and the reference answer stay with the reviewers so submissions are compared rather than reverse-engineered, and the full findings go to the candidate after the decision. That answer is honest, it is checkable against what you actually do, and it treats the request as reasonable, which it is. Refusing without a reason is what makes the withholding look like something else.

Is withholding the weights unfair to candidates?

Not if the criteria are published and the findings are shared afterwards. Fairness in an assessment means knowing what is being asked, being measured on the same standard as everyone else, and being able to see the reasoning behind the outcome. None of those requires the weights in advance. What would be unfair is scoring on a criterion the brief never mentioned, which is the actual failure the transparency advice was written to prevent.

How do I tell whether I shared too much?

Watch the shape of what arrives. Submissions converging on the same structure, the same headings and the same length, each criterion answered under its own subheading, mean the brief has become a template. Two more signals: reviewers agreeing easily but finding nothing to argue about, and the pass rate climbing without the hires getting better. Any of the three is a reason to move a piece of the rubric back into the scoring pack.

Should the brief say whether AI is allowed?

Yes, in one sentence, either way. An unstated rule is still a rule in the reviewer's head, and candidates who guess conservatively are penalised for guessing right about your etiquette rather than about the work. If assistants are permitted, say whether you want the candidate to describe how one was used. If they are not, say what you will do about it, because a rule with no stated consequence gets read as advice.

References

  1. 1. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature Personnel Psychology, 67(1), 241-293 (Levashina, Hartwell, Morgeson and Campion), reporting Taylor and Small (2002), 2014. doi.org Supports the claim that anchored rating scales pay off on the reviewer's side: .35 against .26 for validity and .77 against .73 for interrater reliability across 19 past-behaviour interview studies.
  2. 2. Reimagining Assessment in the Age of Generative AI: Lessons from Open-Book Exams with ChatGPT arXiv:2605.12363 (Mahmoud, single author), 2026. arxiv.org Supports the claim that when the tool is permitted, correctness of the final answer stops being sufficient evidence, and the judging becomes the readable part.
  3. 3. Rejected by an AI? Comparing job applicants' fairness perceptions of artificial intelligence and humans in personnel selection Frontiers in Artificial Intelligence, 2025. frontiersin.org Supports the claim that an unexplained automated rejection is the worst case measured, and that an explained one rates alongside an unexplained human rejection.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.