Assessment design
Write the Answer Key Before You Send the Assignment
Every take-home needs an answer key, and writing one starts with doing the assignment yourself inside the time budget you set. That draft is the key. Around it, write the three or four decisions where strong answers legitimately diverge, the specific errors that would matter downstream, and a short list of what you are not grading. If you cannot finish inside the budget, neither can the candidates, and that is the first thing the key tells you.
The takeA rubric on its own is half an instrument. It tells a reviewer what to look at and nothing about what good looks like, so the first submission of the round quietly becomes the standard and every later candidate is graded against a stranger. The advice stops at the rubric because a template is a download and a key is an afternoon. Spend the afternoon. It outlives the round, and it is the only part of grading that survives a reviewer leaving the team.
Where Olive fits
Open a role and see what the work shows
Every Olive report is written by a person reading the session against the bank's answer key, and the six findings come back marked demonstrated, partly demonstrated or not demonstrated rather than as a number. The candidate is granted the identical report, free, on every tier.
Rank your shortlistDo the assignment yourself first
Sit down and do it, inside the budget you are about to hand a candidate, with no interruptions and nobody to ask. The draft you produce is the first version of the key, and the hour or two it costs is the cheapest quality check in the process. Two answers fall out immediately: whether the task fits the time, and which parts of it are genuinely hard.
Keep a running note while you work. Where you paused, what you had to look up, the assumption you nearly got wrong, the point where you decided the task was finished. Those notes become the divergence list and the error list later, and they are almost impossible to reconstruct afterwards from the finished draft.
Have one other person do it too, ideally somebody at the level you are hiring for rather than the person who wrote the assignment. An author cannot see their own ambiguity. If the two of you produce materially different work, the gap is in the brief rather than in the key, and it has to be closed there before either document is worth writing.
The budget test is the harshest and the most useful. Finish at the cap, hand in whatever exists at that moment, and look at it. That is roughly the ceiling of what you can expect, and if it embarrasses you, the assignment is too big rather than the candidates too weak.
What goes in the key?
Four things: a reference solution, the decisions where strong answers legitimately diverge, the specific errors that matter downstream, and an explicit list of what you are not grading. The last one does more work than it looks like, because it is what stops a reviewer marking somebody down for a formatting preference that nobody ever wrote down.
- The reference solution. Your own draft, cleaned up only enough to read. Not an ideal answer, a realistic one produced under the same constraint, or the bar drifts upward the moment it leaves the page.
- The legitimate divergences. Three or four decisions where a good candidate could reasonably go either way: which segment to focus on, whether to model the edge case or flag it, how much of the budget to spend on validation. For each, write what makes either choice defensible. This is the list that stops a reviewer penalising an approach they would not have taken.
- The errors that matter. Not every mistake, only the ones with a downstream cost: the misread column, the assumption that breaks at scale, the recommendation the data does not support. Name them concretely. A key that says "analytical errors" is a rubric again.
- What you are not grading. Prose polish, chart styling, file naming, whether they used your preferred library. Write it down and reviewers will actually honour it.
The US Office of Personnel Management's guide to structured interviews defines structure by three properties: every candidate gets the same questions in the same order, every candidate is rated on a common scale, and the interviewers agree in advance on what an acceptable answer looks like 1. That third property is the key, and it is the one almost every take-home process skips. The guide is federal HR practice advice from 2008 rather than a rule that binds anyone, and it predates every AI question by more than a decade, but the definition has not aged.
Keep it to two pages. A key nobody rereads under time pressure is a document you wrote for yourself.
Why does a rubric drift without one?
Because a rubric names dimensions and leaves the scale empty. "Depth of analysis, 1 to 5" is a question rather than an answer, and a reviewer with no reference fills that scale from the submissions in front of them. Whichever one is read first becomes the middle of the scale, and everything after it is positioned relative to a candidate instead of relative to the job.
Reviewers are also very good at making sense of whatever they are given, which is the part that makes drift invisible. In a controlled study of unstructured interviewing, predictions of a classmate's GPA made after an interview correlated .31 with the real figure, against .65 for prior cumulative GPA alone, and 96 of 169 participants chose to run an interview in which the answers were generated at random over running no interview at all 2. That is undergraduates in a lab, not managers grading work, and the effect size does not transfer. The mechanism does: given material with little signal in it, a reviewer builds a coherent story anyway, and confidence rises while accuracy does not.
Polish is the version of this that shows up in every AI-era take-home. It reads like a signal and it is not one. Asked to tell GPT-3 text from human writing across stories, news and recipes, untrained evaluators performed at chance, and three quick training methods lifted them only to about 55%, inconsistently across the three domains 3. That study is from 2021, on short passages, with crowdworkers rather than domain reviewers, so it is not a measurement of your team reading work in your own field. It is still enough to keep "this reads like a model wrote it" out of the key, which is exactly the trap when every submission comes back polished.
The key closes the gap by fixing the standard before the first candidate exists. It cannot make a bad assignment informative, and it will not settle every disagreement. What it does is make the disagreements about the work rather than about the order the files were opened in.
Grade the first submission twice
Grade it, put it away for a day, and grade it again from the key without looking at the first pass. Agreeing with yourself is the minimum bar a key has to clear. Disagreeing by more than one band means it is not specific enough yet, and the repair is to write down what the second pass noticed that the first did not, then add that to the key.
Once the key survives you, run it past a second reviewer on the same submission, blind. The conversation afterwards has one rule that makes it worth having: the disagreement edits the key, not the score. A reviewer who marked lower because the candidate skipped a validation step has either found a missing line in the errors list or an assumption of their own, and the discussion is about which. Two reviewers who agree without a key have usually agreed about the candidate's writing.
This is where an AI-open assignment needs slightly more from the key than a traditional one, because the thing worth grading moved. Once an assistant can produce the deliverable, what separates submissions is which claims got checked, which suggestions got rejected, and what the candidate did with the parts the assistant handled badly. Those need their own lines in the key, written as concretely as the error list, or reviewers fall back on prose quality. The mechanics of getting two people to score that consistently are worth their own treatment: see writing an AI-use rubric two reviewers score the same way.
Rewrite the key at the end of every round. A round you actually graded will have surfaced divergences you did not anticipate and errors you never thought to list. Fifteen minutes after the last decision is the only moment you will remember them.
Common questions
How long should writing an answer key take?
Roughly the time budget of the assignment plus an hour. A two-hour exercise costs about three hours to key properly, most of it spent doing the work rather than writing the document. That is a one-time cost per assignment, not per candidate, so it pays back on the third submission and every one after. If keying an assignment takes a full day, the assignment is too large to be a screening step.
Should the answer key be shared with candidates afterwards?
Sharing the reference solution after a round closes is generous and usually safe, as long as you plan to retire that assignment. Sharing it while the assignment is still in rotation guarantees it circulates. A better middle path is to send each candidate the two or three specific observations their submission produced, which is more useful to them than a model answer and costs a reviewer five minutes.
What if a candidate solves it a way the key does not cover?
Grade the approach on its own terms and add it to the divergence list the same day. A key is a record of what you have seen work, not a boundary on what can. The failure mode to watch is a reviewer marking an unfamiliar approach down for unfamiliarity, which is why the divergence list exists at all. If the new approach is better than the reference solution, replace the reference solution.
Does an answer key work for open-ended assignments?
Yes, and it matters more there. For a strategy memo or a design exercise there is no single right answer, so the key holds the questions a strong answer has to engage with, the tradeoffs it has to name, and the claims it cannot make without support. That is a harder document to write than a worked solution, and writing it is usually the first time anyone discovers what the assignment is really testing.
Can one reviewer grade everything if the key is good?
For a small round, yes, and a good key makes a single reviewer far more consistent than two without one. Add a second reader at the point where the decision is close or the round is large enough that the first and last submissions are separated by a week. Double grading buys reliability rather than accuracy, so spend it where the ordering of two candidates is actually in question.
References
- 1. Structured Interviews: A Practical Guide opm.gov Supports the definition of structure as same questions, common rating scale, and agreement in advance on what an acceptable answer looks like, which is what an answer key records.
- 2. Belief in the unstructured interview: The persistence of an illusion sjdm.org Supports the claim that reviewers construct a coherent judgment from low-signal material, including answers generated at random, which is the mechanism behind rubric drift.
- 3. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that untrained readers cannot separate model-written from human-written text, so surface polish does not belong in an answer key.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.