Assessment design

Use a Real Task You Already Solved, and Keep the Outcome

Source a hiring work sample from a real problem your team solved six to twelve months ago, strip the identifying detail, and keep the decision that was actually made as your comparison point. That gives you two things an invented scenario cannot: a task specific enough to resist an off-the-shelf answer, and a known outcome to hold submissions against without pretending one answer was right. A candidate's past project supplies neither, because you cannot establish who did what. Never use a live problem you still need solved.

The takeThe advice to invent a fictional scenario optimizes for one risk, the accusation of free work, and it pays for that with the only property that still matters. Fiction has to be generic to be readable in a page, and generic is the shape a model completes without breaking stride. A solved problem answers the free-work objection more cleanly anyway, because the brief can say the work is done and the decision is made.

Where Olive fits

Open a role and see what the work shows

Sourcing a case from real work is the part that does not scale. Olive's assignments are grounded in one occupation and its SOC code, the assistant beside the candidate is expected to produce the work, and a human reviewer writes all six findings from what the candidate did with it.

Rank your shortlist

Where should the work sample come from?

From a decision your team already made and has stopped arguing about, usually six to twelve months old. Ask the hiring manager for the last call on their team that took a week and could have gone either way. That is the raw material: real constraints, a real trade-off, and an outcome you can hold a submission against.

Three sources are on offer and they are not equivalent.

  • A real task the team already solved. Specific, defensible, and it comes with a comparison point. This is the default.
  • A candidate's past project. Good interview material. As evidence it fails on attribution: you cannot establish who was in the room, what was handed off, or what got thrown away.
  • An invented scenario. The standard recommendation, and the weakest of the three for reasons the next section covers.

Content does more work here than format. In the job knowledge meta-analysis behind the 2022 validity revision, all 164 studies together produced a mean observed validity of .22, while the 59 that used knowledge tests built for the job in question produced .31 1. What separates those two subsets is relevance rather than method.

That comparison sits between two subsets of one older meta-analysis with no controlled test underneath it, and job knowledge tests assume candidates who already have the knowledge. Do not stretch it into a law. Take the direction from it: an exercise about the work you actually do carries more than an exercise about work in general.

Why does an invented scenario fail now?

Because a fictional brief has to be generic to be readable in a page, and generic is exactly what a model completes without breaking stride. The public exercise banks have already been measured against a model. Evaluated against 115 Python problem statements taken from HackerRank, Codex solved 96% of them zero-shot in 2022, and the authors report clear signs the model was reproducing memorized code rather than synthesising it 2.

The memorization half is the part usually dropped, and it is the half that matters here. A problem sitting in a public bank was plausibly in the training data, so the exercise tests a lookup whether or not a candidate uses an assistant. That was a 2022 model on a curated benchmark, so treat the figure as a floor under today's capability.

Neither study measured a brief invented for one role, and a brief written last week sits in no training set. The overlap is shape. A self-contained problem statement with everything it needs inside it is the easiest thing an assistant is ever handed, whether or not it has seen that exact one.

Raising the difficulty buys a shrinking margin rather than a wall. Submitting model-written solutions to LeetCode, one study recorded 92% of easy, 79% of medium and 51% of hard problems solved by a 2023-era model, with further gains from feeding failed test cases back in a second prompt 3. Those are public algorithmic problems and that model is superseded, so the useful part is the gradient.

The reporting line that complicates the obvious fix, the metric two teams define differently, the contract that rules out the clean option: those are what an off-the-shelf answer cannot reach, and a harder generic problem still has none of them. That is also the material a work sample has to test now that AI can produce the work sample.

Turn a closed decision into a brief

Start from the decision, not the deliverable. Write down what was decided, which two options were live, what evidence existed at the time, and what was missing. Then wind the exercise back to the moment before the call, hand the candidate the same incomplete evidence the team had, and ask for the decision plus the reasoning that got there.

The steps, in order:

1. Pick the decision. Six to twelve months old, closed, and one you would still defend. Anything younger is often still contested; anything older has usually stopped resembling the job. 2. Strip the identifiers. Names, customers, figures that matter. Rebuild the numbers so they carry the same shape without the real values, which is the same discipline as deciding whether real company data belongs in the exercise at all. 3. Cut the context to one page. If the brief needs three pages of background, the exercise has become a reading comprehension test. 4. Write down the actual outcome and keep it back. Hold it as a comparison point. Grading against it is the failure mode, because a candidate who reaches a different conclusion well is a better signal than one who guesses yours. 5. Name what a strong answer engages with. Usually two or three constraints. A submission that ignores all of them has told you something.

Then maintain it. Re-read the brief each quarter, and retire it when the original answer is no longer one you would argue for, or when the same three submissions start arriving. Sourcing from real work makes an exercise perishable, which is a feature, and the answer key you write before the assignment goes out is what tells you the ground has moved.

When does a past project or a live problem make sense?

A past project belongs in the interview and never in the work sample, and a live problem you still need solved belongs in neither. Nobody can reconstruct who did which part of it, and no amount of reading the artifact recovers that. Use it in conversation: ask what was rejected, what got checked against something outside the room, what would be done differently now. Those questions survive an assistant sitting beside the candidate.

The live problem is where the free-work objection is real. That is the line between an exercise and unpaid consulting, and it does not move because the brief calls the work hypothetical. The solved-problem source disposes of the objection outright: the brief can state that the decision was made months ago and the exercise exists to watch somebody reason, which is both true and checkable.

Two edge cases deserve naming. Regulated or safety-critical work sometimes cannot be reconstructed at all, in which case a live session with a person in the room does the job a take-home cannot. And in high-volume early-career hiring a bespoke case per role is not affordable, so the realistic move is one carefully built case per job family, refreshed on a schedule, in place of an off-the-shelf test that a model already solves in thirty seconds.

One test before anything goes out. Could a competent stranger with an assistant produce a credible submission while knowing nothing about your business? If yes, the exercise is measuring the assistant. If no, it is measuring the part of the job you were hiring for, and the free-work objection never arrives, because the work was finished before the candidate ever saw it.

See what gets scored

Common questions

Can I use a problem the team is working on right now?

No. A live problem you still need solved is unpaid consulting whatever the brief calls it, and the incentive it creates is bad in both directions: the team starts hoping for a usable answer and the reviewer starts scoring usefulness rather than reasoning. Wait until the decision is closed. If a current problem is genuinely the only material available, take it into a live session where the conversation is plainly the point, or pay for it on a contract. Whether an unpaid exercise ever becomes compensable work is a wage-law question that varies by jurisdiction, so put a long one to counsel before it goes out.

How old should the source problem be?

Six to twelve months is the usual window. Younger than that and the team is often still arguing, which means reviewers score submissions against their own side of the argument rather than against the rubric. Older than about eighteen months and the constraints tend to have changed enough that a strong candidate's answer looks wrong for reasons that have nothing to do with them. Re-read the brief each quarter and retire it when the original decision stops being one you would defend.

Do I have to tell candidates the problem is real?

Say that it is drawn from real work, that the decision has already been made, and that the material has been altered. That single sentence does most of the work the fictional framing was invented to do: it explains why the task is specific, it removes any suggestion the business is harvesting free labour, and it sets the expectation that there is a defensible answer without promising a correct one. Do not name customers, and do not imply the candidate's answer will be used.

What if the team cannot agree on what the right answer was?

Then it is a strong source and a weak rubric, and the rubric is the part to fix. Write down the two positions and score submissions on whether they engage with the constraint the disagreement turns on. A live disagreement inside the team also predicts a live disagreement in the debrief, so settle it before the first candidate sees the brief. If it cannot be settled, pick a different decision instead of sending an exercise your reviewers will score two different ways.

How many exercises does one role need?

One well-built case per job family, refreshed on a schedule, is enough for most teams. Multiple instances of that case are useful, but they should be instances rather than different tasks: the same question against different data. Building a second exercise is worthwhile only when a second job family genuinely asks for different judgment. A team with more exercises than reviewers who can calibrate them has bought variety at the cost of ever knowing what any of it measured.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the claim that relevance carries more than format: job-specific knowledge tests produced .31 against .22 across all 164 studies in the meta-analysis.
  2. 2. Codex Hacks HackerRank: Memorization Issues and a Framework for Code Synthesis Evaluation arXiv:2212.02684 (Karmakar, Prenner, D'Ambros and Robbes), 2022. arxiv.org Supports the claim that an exercise built from a public problem bank tests a lookup: 96% of 115 HackerRank problems solved zero-shot, with clear signs of memorized code.
  3. 3. Evaluating ChatGPT-3.5 Efficiency in Solving Coding Problems of Different Complexity Levels: An Empirical Analysis arXiv:2411.07529 (Li and Krishnamachari), 2024. arxiv.org Supports the claim that raising difficulty on a public-style exercise buys a shrinking margin: 92% easy, 79% medium and 51% hard solved by a 2023-era model.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.