Assessment design

Write the Assignment Brief So Two Submissions Can Be Compared

A take-home brief that makes two submissions comparable names four things most templates leave out: the time budget, who the output is for, what is explicitly out of scope, and one line on whether AI is allowed. Add the words your rubric uses, so nobody has to guess at the bar. Then read it back. If a candidate would need to email a clarifying question before starting, the submissions are already incomparable, and the variance you are about to grade is yours.

The takeMost scoring disagreements are brief failures wearing a rubric's clothes. Two reviewers arguing about whether a submission went deep enough are usually arguing about a depth the brief never specified, and whichever candidate happened to guess the reviewer's preference wins a coin toss nobody agreed to run. Fix the brief and most of the calibration meeting disappears. A brief that survives four candidates without producing a single clarifying email is worth more than a fifth reviewer.

Where Olive fits

Open a role and see what the work shows

Olive runs one authored assignment per role rather than a prompt each candidate scopes for themselves: 40 to 60 minutes of work with the brief, the materials and the deliverable fixed in advance. What comes back is six findings, each carrying the timestamped excerpt a human reviewer wrote it from.

Rank your shortlist

What four things does the brief have to name?

Time budget, audience, out of scope, and the tool rule. The prompt says what to make. Those four say how long, for whom, how far, and with what. Templates ship the prompt and stop there, so every constraint left open is one each candidate closes privately, which is the moment two submissions stop being the same question.

  • The time budget, in hours, stated as a cap rather than a suggestion. Two hours means stop at two hours, and the brief should say what to do with the part that is unfinished: leave a note about what you would do next. Without that sentence, the candidate with a free Saturday hands in twice the work and the rubric reads it as twice the ability.
  • Who the output is for. A memo for the head of finance and a memo for a fellow analyst are different documents. Name the reader, their seniority, and how much context they already have.
  • What is out of scope, by name. The cleanest version is a short list: no visual design, no production-ready code, no external research beyond the attached files. Out-of-scope lines save more candidate hours than any other sentence in the brief.
  • The tool rule, in one line, pointing at the policy you already have. Whatever your position on AI in the take-home turns out to be, the brief has to state it plainly, because silence is not a rule and candidates will resolve it in opposite directions. If AI is allowed, the same line has to answer whose AI account the candidate uses.

One addition earns its space when the assignment involves data or a codebase: say what the candidate may assume is correct. Otherwise somebody emails to ask whether a broken field in the sample data is the puzzle or an accident.

Why do identical prompts produce incomparable work?

Because a prompt is not a specification. A candidate reading "analyse this dataset and write up what you find" has to decide the depth, the reader and the hours before writing a word, and four candidates decide differently. What comes back then varies as much by those private decisions as by anything you meant to measure, and a rubric applied afterwards cannot separate the two.

The same variable separates a good interview from a bad one. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42 against .19 for unstructured ones, pooled from two earlier meta-analyses of interview validity 1. Those are corrected correlations with supervisor performance ratings, pooled across many jobs, not accuracy rates, and none of the underlying studies is about a take-home. What travels is the mechanism rather than the number: the difference between two interviews of the same method is larger than the difference between two methods, and structure is the thing that difference is made of.

A take-home is not an informal step outside the rules. The Uniform Guidelines, the 1978 US federal rules on selection, define a selection procedure as any measure or procedure used as a basis for an employment decision, covering the full range from paper-and-pencil tests through performance tests to informal interviews and unscored application forms 2. A written assignment sits squarely in that definition. The Guidelines impose a validation burden only where adverse impact appears, so this is not an audit mandate, but it does settle the category question: the brief is test documentation, not a friendly email.

The practical tell arrives at grading. When every reviewer's notes start with a sentence explaining what that candidate seems to have assumed, you are reading four specifications, and the polish on the surface will not tell you which assumptions were reasonable.

Write the rubric's words into the brief

State what you will read the work for, in the same words the rubric uses. If a rubric line says "checks the load-bearing claim against something outside the prompt," the brief says that too. Candidates stop guessing at a hidden bar, reviewers read for something they announced in advance, and the exercise stops quietly rewarding whoever guessed your taste.

The objection is that publishing the criteria lets candidates perform to them. Take that seriously and then look at what performing to them costs: a candidate who reads "we look for a stated assumption behind every estimate" and then states their assumptions has done the thing you wanted. Criteria you can game by doing the work are criteria worth publishing. The ones you cannot publish are usually the ones that were never criteria, just preferences.

What stays private is narrower than people expect. Keep the weights, keep the reference solution, and keep the specific errors you are watching for. Publish the dimensions and the words. If writing the key reveals that a dimension cannot be described without giving away the answer, that is a finding about the assignment rather than a reason to hide the rubric.

One format note. Put the criteria at the end of the brief, under a heading that says what they are, and cap them at five lines. A brief that opens with a rubric reads as an exam paper and lowers the quality of everything after it.

Test the brief before it goes out

Hand it to two colleagues who have not seen it, and ask each what they would produce and how long they would spend. Two different answers means the brief is still open somewhere. Then do the assignment yourself inside the stated budget. A brief nobody on your side can finish in the time you gave is not a brief, it is a request for unpaid overtime with a deadline attached.

Three cheap checks, in the order they catch things:

1. The clarifying-email test. Read every question the last cycle produced. Each one is a sentence the brief should have contained, and adding it costs nothing. If you are running the assignment for the first time, ask your two readers to write down the questions they would send. 2. The out-of-scope test. For each thing a strong candidate might reasonably do, decide whether you want it. Anything you do not want goes in the out-of-scope list, or someone will spend a third of their budget on it. 3. The stop test. Ask both readers what they would hand in if they ran out of time at the cap. If neither can answer, the brief has no graceful exit and it is going to punish the candidates who respected the clock.

Freeze the brief once the first candidate has it. If you must change it mid-cycle, send the change to everyone who already received it, extend their deadline, and note in the file which version each submission answered. An unrecorded edit halfway through a round makes the early and late submissions two different tests, which is the failure this whole exercise exists to prevent. The same discipline applies when an assignment gets redesigned between cycles: the old and new submissions do not belong in one comparison, however similar the prompts look.

See what gets scored

Common questions

How long should a take-home assignment be?

Under three hours for most roles, and the honest way to set it is to do the assignment yourself and double what it took. The number matters less than the fact that it is stated and enforced, because an unstated budget is what turns a two-hour exercise into an eight-hour one for whoever has the weekend free. If the work genuinely needs a day, that is a paid project rather than a screening step.

Should the brief say how the submission will be reviewed?

Yes: how many people read it, roughly when the decision lands, and whether there is a follow-up conversation about the work. A candidate deciding whether to spend three hours is making a scheduling decision, and vagueness there costs completions among exactly the people with the least slack. It also removes the most common cause of a chasing email, which is silence rather than delay.

What if a candidate asks a clarifying question anyway?

Answer it, then send the same answer to every other candidate in the round and add it to the brief. Answering privately hands one person a better specification than the rest, which is the incomparability you were trying to remove, arriving through the front door. Log the question either way, because it is free evidence about which sentence in the brief is not doing its job.

Can one brief work across seniority levels?

The task can, the constraints cannot. Keep the same prompt and dataset, then change the time budget and the audience: a junior writes a recommendation for their manager, a senior writes the same recommendation for a director and names the tradeoff they would defend. Grading against one rubric with different expectations per level is far more workable than running two unrelated assignments and pretending the results are on one scale.

Does a longer brief hurt completion rates?

Length is not the problem, ambiguity is. A one-page brief with a clear cap and an explicit out-of-scope list reads faster than a four-paragraph prompt that leaves a candidate calculating how much work is enough. Keep it to a page, use headings, and put the deadline and the time budget where they cannot be missed. Everything you cut from the brief reappears as email.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 and .19 corrected validity estimates for structured and unstructured interviews behind the claim that structure, not method, is what makes two assessments comparable.
  2. 2. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that a written take-home assignment is a selection procedure under federal selection law, in the same category as a test or an interview.

2 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.