Assessment design
Every Requirement Needs a Stage That Tests It
Put the requirement list in two columns before the posting goes out: the requirement on the left, the stage that will produce evidence for it on the right. Any line with an empty right column either gets a stage or comes off the list. A requirement becomes testable when you can name what a person would do or show that satisfies it, which usually means rewriting an attribute into a behaviour or an artifact.
The takeMost requirement lists are twice as long as they should be, and the cure is subtraction rather than another round. Four requirements with real evidence behind them beat twelve with none, because twelve produce a debrief where nobody can say which line the candidate failed. The list is also the one part of a job description that gets reread, which makes it the cheapest place in the whole process to prevent a bad hire.
Where Olive fits
Open a role and see what the work shows
If you build the evidence side of this in-house, the expensive parts are the answer key and the trail behind each finding. Olive runs a role-grounded assignment in one of twelve occupations and returns six findings, each anchored to a timestamped excerpt from the session rather than to a level.
Rank your shortlistWhat makes a requirement testable?
A requirement is testable when the evidence for it has a name and a place: what a person would do or show that satisfies it, and the stage where they will do or show it. Attributes fail that test. *Strong communicator* names nothing observable, so two interviewers can rate it opposite ways and both be right. *Explains a technical decision to somebody who disagrees, in round two* gives the interviewer something to listen for.
Take the attribute, ask what the person is actually going to do with it in the first quarter, and write that down instead.
- Strong communicator becomes *explains a decision to a colleague who disagrees with it*, tested in round two by handing over a real disagreement from last quarter.
- Detail-oriented becomes *finds the error in a document that looks finished*, tested with a short exercise built from a document that had an error in it.
- Self-starter becomes *picks what to work on when the brief is ambiguous, and says why*, tested in the screen with a question about one specific week.
- Five or more years of experience becomes whatever you thought the five years would buy, because the years themselves carry very little. A meta-analysis of 81 independent samples puts the corrected correlation between prehire work experience and later job performance at .06, and experience with relevant tasks, jobs or occupations does no better, at roughly .07 1.
The experience line is the one people fight hardest to keep, so read the finding carefully. It covers experience measured at hire, which is exactly what a resume screen reads, and it says nothing about licensure or a legally required minimum. Within that scope, the crude count you are screening on is not carrying the signal you think it is.
Relevance is what carries signal. In the job knowledge meta-analysis behind the 2022 revision of the selection literature, all 164 studies together averaged an observed validity of .22, while the 59 studies using knowledge tests built for the job in question averaged .31, rising to .40 once corrected for unreliable performance ratings 2. That is a comparison between subsets rather than a controlled trial, and job knowledge tests assume candidates who already hold the knowledge. Within that one meta-analysis, the job-specific subset predicted better than the rest, which is the reason to write a requirement around the work itself.
Write the list as two columns before the posting goes out
Two columns, one page, before anyone opens the posting template: the requirement on the left, the stage that will produce evidence for it on the right. Fill the right column out loud with whoever is going to run that stage, because a stage nobody has agreed to run is not a stage. Then read down the right column and count how many times each name appears.
| Requirement | Stage that produces evidence | What the evidence looks like |
|---|---|---|
| Reconciles a forecast against source data | Work sample, 45 min | The candidate finds the two figures that do not tie and says which one they trust |
| Explains a decision to someone who disagrees | Round two, hiring manager | Holds or changes position, and names what would change it |
| Writes a brief a client can act on | Screen, one question | Describes the last brief they wrote and who acted on it |
| Works with an assistant without shipping its errors | Work sample, same session | Names something the assistant produced that they threw out, and why |
Three patterns show up the first time anyone does this. One stage carries eight requirements and the rest carry none, which means the loop has one real stage and three social ones. Two requirements map to the same stage and contradict each other, usually speed against thoroughness, and the interviewer has been silently choosing between them for months. And several lines map to nothing at all, which is the point of the exercise.
The right column also settles arguments the left column cannot. Whether an AI-skills line belongs in the posting stops being a question of principle once you have to name the stage that will check it, which is also most of the work in writing an AI-skills requirement that isn't legally vague. A requirement with a stage behind it is a commitment. A requirement without one is a preference wearing a requirement's clothes.
Why does an untested requirement cost more than it used to?
Because it comes back to you confirmed. A requirement nobody checked used to be decoration: it sat in the posting, nobody asked about it, and the loop carried on without it. Now the posting is the brief that applications get assembled against, so every line in it is answered on the way in, and the process has no way to tell a confirmation apart from evidence.
The gap between what a posting says and what a process does has been measured directly, in the one case where an employer's stated requirement changed en masse. Matching 11,332 roles at large firms that dropped a degree requirement against the career histories of roughly 65 million US workers, the Burning Glass Institute and Harvard Business School found those firms raised the share of hires without a bachelor's degree into those roles by about 3.5 percentage points. Because only 3.6% of roles dropped a requirement, the net effect across the labour market was 0.14 percentage points, which the authors put at roughly 97,000 workers a year and fewer than 1 in 700 hires 3.
The interesting half is who moved. Nearly all the real change came from 37% of firms, which raised their share of hires without a degree in the analysed roles by nearly 20% in relative terms; about 45% changed the posting with no meaningful difference in hiring behaviour, and about a fifth made gains that did not stick 4. Education there is inferred from career histories rather than verified, and the population is only firms posting more than 500 job ads a year, so these are estimates about large employers. Within that population, editing requirement text changed hiring only where the rest of the process changed with it, which is the argument for dropping the degree line only in the same change that replaces it.
The right-hand column exists for exactly that reason. Without a stage behind it, the edit stops at the document, and the document is the surface applications are now written toward, which is also why keyword screening on that document stopped working.
Cut the list until every line has a stage
Delete every line whose right column is empty, then check whether the stages you named can carry what you assigned them. Most loops have one stage doing nearly all the filtering and several that mostly ratify it, so a requirement parked in a late round is a requirement almost nobody is testing. Either move it earlier or cut it.
Across more than 54 million applications and 93,000 jobs in one ATS vendor's customer base, recruiter screens passed roughly 35% of the candidates who reached them, while post-onsite stages converted at 95% and offer stages at 81% 5. Those are passthrough rates among candidates who got that far, not end-to-end rates, and the customer base skews to venture-backed technology companies, so treat the levels as illustrative. The pattern is the useful part: by the time a candidate reaches the last stage, the decision is effectively already made, so a requirement you assigned to the final round gets ratified on the way past.
On Monday, take the open req and do four things in this order.
1. List the requirements as they are actually written, including the ones copied from the last version of this role. 2. Beside each, write the stage and the name of the person who will run it. A round number is not an owner. 3. Strike every line with a blank beside it. Read the struck lines aloud to the hiring manager; the ones they fight for are the real requirements, and now they need a stage. 4. Count what is left. If more than five survive, go back to the stages: a loop that cannot cover five requirements will cover four and guess at the rest.
What this does not settle is where the bar sits on the lines that survive, which is a separate and harder piece of work covered in setting a defensible bar for good enough in a specific role. Two columns tell you which requirements the process can speak to. They do not tell you how much is enough. But a bar on a requirement nobody tests is not a bar at all, so the columns come first.
Common questions
How many requirements should a job description have?
As many as the loop can produce evidence for, which in most processes is four or five. Do not pick the number first. Count the stages, count how much each one can realistically cover in the time it has, and that is your ceiling. A twelve-line list on a three-stage loop is not a stricter standard than a five-line list. It is the same five requirements plus seven lines that will be judged on whatever the interviewer happened to notice.
What about requirements that are legally or contractually required?
Those stay, and they get a stage too, usually verification rather than assessment. A licence, a clearance, a certification a client contract names, a statutory minimum: the stage is a document check at a defined point in the process, run identically for everyone. Write it in the right-hand column like any other line so nobody discovers at the offer stage that it was never confirmed. Check anything touching a statutory or contractual minimum with counsel.
Can one stage test more than one requirement?
Yes, and a well-built work sample usually tests three or four at once. The limit is attention: an interviewer tracking six things tracks none of them well, and the scorecard turns into a general impression with six boxes. Two or three requirements per stage is a workable load. If a stage is carrying more than that, either the stage is longer than anyone admits or the requirements are being scored from the same overall feeling under different names.
The hiring manager insists on a requirement I cannot test. What then?
Make them name the evidence that would satisfy them, and write down whatever they say. Usually the answer names a stage nobody had thought of, such as a reference call with a specific question, or a fifteen-minute conversation with the person currently doing the work. Sometimes the honest answer is that they would accept nothing, which means the line is a preference. Preferences are allowed to exist. They are not allowed to sit in a list the process is being judged against.
Does this apply to a role I have never hired for before?
It applies harder, and it will be wrong the first time. For a first-of-kind role, expect the requirement list to survive contact with about five screens before you find out which line was imaginary. Build the two columns anyway, run five candidates, then edit the list against what those five actually showed you. The point of writing it down is not that the first version is right. It is that the second version can be different on purpose.
References
- 1. A meta-analysis of the criterion-related validity of prehire work experience digitalcommons.unf.edu Supports the claim that a years-of-experience requirement carries almost no information about later job performance, at .06 corrected across 81 samples and .07 for task-relevant experience.
- 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range static1.squarespace.com Supports the claim that job-specific content raises what an assessment predicts: .22 across all 164 job knowledge studies against .31 for the 59 job-specific ones, .40 corrected.
- 3. Skills-Based Hiring: The Long Road from Pronouncements to Practice static1.squarespace.com Supports the claim that changing requirement text in a posting moved actual hiring by 0.14 percentage points net, or fewer than 1 in 700 hires.
- 4. Skills-Based Hiring: The Long Road from Pronouncements to Practice static1.squarespace.com Supports the split between the 37% of firms whose hiring changed and the 45% that changed the posting text only, which is the evidence that a requirement moves hiring only when a stage moves with it.
- 5. Recruiting Operations Benchmarks | 2026 Talent Trends Report ashbyhq.com Supports the claim that early stages do the filtering (about 35% passthrough at recruiter screen) while late stages ratify (95% post-onsite, 81% at offer).
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.