Assessment design
Do Not Require a Skill Nobody Can Score: Borrow the Evaluator First
Don't write a requirement for a skill nobody on your team can judge; borrow the evaluator first. Find one practitioner who does the work well, from inside the company, an advisor, a contractor or a peer elsewhere, and have them write three things: what a strong answer to a real task looks like, what a plausible but weak one looks like, and the two questions that separate them. That artifact is the requirement. The line in the posting is only its short form.
The takeTwo industries are built around this gap and neither one closes it. Recruiting content answers a different question, how to word a requirement so it reads well, and assessment vendors sell a test whose key the buyer also cannot read. Both leave you holding a bar you cannot apply. The cheapest fix is embarrassingly old: one paid afternoon of a practitioner's time, in exchange for a page you can hand to whoever runs the interview.
Where Olive fits
Open a role and see what the work shows
The two hard parts of building this in-house are the answer key and a person who can read it. Olive ships authored cases grounded in one occupation and returns six findings written by a human reviewer, each quoting the moment in the session it rests on.
Rank your shortlistGet the evaluator before you write the requirement
Start with the person, not the wording. One practitioner who does this work well, for a paid afternoon, will tell you more about what the role needs than a month of reading other companies' job descriptions. Sources, roughly in order of how fast they answer: an investor's network, a former colleague, a contractor you have paid before, and a peer at a company one size up.
Be specific about what you are buying: a description of the work, and a way to tell good from plausible. A referral, a read on the market, or an opinion about whether a candidate seems smart will not get you there. Two hours is enough if you arrive with a real task and your questions already written.
Relevance is what you are paying for, and a measured gap sits behind it. In the job knowledge literature the 2022 re-analysis relies on, all 164 studies together produced a mean observed validity of .22, while the subset using knowledge tests built for the job in question produced .31, rising to .40 once corrected for unreliable performance ratings 1. Those are two subsets of one meta-analysis with no controlled experiment behind them, and job knowledge tests assume candidates who already have the knowledge, which makes this evidence about hiring experienced people. The direction still matters for a buyer with no expertise: what an assessment is about does more work than what format it takes, and only somebody who knows the work can set that.
One rule while you shop. If a vendor cannot tell you who wrote their key and what a wrong answer looks like on it, you have paid somebody else to have your problem. That is the substance of the build-or-buy decision for an AI exercise, which applies to any domain you cannot read.
What a borrowed evaluator has to hand you
Three artifacts, and one page holds all of them. First, a real task from the next quarter of the job, written the way a candidate would receive it. Second, a strong answer and a plausible-but-weak answer described side by side, in enough detail that a non-practitioner can tell them apart. Third, the two follow-up questions that separate the two.
The second artifact is the one people skip and the one that does the work. A description of a strong answer on its own is unusable, because almost everything looks strong to a reader who cannot see what is missing. The pair makes the difference legible: one answer names the constraint and checks its number against a second source, the other asserts both.
Then turn the pair into a written scale, and decide in advance how the marks get added up. A meta-analysis of selection and admissions studies found applicant data combined by a stated rule correlated .44 with job performance, against .28 when the same kinds of data were combined by expert judgment, which the authors describe as a population-level improvement in prediction of more than 50% 2. Read that as a relative gain in a correlation: the job-performance comparison rests on nine studies and the gap is smaller on other criteria, so do not oversell it. A mechanical rule here can be as plain as adding unit-weighted marks on a written scorecard, which is exactly what a borrowed evaluator page makes possible.
What the evaluator should not hand you is a verdict. Having the expert sit in and give a gut read at the end is the one move the evidence argues against: adding an intuitive judgment on top of two mechanical predictors moved the correlation with the criterion from .45 down to .35, and the review reporting it notes that the idea of a naturally better interviewer is not supported by the evidence on variance in interviewer validity 3. That study drew on 1943 university admissions data and was never about hiring, so do not overclaim it. The division of labour still holds: experts collect and define, a written rule combines. Getting non-practitioner managers to apply that rule consistently is its own problem, worked through in how to get managers who do not use AI to judge AI-assisted work.
Why an unjudgeable requirement is also an undefendable one
Because the standard a selection procedure has to meet is stated in terms of the job, and nobody in the room can say what the job requires. Under Title VII in the United States, once a complaining party shows a particular practice causes disparate impact, the employer has to demonstrate it is job related for the position in question and consistent with business necessity 4. That demonstration starts from a job analysis, and a job analysis is what is missing.
The principle is older than the statute. In Griggs v. Duke Power Co., decided in 1971, a unanimous Supreme Court held that Title VII reaches practices that are fair in form but discriminatory in operation, that the touchstone is business necessity, and that nothing forbids testing, only giving a test controlling force unless it is demonstrably a reasonable measure of job performance 5. The Court's summary line is the one to keep: a test must measure the person for the job and not the person in the abstract.
Calling the stage informal does not move it outside any of this. The Uniform Guidelines define a selection procedure as any measure, combination of measures, or procedure used as a basis for any employment decision, and say so explicitly for the full range from paper tests through informal or casual interviews and unscored application forms 6. The same Guidelines say a validity study rests on a review of information about the job and should include a job analysis, which for a content-valid procedure has to identify the important parts of the work required for successful performance and their relative importance 6. So a founder who decides on impression in a conversation has not avoided running a test. They have run one with no record of how it was applied and nothing written down about the job it was for.
What follows is public legal fact rather than advice: the Uniform Guidelines date from 1978, the burden-shifting text is 42 U.S.C. 2000e-2(k) as added in 1991, and Griggs is 1971. Title VII covers race, color, religion, sex and national origin; age and disability claims run under separate statutes. Whether a requirement is defensible for a particular role belongs in front of counsel. The operational point needs no lawyer: the artifact that makes a requirement judgeable is the one that makes it defensible.
The AI-era version of this runs in one direction only. Plausible expert-looking work is now cheap to produce in a domain nobody on your side can check, which converts an evaluation gap from a slow hire into a wrong one. Nobody in the room can tell that the architecture is conventional, the model misspecified, or the filing headed for rejection. That failure surfaces in month four, long after the interview.
Run the task when no evaluator exists at any price
Write the task instead of the requirement, pay for it, and judge the work against the outcome it has to produce. If nobody reachable knows what good looks like, the only honest evaluator left is reality: a scoped, paid piece of real work with a stated outcome, a deadline and a definition of done a non-practitioner can check without an opinion about craft.
Pick an outcome with an external verdict attached. The migration ran and the numbers reconciled. The filing was accepted. The integration passed the partner's test suite. Each is checkable by somebody who could not have produced the work, which is the point of choosing it.
Be clear about what this does not buy. It is a sample of one on a task you scoped, so it says little about how the person handles a task you did not anticipate. It costs money, and it should: unpaid multi-hour work is a different conversation and a worse one. And it tells you only whether each candidate cleared the outcome, which is a narrow basis for comparing two of them. That is a real limitation, and still better than a bar nobody can apply.
One last pass before the requisition opens. Name the evaluator for every line on it, with a person's name. Any line with no name is a preference, and the posting should say so. The one- or two-person version of this is worked through in making the first hire into a role you cannot judge, and the answer is the same: borrow the reader, or post the task.
Common questions
Who can I ask to be an evaluator if I have no network in that field?
Paid experts are easier to reach than most first-time hirers expect. A contractor from a previous project, a consultant billing by the hour, a professional-association member, or a practitioner two levels senior at a company that is not a competitor will usually take a two-hour engagement. Say exactly what you want produced: one task, a strong and a weak answer described side by side, and two follow-up questions. A defined deliverable gets a yes far more often than a request to help with hiring.
Can I use an AI model to write the requirement and the rubric instead?
It will produce something plausible, which is precisely the failure mode you are trying to avoid. A model can draft the task and the two answer descriptions quickly, and that is a reasonable starting point, but nobody on your side can tell whether the key is right, so you have moved the unjudgeable artifact one step upstream. Use it to prepare a draft that a practitioner marks up in twenty minutes instead of writing from scratch. The human check is the part that cannot be skipped.
What should the posting line say if the rubric is a whole page?
Write the task, not the skill. A line reading you will own the monthly close and the first thing you will do is rebuild the reconciliation tells a candidate what the work is and lets a qualified reader self-identify. Skill words like expert-level and deep experience mean whatever the reader wants them to mean, and they attract applications tuned to the wording of the posting. The page-long rubric stays internal, where it does its work in the interview and the debrief.
Is it fair to run a paid trial task instead of an interview?
It is fair when it is scoped, paid at a real rate, and optional in the sense that a candidate knows what they are agreeing to before they start. It stops being fair when it is long, unpaid, or delivers production value that you would otherwise have bought. Say the hours, the fee and the decision it feeds before anybody agrees. And run it as the last step rather than the first, because asking several finalists for a day of work is a much smaller ask than asking twenty applicants.
What if the evaluator and my instinct disagree about a candidate?
Write down which one of you can point at evidence in the work, and go with that. The pattern the research warns about is layering a holistic impression on top of a structured signal, which can lower accuracy rather than raise it. That does not mean an instinct is worthless: it usually means something was observed that the rubric did not cover. Say what it was out loud, and if it turns out to be a real requirement, write it into the rubric before the next candidate, and decide this one on the evidence the rubric already covers.
References
- 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range static1.squarespace.com Supports the job knowledge comparison: 164 studies at a mean observed validity of .22 against the job-specific subset at .31, rising to .40 once corrected for unreliable performance ratings.
- 2. Mechanical Versus Clinical Data Combination in Selection and Admissions Decisions: A Meta-Analysis gwern.net Supports the .44 against .28 comparison between mechanical and holistic combination of applicant data and the authors described population-level improvement of more than 50%.
- 3. Stubborn Reliance on Intuition and Subjectivity in Employee Selection edbatista.com Supports the drop from .45 to .35 when intuitive judgment was added to two mechanical predictors in Sarbin's 1943 admissions data, and the review's statement that variance in interviewer validity is attributed to sampling error.
- 4. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases uscode.house.gov Supports the statutory burden-shifting: once a complaining party demonstrates that a particular practice causes disparate impact, the employer must demonstrate the practice is job related for the position in question and consistent with business necessity.
- 5. Griggs v. Duke Power Co., 401 U.S. 424 (1971) law.cornell.edu Supports the holding that Title VII reaches practices fair in form but discriminatory in operation, that the touchstone is business necessity, and that a test must measure the person for the job rather than in the abstract.
- 6. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q), 1607.3(A) and 1607.14(A) and (C)(2) govinfo.gov Supports the definition of a selection procedure as any measure or procedure used as a basis for an employment decision, explicitly covering informal or casual interviews and unscored application forms, and the requirement that a validity study rest on a review of information about the job and include a job analysis.
6 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.