Assessment design

Borrow the Bar, Not the Verdict, When You Cannot Judge the Work

You can hire well for a job you have never done, because the bar and the verdict are two different purchases and only one is scarce. What good looks like in the role is borrowable: two practitioners, thirty minutes each, on what a bad hire gets wrong in month one and what a good one has shipped by month three. The verdict stays with you, and you make it judgeable by setting an exercise that asks for a decision from your own business with the reasoning shown.

The takeThe dangerous move is the newest one. Asking a model to write the interview questions and grade the answers turns not knowing the function into a number that reads like knowledge, and you cannot audit either the questions or the grade. Point the same tool the other way: use it to attack your own brief and to produce a mediocre answer you can hold the real ones against. Keep the grading.

Where Olive fits

Open a role and see what the work shows

Building this in-house, the expensive parts are the answer key and the evidence behind each finding. Olive runs a role-grounded assignment in each of twelve occupations and returns six findings, each anchored to a timestamped excerpt from the session that a first-time hirer can read for themselves.

Rank your shortlist

What can you borrow, and what can you not?

You can borrow the standard and you cannot borrow the call. Two people who do this job elsewhere will tell a stranger, in half an hour, what a bad hire gets wrong in the first month and what a good one has produced by the third. What they will not do is sit in every interview you run for the next six weeks, which is the part founder advice keeps assuming you can buy.

So make the borrowed half permanent and cheap. Ask each practitioner four things and write the answers down once:

  • What does someone in this role get wrong in their first month that a good one never does?
  • What has a strong hire actually produced by the end of month three?
  • Which judgment in this job should never be handed to a model, and why?
  • What does a portfolio piece look like when the person did not really do the thinking?

Two calls, an hour in total, and you own a standard you can reuse for every candidate on this search and the next one. It also survives your ignorance, because the answer to the first question describes something a person does, and you can watch for that without knowing the field.

What you should not do is fall back on the proxies that feel like expertise. Years of experience is the usual one, and it is close to empty: across 81 independent samples, prehire work experience correlated .06 with later job performance and .00 with turnover, and experience with relevant tasks, jobs or occupations did no better 1. Those are corrected correlations, so this is not a real signal being under-measured. Portfolios are the second fallback, and their problem is newer: the artefact and the ability behind it have come apart. Evaluating a portfolio when AI clearly made half of it is a procedure of its own.

Write the bar in three lines before you meet anyone

Three lines, written once, reused for every candidate: the worst mistake a bad hire makes early, the thing a good one has produced by month three, and the one judgment that must never be handed to a model in this role. That last line is the one your two practitioners will argue about, and the argument is the useful part.

Write them as behaviors. "Senior-level judgment" asks you to compare a candidate against a population you have never seen. "Notices when a number in the brief contradicts the summary above it" either happened in front of you or it did not, and you can see which without knowing the field.

Naming the specific capability you saw also seems to carry signal. Text-mining post-interview notes on 7,650 candidates hired at one large technology company, researchers found that the number of job-related capabilities an interviewer named in the notes tracked later performance and promotions, and ran the other way against turnover, with roughly 2 percent higher performance per standard deviation of match to the job analysis 2. Small effect, one firm, and measured only on people who were hired, so treat it as a reason to write specifically. It is not proof that note-taking predicts anything. The three lines are what your notes have to name.

Keep them visible while you interview and do not edit them mid-search. A standard that moves as candidates arrive is not a standard, it is a record of who you liked. The harder version of this, where the skill is AI use and you may have to defend the threshold afterwards, is setting a defensible bar for good enough in a specific role.

Design an exercise you can judge without the domain

Ask for a decision, not a deliverable. A deliverable needs a specialist to grade. A decision comes with reasons, tradeoffs and something the candidate chose not to do, and all three are legible to whoever runs the business the decision belongs to. Take a real call you made last quarter, strip your answer out, and hand over the information you had at the time.

Then plant something wrong in the brief. A stale figure, a claim that contradicts a table two pages later, a constraint that makes the obvious plan impossible. Noticing takes no domain expertise on your side, which is why this is the part of the exercise you can grade cold. One planted error is only one observation, so read the checking as well as the catch: a candidate who missed it and can still say which figure they would have opened has shown you the habit you are hiring for.

Relevance is doing the work here. In the job knowledge meta-analysis behind the current validity estimates, all 164 studies together averaged an observed validity of .22, while the 59 that used knowledge tests built for the job in question averaged .31, rising to .40 once corrected for unreliable performance ratings 3. That comparison sits inside one older meta-analysis, was never designed as an experiment, and presupposes candidates who already hold the knowledge. It still points where the practical advice points: build the exercise out of your own material.

Do not let it be the only evidence you have. On the corrected estimates, a mechanically weighted composite of six standardized predictors reaches about .61, and giving cognitive ability a weight of zero inside it costs .05 4. That is a modelled combination of scored instruments, so no founder reproduces the number by stacking conversations. What transfers is the direction: two kinds of evidence beat one, and a reference conversation with someone who watched this person work costs you twenty minutes. Whether to build the exercise yourself or buy an assessment is a separate call, and it turns on how often you will hire into this function.

Why shouldn't a model grade the answers?

Because a grade from a system that has never seen your business is a guess wearing a number, and you are the one person in the room who cannot audit it. The same tool earns its place pointed the other way: at your brief, hunting for what is ambiguous, and at the task itself, producing the mediocre answer you hold the real ones against.

Run it in that order on the Friday before your first interview.

1. Paste the brief in and ask what a capable person would still have to ask you before starting. Every question it raises is a hole a candidate would have hit. 2. Ask it to do the task, then read what comes back. That is your floor. Any candidate answer that reads like that one has told you something about the candidate. 3. Ask it to argue against the decision you actually made last quarter. If that argument is strong, your planted error is in the wrong place.

What you never do is hand it a candidate's answer and ask how good it is. You cannot check that grade, which is the whole reason you are reading this, and an unauditable number is worse than none because it closes the conversation you should still be having with your practitioners.

The honest limit: this gets you a competent hire rather than a great one. Great requires taste you do not have yet in this function. The fix is the hire itself. Whoever you bring in should spend part of their first month writing down the real bar, in their own words, for the second person you hire into this team.

See what gets scored

Common questions

Where do I find two practitioners who will talk to me?

Ask for thirty minutes on a narrow question rather than for advice. People who do a job will describe what a bad first month looks like in their field to almost anyone, because it costs them nothing and they enjoy it. Investors, your existing customers, former colleagues one function away, and anyone in your network who has managed this function before are all reachable. Do not ask them to review candidates or to sit in on interviews, which is a real favour with a real cost, and is the ask that usually gets declined.

What if the two practitioners disagree about what good looks like?

That is information, not a failure of the method. Disagreement usually means the role is really two roles, or that the level is unsettled, and both are things you needed to know before writing the posting. Ask each of them what the other's answer would cost you if it were wrong. Pick the version that matches the work sitting in front of your company this year, write it down, and note the other one as the thing you may need to hire for next.

Should I hire a recruiter or a fractional lead instead?

A fractional lead solves this properly if you can afford one and can find one for your function, because they bring the standard and the judgment together. A contingency recruiter does not: they source and screen against a specification you supply, so a vague specification produces a plausible shortlist you still cannot evaluate. If you use one, hand them the three-line bar rather than a job title, and keep the exercise and the decision inside the company.

How long should the exercise be?

Long enough to require a decision and short enough that a working candidate will do it. Sixty to ninety minutes is usually the ceiling for something unpaid, and a live session where you watch the reasoning happen is often better than a take-home you have no expertise to grade. Pay for anything longer. The length matters less than the shape: one real decision with the reasoning shown beats a polished artefact you cannot assess.

Can I just ask the candidate to teach me the function?

It is a good interview question and a bad evaluation method. Asking someone to explain how they would set up the function tells you how they think about scope, sequencing and what they would refuse to do first, all of which you can judge. What it cannot tell you is whether their answer is right, because the whole premise is that you have no way to check. Use the question, then check the answer against the three lines your practitioners gave you.

References

  1. 1. A meta-analysis of the criterion-related validity of prehire work experience Personnel Psychology, 72(4), 571-598 (Van Iddekinge, Arnold, Frieder and Roth); record and abstract at the University of North Florida Digital Commons, 2019. digitalcommons.unf.edu Supports the claim that years-of-experience proxies carry almost no predictive information (.06 with job performance, .00 with turnover across 81 samples), which is what a first-time hirer falls back on.
  2. 2. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company Frontiers in Psychology, Volume 11, Sec. Organizational Psychology (Shanshi Liu, Yuanzheng Chang, Jianwu Jiang, Haigang Ma and Huaikang Zhou), 2021. frontiersin.org Supports writing the bar as named job capabilities rather than levels: notes that named the capabilities the job analysis called for tracked later performance, promotions and turnover.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the claim that content relevance, not assessment format, is what moves validity: job-specific knowledge tests averaged .31 against .22 for the full set.
  4. 4. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300, doi 10.1017/iop.2023.24 (Cambridge University Press), 2023. cambridge.org Supports the claim that a hiring process should combine kinds of evidence rather than rest on one exercise: a mechanically weighted composite of six standardized predictors reaches about .61, and zero-weighting cognitive ability costs .05.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.