Assessment design

Ask for the Instance First, the Hypothetical Only Where There Is None

Ask for the instance first: a behavioral question carries the screen and the middle rounds, because what the candidate actually did holds up through three rungs of follow-up and a rehearsed answer runs out. Save the hypothetical for the final round, where no comparable experience exists yet, such as a step up in scope, and read it as evidence about reasoning rather than track record. A technical question earns a slot in the craft round only when the bar for a passing answer is written down beforehand.

The takeThe comparison is usually settled by an argument that no longer holds. Every guide prefers behavioral questions because a candidate cannot easily fabricate a past experience, and keeps hypotheticals for showing how someone thinks. Both halves moved: a plausible past is cheap to produce, and a hypothetical with a known good answer is the most retrievable question in the set. The recommendation survived while its stated reason did not. What decides now is how much of the first answer you let count.

Where Olive fits

Open a role and see what the work shows

If you are building the middle round yourself, the expensive parts are the pass bar and the evidence trail under each judgment. Olive ships authored cases grounded in one occupation across twelve item banks and returns six separately-evidenced findings, each anchored to a moment in the session rather than to a number.

Rank your shortlist

Which question type belongs in which round?

One instance question in the screen, instances probed against written anchors through the middle rounds, and a technical question only where a written pass bar already exists. The type is a smaller decision than it looks. What a round can verify is set by how much time it has and whether an anchor exists to judge against.

A working allocation for a four-stage loop:

  • Recruiter screen. One instance question tied to the requirement most likely to disqualify, plus logistics. Thirty minutes buys one ladder, not four topics.
  • Hiring manager round. Two or three instance questions on the competencies that carry the role, each with planned follow-ups and written anchors. This is where the loop earns most of what it will learn, and what the manager covers that a recruiter cannot is the other half of what the round is for.
  • Craft round. The work itself, or the closest thing to it, with the pass bar written before anyone sits down. A technical question that is really a quiz belongs here or nowhere.
  • Final round. Scope and judgment, which is where a hypothetical is legitimate, because the candidate genuinely has not done the bigger version of the job yet.

The thing to resist is spreading all three types thinly across every round so each panelist gets a taste of everything. That produces four opening answers per competency and no depth in any of them, and it is the most common way a five-stage loop returns less evidence than a two-stage one.

What does each type actually predict?

Less than the comparison posts imply, and by margins narrower than a ranking suggests. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42, job knowledge tests at .40, work samples at .33 and unstructured interviews at .19 1, while situational judgment tests sit at .26 whether they ask what a candidate should do or what they would do 2. These are corrected correlations with supervisor ratings of job performance. None of them is an accuracy rate.

Two cautions travel with those numbers, and both cut against using them as a league table. The first is that the ranking is assembled across separate literatures. The paper itself cites a head-to-head test in which cognitive ability tests beat assessment centres across separate meta-analyses (.51 against .37), and the result reversed when the same people took both and were rated on the same criterion (.44 for assessment centres against .22 for ability) 2. Comparing coefficients drawn from different literatures is exactly what the famous ranking tables do. The second is that these are corrected numbers. Before any correction, the mean observed correlation between structured interview scores and rated job performance was .28 in one of the two meta-analyses behind the estimate and .36 in the other 2. The corrections are legitimate, since unreliable performance ratings genuinely drag a raw correlation down. The two kinds of number are still not interchangeable, and an employer checking its own interview scores against its own performance ratings should expect the observed size.

Content matters more here than format. In the job knowledge meta-analysis behind that table, all 164 studies together produced a mean observed validity of .22, while the 59 studies using knowledge tests built for the job in question produced .31, rising to .40 once corrected for unreliable performance ratings 2. That is a comparison between subsets rather than a controlled test, and job knowledge tests presuppose candidates who already have the knowledge, so it is evidence about hiring experienced people. Taken carefully, it still points the same way: a question about the work this role does is worth more than a well-formed question about work in general.

So the type is not where the decision is. What makes an interview question a good one turns on whether the answer depends on something only that person did and whether two interviewers would rate it the same way, and all three types can pass or fail that test.

Keep the technical question only if the pass bar is written down

Keep a technical question only if you can write down, before the interview, what a passing answer contains. Without that it is a topic rather than an assessment, and two interviewers asking it will bring back two different verdicts about the same candidate. Write three or four sentences describing a strong answer, an adequate one and a weak one, and if you cannot, the slot is better spent elsewhere.

The pass bar also decides whether the question is about the job or about the interviewer's favorite corner of it. Writing the anchor forces the question, quietly, toward things the role actually needs, because it is hard to describe a weak answer to a trivia question without noticing that a weak answer would not hurt anybody on this team.

Two further tests before a technical question keeps its slot. Does it resemble the work, or does it resemble a puzzle with a trick? Puzzles measure exposure to puzzles. And could a competent person answer it in ten seconds with a search box or an assistant? If so, you are testing recall of something nobody recalls at work, which is fine as a warm-up and misleading as a signal. Running a coding or case interview when the candidate has an assistant open is the version of this problem that shows up live.

A question that survives all three is usually one where the candidate has to choose between two defensible approaches and say why. That is a technical question doing the work of a judgment question, and it holds up under follow-up in a way a recall question never has.

Put the retrievable material outside the conversation

Anything a candidate could retrieve belongs in an exercise where retrieving it is allowed and visible, not in a conversation where you are guessing whether they knew it. Public problem banks stopped functioning as work samples some time ago: evaluated against 115 Python problem statements taken from HackerRank, Codex solved 96% of them zero-shot, and the authors report clear signs the model was reproducing memorized code 3.

The memorization half is the important half and it usually gets dropped. HackerRank problems are public, so they were plausibly in the training data, which is the same reason the exercise was already weak: a candidate who had seen the problem was being scored against one who had not, long before any model was involved. That is a 2022 model on a curated benchmark, so read the number as a lower bound on what a model can do today, and note that it says nothing about a bespoke task in a real repository.

What to do with the material instead:

1. Move recall into a short written stage the candidate completes with whatever tools they use at work, and ask them to mark what they checked and what they discarded. 2. Keep the conversation for the parts that need a person present: the choice between two approaches, the constraint they were working against, what they would do if the requirement changed. 3. Score the second stage on judgment, since the finished artifact arrives clean regardless of who made it.

The reshuffle in one sentence: instances in conversation, retrievable material in an exercise, and the score taken from what came after the first answer. Live coding, a take-home, or an AI-allowed work sample is the choice waiting at the end of that move, and whether behavioral questions still predict anything is the one waiting at the start of it.

See what gets scored

Common questions

Are behavioral questions better than situational ones?

For a candidate who has done comparable work, yes, and for one who has not, no. Both opening answers are preparable now, so the difference sits in the follow-up. A behavioral question lets you push an instance into detail that only exists if the person was there. A situational question has a knowable good answer, which makes it the easiest type to prepare for and the quickest to run out. Use the behavioral version wherever the experience exists, switch to the hypothetical where it genuinely does not, and note in the scorecard which one you asked so ratings stay comparable.

What is a situational judgment test, and does it work?

It is a format rather than a construct: a set of scenarios with response options, usually scored against expert consensus. In the 2022 re-analysis these estimate at .26 against supervisor ratings of job performance, the same whether the instructions ask what a candidate should do or what they would do. That places them in the middle of the evidence. Two such tests can measure entirely different things, so the average tells you little about any particular instrument, and the studies behind the number measured people already in the job.

Should the technical round allow AI tools?

Decide it before the round, write it into the invitation, and apply it to every candidate. Both rules are defensible. Allowing tools makes the exercise resemble the work and moves the interesting evidence to what the candidate rejected and verified. Forbidding them makes a recall-heavy exercise interpretable but measures something the job no longer contains. What is not defensible is leaving it unstated, because candidates then guess differently and you end up comparing people who followed two different rules.

How many question types should one interview mix?

One, usually. An interviewer switching types mid-round switches what they are judging, and the ratings stop being comparable across candidates unless the guide fixes the sequence for everyone. Give each round a single job: instances here, judgment under uncertainty there, craft in the craft round. Mixing is fine at the loop level and expensive inside a single half-hour, where it mostly costs the depth that would have made any of the answers worth something.

Does the STAR framework still help?

It helps the candidate more than it helps you now. STAR is a way of organizing an answer, and organizing answers is exactly what preparation tools do best, so a well-formed STAR response is weak evidence of anything except preparation. Keep it as a prompt for candidates who ramble, and stop treating structure in the answer as a sign of quality. Take the evidence from what the person can say about the parts STAR leaves out: what they nearly did instead, who disagreed, and what happened after the result.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the corrected validity estimates used to compare question types: structured interviews .42, job knowledge tests .40, work samples .33, unstructured interviews .19.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection (accepted manuscript) Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the situational judgment test estimate of .26 under either instruction type, the head-to-head reversal that undermines cross-meta-analysis rankings, and the job-specific knowledge test comparison (.22 overall against .31 job-specific, rising to .40 corrected).
  3. 3. Codex Hacks HackerRank: Memorization Issues and a Framework for Code Synthesis Evaluation arXiv:2212.02684 (Karmakar, Prenner, D'Ambros and Robbes), 2022. arxiv.org Supports the claim that an exercise assembled from a public problem bank has stopped functioning as a work sample: 96% solved zero-shot across 115 HackerRank Python problems, with clear signs of memorization.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.