Assessment design
Pick One Load-Bearing Claim Before the Call, Then Stay on It
A screening call should test the one claim the hire rests on: the claim that, if false, makes the rest of the application irrelevant. In practice that is a claim about scope rather than a tool or a title, such as the size of the thing the candidate ran, the decision they owned alone, or the part that went wrong. Write it down before the call, together with the specific detail a real owner would know and a plausible outsider would not. That detail is the test.
The takeQuestion banks are the wrong shape for this round, and their popularity is inertia. A list of screening questions covered a resume back when the resume was a rough draft the person wrote themselves and the risk was missing a line. A prepared answer now exists for every line, so breadth costs the whole call and buys the document back. Depth on one claim is the only part of a screening conversation nobody can prepare from the posting.
Where Olive fits
Open a role and see what the work shows
Build this in-house and the expensive parts are the answer key and the evidence trail behind each finding. Olive ships twelve item banks, each grounded in one occupation, and returns six separately-evidenced findings, each anchored to a moment in the session.
Rank your shortlistWhich claim should a screening call test?
The one the hire rests on. Read the application and ask which single statement, if it turned out to be false, would make the rest of it irrelevant. That is usually a claim about scope: what they ran, how big it was, which call was theirs alone. Rarely a tool, never a title, because both of those are cheap to write and cheap to say.
The test has a useful negative form. If a claim could be false and you would still want to interview the person, it is not the load-bearing one. "Familiar with Python" fails that test for most roles. "Owned the migration off the old billing system" passes it for a role that exists because of a migration. Run the question over the application and one claim usually stands out inside a minute.
For a role with no single dominant claim, the choice is between the largest scope claim and the most recent one. Prefer the most recent, because memory is better and because the candidate's story about it has had less time to harden. Prefer the largest only when the requisition exists specifically because of scale.
Write the chosen claim on the form in the application's own words. Paraphrasing it into your language is where the test starts to slip, since you will unconsciously ask about the version you understand.
Why choose scope over tools and titles?
Because a title is a claim about prior experience, and prior experience measured at the point of hire predicts almost nothing. A meta-analysis of 81 independent samples found prehire work experience correlating .06 with later job performance and .00 with turnover, and experience with tasks, jobs or occupations relevant to the current position did no better 1. Nothing in that evidence covers a tool list, which is the cheapest line on a resume to write in the first place.
Those are corrected correlations, so the near-zero result is not an artefact of under-correction, and the claim is narrower than it sounds. It is about experience measured at hire, which is what a resume reports. It is not about tenure in the job once someone is in it, and it says nothing about licensure or roles where experience is legally required.
What does carry weight is relevance to the actual work. In the job knowledge literature the 2022 selection re-analysis relies on, 164 studies produced a mean observed validity of .22, while the 59 using knowledge tests built for the job in question produced .31, rising to .40 once corrected for unreliable performance ratings 2. That comparison sits between subsets of one meta-analysis and was never a controlled experiment, and job knowledge tests presuppose candidates who already have the knowledge. The direction still points the same way: what an assessment is about does more work than what shape it takes.
A scope claim is the closest thing a twenty-minute call has to job-relevant content. It is about the work, it has edges, and its details sit with whoever did it. A tool list has none of those properties, which is why a candidate listing six AI tools needs handling on its own terms and will not carry a round.
Decide what a true answer would contain
Write the detail before you dial. A person who genuinely owned the thing knows a number, a constraint, a name, or the specific way it went wrong; a person who read the posting knows the shape and not the grain. Pick which grain you expect. One of the three properties that define a structured interview in the US Office of Personnel Management's guide is interviewers in agreement on acceptable answers, settled while the interview is built 3.
The other two properties are the same questions in the same order and a common rating scale. That guide was written for US federal hiring in 2008; it is guidance, not statute, and it binds no private employer. The agreement is the cheapest of the three to adopt, since it needs one sentence on a form.
Worked examples, so the abstraction has edges:
- Claim: "led the replatform of the checkout flow." Expected grain: what broke in the first week after cutover, and who decided to ship anyway.
- Claim: "managed a team of six." Expected grain: the last hire and the last exit, and what changed about the team's work between them.
- Claim: "owns the monthly close." Expected grain: which account is the one that always needs a manual adjustment, and why.
- Claim: "built the reporting the exec team uses." Expected grain: the metric someone argued with, and how it was resolved.
None of those can be answered from a job posting, and all of them can be answered on the spot by whoever did the work. That asymmetry is the entire instrument, and it disappears the moment you accept the first general answer and move on.
How far down should you follow one answer?
Until you get the detail or a second layer of generality, whichever comes first. Three follow-ups is usually enough for both outcomes. The pattern to watch is direction: a real owner gets more specific under pressure, moving toward names, numbers and things that went badly, while a rehearsed answer gets broader, moving toward principles and lessons learned.
The follow-ups are built out of the answer as it comes, which is why a question bank cannot hold them. Three shapes cover most of it: ask for the alternative that was rejected, ask what went wrong, ask who disagreed. Each one requires the answer to have had a real second option, a real failure, or a real other person in it. What follow-up questions expose about understanding goes further into the mechanics, including what preparation does and does not cover.
Two qualifications keep this honest. A second layer of generality is evidence the claim is thinner than the resume implied; it is not evidence of bad faith, and plenty of people describe their own work badly under time pressure. And nerves, an unfamiliar accent, a bad line and a first interview in three years all produce vagueness that has nothing to do with ownership. Ask again in a different direction before drawing a conclusion, and apply the same number of chances to everybody.
The last boundary is the important one. This establishes whether the claim belongs to the candidate. It does not establish whether the candidate is good at the job, which is a separate question that a twenty-minute call cannot reach and should not pretend to. It sits before an assessment and never in place of one, which is also why the length of the call is best budgeted in turns.
Common questions
What if the application has no claim worth testing?
The screen has nothing to do, and that is a finding about the top of the funnel rather than about the candidate. A resume with no scope claim on it is usually early-career, deliberately generic, or written entirely from the posting. For early-career candidates, substitute a claim about something they built or ran outside work and probe that the same way. For the generic case, the fix is upstream: the application form should be asking for one specific thing that a posting cannot supply.
Should I tell the candidate which claim I am testing?
Yes, at the top of the call, and it makes the round better rather than weaker. Saying that you want to spend most of the time on one project removes the pressure to cover everything and gives the candidate a fair chance to pick the right memory. The follow-ups are what carry the test, and they are built from the specific answer given, so announcing the topic costs nothing. It also produces a much better experience for people who prepared for a resume tour.
Is one claim enough to make a decision on?
It is enough for the decision this round is making, which is whether to spend an hour of somebody's time. It is not enough to make a hire on, and no screening call is. Treat the outcome as one input with its own known limits: it establishes ownership of one claim, on one project, as described by the person who did it. Everything about capability, judgment and how someone works belongs to rounds designed to observe those directly.
How do I keep this fair across candidates for the same role?
Pick the claim type in advance for the requisition rather than per candidate, and use the same number of follow-ups for everyone. The claim itself differs because each application differs, but the rule generating it should not: for this role, probe the largest scope claim in the last two years, expect one concrete detail, allow three follow-ups. Written that way the rule can be shown to a colleague, applied by a second screener, and defended later, none of which is true of picking whatever looked interesting.
Does this work for internal candidates?
It works, and the claim usually has to change. For someone already inside the company, ownership of the work is often verifiable without asking, so probing it wastes the call. Move the test to the part nobody can see from the org chart: which decisions were actually theirs, what they would do differently, and where they needed help. The structure is the same, one claim and three follow-ups, but the load-bearing claim is about autonomy rather than about scope.
References
- 1. A meta-analysis of the criterion-related validity of prehire work experience digitalcommons.unf.edu Supports the claim that prehire work experience, which is what a resume reports, predicts later performance and turnover close to zero.
- 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range static1.squarespace.com Supports the job-knowledge comparison used to argue that relevance to the job carries more than assessment format.
- 3. Structured Interviews: A Practical Guide opm.gov Supports the definition of a structured interview, including agreeing in advance what an acceptable answer contains.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.