Assessment design

A Question Bank Row Is Four Fields, and One Person Can Delete

Keep the shared interview question bank and change the unit. A row is four fields, not one: the question, what a strong answer must contain stated as evidence rather than adjectives, one follow-up only somebody who did the work can answer, and the date the row was last reviewed. One named owner in recruiting operations holds delete authority. Assume every question is already public, and retire a row when the spread of answers collapses.

The takeThe bank's real failure mode is silent. Nobody complains when a question stops working; answers simply keep getting better, and a panel reads the improvement as a stronger pool. A library measured by how many questions it holds cannot see that, because the count rises either way. Measure the bank by how many rows carry an expected-evidence line and a review date inside the last year, and be willing to end a quarter with fewer questions than you started with.

Where Olive fits

Open a role and see what the work shows

Building this in-house, the expensive fields are the expected-evidence line and the follow-up, since both have to hold up once the question is public. Olive's cases are authored per occupation against that occupation's own work, and each report returns six findings a human reviewer wrote, every one anchored to a timestamped moment in the session.

Rank your shortlist

Should interview questions live in a shared bank?

Yes, and the reason is consistency rather than convenience. Structured means three things: every candidate gets the same questions in the same order, the answers are rated on a common scale, and the interviewers agree in advance on what an acceptable answer looks like 1. The third is the one a bank is supposed to carry, and a common scale means nothing without it. A bank holding only the questions has shipped one third of the thing.

What a vendor library sells is that first third. Hundreds of questions tagged by competency and role, on the premise that more is better, transfers the half that was never scarce. The scarce half is the bar an experienced interviewer carries unwritten: what a good answer has to contain, and which follow-up reaches past a prepared one. That half stays where it is, so a team can adopt a thousand-question library and run exactly the interviews it ran before, with better tagging.

The bank a loop actually uses is small. Five to eight questions per competency, each used often enough that somebody would notice if one stopped working, beats four hundred nobody has read. Count is the wrong measure and it is the measure every library reports, which is worth remembering when a demo opens with the size of the catalogue.

Write four fields per row, not one

Four fields, and the second one does most of the work: the question, what a strong answer must contain stated as evidence, one follow-up only somebody who did the work can answer, and the date the row was last reviewed. The second field fills with adjectives when nobody checks. Strategic thinking is not an expected-evidence line. Names a tradeoff they made and says what it cost is one.

A worked row, for a competency most teams have and few have written down:

  • Question: Tell me about a time you shipped something you knew was not right yet.
  • Expected evidence: names the specific defect they accepted, who they told, and what the decision cost. A candidate who describes the pressure but not the defect has not cleared it.
  • Follow-up: what would you have needed to know to make the opposite call? Only somebody who sat with the decision can answer that. Further from it, the answer returns to the pressure.
  • Last reviewed: a date, and the name of whoever looked.

Content is where the predictive weight sits, and one measured comparison shows how much. In the job knowledge meta-analysis the 2022 revision of the selection literature relies on, all 164 studies together produced a mean observed validity of .22, while the 59 studies using knowledge tests built for the job in question produced .31, rising to .40 once corrected for unreliable performance ratings 2. Read that narrowly. It compares subsets of one meta-analysis, the job-specific subset is smaller and may differ in other ways, the .40 carries no correction for range restriction, and job knowledge tests are given to people who already have the knowledge, so the evidence is about hiring experienced candidates. That is a finding about knowledge tests, and applying it to interview questions is an extension. The direction it points is clear enough: write the bank against the work this role actually does, and keep it small enough to write that way. The follow-up field is the one most people underestimate: what follow-up questions expose about real understanding is the craft underneath it.

Who owns the bank?

One named person in recruiting operations, with the authority to delete. Not a committee, and not everyone. A bank anybody can add to and nobody can remove from becomes an archive inside a year and a liability shortly after, because a stale question still gets asked by whoever opens the guide, and it still produces a rating that goes into a debrief.

Ownership splits cleanly if you say it out loud. The hiring manager owns the expected-evidence line for their own role, because they are the only person who knows what a strong answer has to contain in that job. The bank's owner owns the shape: every row has four fields, every row has a review date, and no row enters without the second and third filled in. Interviewers propose. The owner accepts or declines, and declining is not an insult when the standard is written down.

That standard is also the cheapest quality control available. A proposed question with no expected-evidence line is usually one somebody enjoyed being asked, and asking the proposer to write the line surfaces that in about two minutes. What makes an interview question worth asking is the longer version of the test the owner applies.

When does a question retire?

Retire on evidence, not on a schedule: when the spread of answers collapses and nearly everyone clears the bar, the question has stopped separating people and it goes, whatever the reason. The signal is easy to miss because it arrives as good news, so somebody has to look on purpose. The cheapest place to look is the debrief: if the last six candidates all cleared a question, say so out loud and put the row on the review list.

So assume the question is public and select for questions that survive being known. There is a measurement of what a public bank is worth, and it comes from code: evaluated against 115 Python problem statements taken from a popular competitive programming portal, OpenAI's Codex solved 96% of them zero-shot and 100% few-shot, and the authors report clear signs of the model reproducing memorized code rather than synthesising it 3. That is a 2022 model on a curated public benchmark, so read the number as a floor rather than a current capability, and nothing in it measures a bespoke task with real constraints. The memorization half is the part that transfers: a question pulled from a public pool tests a lookup, and it did so before any candidate opened a model.

What survives being known is a question about a specific thing this person did, because the answer sits in their own history and nowhere else. Rotation is the usual reflex and it is expensive, since a new question arrives with no expected-evidence line and no calibration behind it. Why every candidate suddenly gives the same polished STAR answer is the symptom, and whether to keep rotating questions every quarter is the decision it forces. Retirement is the narrower move: one row out, one row in, both with four fields filled.

See what gets scored

Common questions

How big should the bank be?

Small enough that the owner has read every row this year. For most teams that is five to eight questions per competency per role family, which is enough to vary a loop without asking the same three questions to every candidate, and few enough that a dead question gets noticed. Size is the metric libraries advertise and the one that predicts least. A bank of forty live rows with expected-evidence lines beats four hundred rows nobody maintains.

Does it matter that our questions are already on Glassdoor?

It matters less than the design of the question. Anything used in a live loop is public within weeks, and there are sites whose entire business is republishing it, so secrecy is not an available strategy. What changes under publication is which questions still work. A hypothetical has a best answer that can be rehearsed. A question about a specific thing the candidate did has an answer only they hold, and a follow-up asks for the part of it only somebody close to the work would have.

Can interviewers add their own questions?

They can propose, and the bank's owner decides. Free addition is how a bank grows past the point anyone can maintain it, and a blanket ban is how interviewers end up keeping private lists nobody can calibrate. The standard does the work: a proposed row needs the expected-evidence line and the follow-up filled in before it enters. Most weak proposals fail that on their own, without anyone having to reject a colleague.

Should the bank hold the expected answers, or is that a leak risk?

Hold them, and accept the risk, because the alternative is worse. Without a written expected-evidence line, every interviewer applies a private bar, and two ratings on the same answer become unreconcilable. If a specific answer key would be genuinely damaging in public, that is a signal about the question rather than about the bank: it means the question has one right answer that can be memorized, which is the kind that stops working first anyway.

Do we need a separate bank per role?

One bank, with the expected-evidence line written per role. The questions themselves travel further than people expect, since how someone handles being wrong looks similar across jobs. What does not travel is what a strong answer has to contain, which is entirely a function of the work. So share the rows, fork the second field, and make the hiring manager write it at intake, while it is still cheap to be honest about the bar.

References

  1. 1. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three defining properties of a structured interview used to argue that a bank holding only questions ships one third of what structure requires.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology, 107(11), 2040-2068 (American Psychological Association); accepted manuscript hosted by co-author Filip Lievens, 2022. static1.squarespace.com Supports the claim that job-specific content carries the predictive weight: .22 across all 164 job knowledge studies against .31 for the 59 job-specific ones, rising to .40 corrected for criterion unreliability.
  3. 3. Codex Hacks HackerRank: Memorization Issues and a Framework for Code Synthesis Evaluation arXiv:2212.02684 (Karmakar, Prenner, D'Ambros and Robbes), 2022. arxiv.org Supports the claim that a question drawn from a public pool tests a lookup: 96% zero-shot and 100% few-shot on 115 public Python problems, with the authors reporting memorized code.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.