Interviewing

Anchor the Question to Their Work Instead of Rotating It Every Quarter

You do not have to keep rotating interview questions now that AI can answer the standard ones. Rotation costs you the thing that makes an interview work, which is asking every candidate the same question and rating the answers on one scale, and a model turns your posting into plausible variants in a minute, so the treadmill cannot be won. Publish the competencies, keep a stable bank, and anchor each question to something that exists only in that conversation: the candidate's own work, your real material, the follow-up chain.

The takeThe advice to strip competencies out of the posting deserves a harder look than it gets. It taxes every honest applicant, who now cannot tell whether to apply, in order to slow down a guesser who will produce a decent guess from the duties list regardless. It also breaks the posting's actual job, which is helping the right people apply and the wrong people stop. Any tactic whose cost falls on the people behaving well, and whose benefit is a few days of surprise, is a bad trade.

Where Olive fits

Open a role and see what the work shows

A question that survives being published is one whose answer has to be produced rather than recalled. Olive runs on that footing: a role-grounded assignment worked with an AI assistant, and six findings a human writes out of what the session actually contained.

Rank your shortlist

Why does rotating the bank make the interview worse?

Because sameness is the mechanism. A structured interview is defined by three properties: the same questions in the same order, a common rating scale, and interviewers who agreed in advance what an acceptable answer looks like 1. Rotate the questions every quarter and the second and third go with the first, since nobody has heard enough answers to a new question to know what a good one sounds like.

Calibration is the hidden cost, and it is the expensive one. A question earns its rating anchors over dozens of answers: this is what a 2 sounds like, this is where a 4 stops. Replace the question and every interviewer is back to rating against a private standard, which is the condition under which the same answer gets a 3 from one panelist and a 5 from another. Two quarters of that and the scorecard is decoration.

Most teams are also further from structure than they think. In SHRM's benchmarking survey of a random sample of its member organizations, 34% used structured interviews for executive hires, 37% for middle management and 36% for individual contributors, while in-person interviews ran at 79%, 78% and 76% of organizations 2. That is self-reported against SHRM's own definition, from 2021, among organizations large enough to have an HR function, so read it as an upper bound rather than a census. It still says the common case is an unscored conversation, and rotating the questions in an unscored conversation changes nothing that was ever measured.

What should be anchored instead of rotated?

Three anchors, none of which a preparation session can reach. The candidate's own work, because you did not know what they would bring. Your real material, because it is not published anywhere. And the follow-up chain, because the third question in it depends on the second answer. Keep the questions stable and let the anchors supply the part that cannot be looked up.

The artifact anchor is the cheapest to add: name it in the invitation, ask for a recent piece of the candidate's own work, and spend fifteen minutes on the decisions inside it. The material anchor takes more preparation and pays more back. A redacted case, a real spreadsheet with a genuine inconsistency in it, an actual support ticket, a paragraph of your own documentation that is out of date. Ask what they would do with it, then ask what they checked.

That second anchor is a work sample in miniature, and the evidence supports the family without overselling it. The 2022 re-analysis estimates work sample validity at .33, which is a correlation with supervisor-rated performance and not a hit rate. That figure replaced a .54 that traces back to a 1974 narrative review, and 53 of the 54 studies behind the newer number tested people already doing the job rather than applicants 3. Do not read .33 against .42 for structured interviews as work samples losing; the credibility intervals overlap and the paper argues the pattern is coherent. Adoption is the real gap: work sample interviews ran at 12%, 11% and 9% of organizations in the SHRM data 2.

If the temptation is to go the other way and set a trap instead, whether to hide an instruction in the job posting to catch AI-written applications works through what that actually filters for.

Publish the competencies and let candidates prepare

Put them in the posting. Name the four or five things the round will rate, the format of each stage, and roughly how long it takes. This is the opposite of the advice circulating now, and it is the cheaper position: a candidate who knows what is coming arrives ready to talk about the right work, and one who reads it and self-selects out saves everybody a round.

What to publish and what to keep back divide cleanly. Publish the competencies, the number and shape of the stages, the rating dimensions, and whether AI tools are permitted at each stage. Keep back the specific material a candidate reacts to, the artifact you will hand them, and anything with a deliberate flaw in it. The first list is about fairness and self-selection. The second is about the round staying live.

The worry underneath the hiding instinct is that a well-prepared candidate will look better than they are. Sometimes true, and the follow-up chain is what handles it: preparation improves the first answer and does very little for the third. Whether a structured interview still separates candidates once everyone has been AI-coached takes that question on directly, and what belongs in a published evaluation process covers how much detail to put in the posting itself.

How often should the bank actually change?

When the evidence says so, which is roughly one question a year for most roles. Retire a question for a reason you can state: it no longer separates anybody, the competency behind it left the job, or the answers stopped containing anything checkable. A calendar is not a reason, and a rumour that a question appeared on a forum is not one either, since a question good enough to survive publication was never relying on secrecy.

Three signals worth watching in the scorecard data. First, a ceiling: if almost every candidate scores at the top, the question has stopped discriminating and the anchors need tightening before the question needs replacing. Second, disagreement that does not resolve: if two calibrated interviewers keep landing three points apart, the question is ambiguous rather than hard. Third, drift in the role: if a competency has genuinely changed, the question follows it, and that happens on the timescale of the job changing rather than quarterly. Whoever watches those signals needs the authority to act on one: who owns the question bank, and who is allowed to delete a row is the assignment that turns a signal into a retirement.

When you do change one, change one. Keep the rest of the bank and the same scale, run the new question alongside the old set for a quarter, and compare how it behaves before it counts for anything. That is also the honest way to answer whether a change helped: hold the rest constant, or you will never know which move did it. If the whole bank feels stale at once, the problem is usually the scorecard rather than the questions, and what to put on an interview scorecard when AI use is in scope is the better place to start.

See how it works

Common questions

What if our questions are already posted on Glassdoor?

Then they have been tested against the condition that matters. A question whose value depends on nobody having seen it was fragile before any of this, since candidates have swapped interview questions publicly for well over a decade. Read the thread, check whether the posted answers would score well on your own anchors, and fix the anchors if they would. A question that still separates people after publication is one worth keeping.

Should we use AI to generate new interview questions?

It is fine for drafting and useless as the final step. A model will produce forty grammatical questions from a job description in a minute, and the work that matters starts after that: mapping each to a competency, writing rating anchors, and testing them on people whose performance you already know. The generation was never the expensive part, which is why generating more of it does not help.

Does a bigger question bank solve this?

It makes it worse. A large bank means each candidate faces a different subset, which destroys comparability and multiplies the calibration problem by however many questions are in rotation. Six to ten well-anchored questions, used consistently and retired on evidence, beat a bank of sixty that nobody has heard enough answers to rate reliably.

How do we keep a panel from all asking the same thing?

Assign competencies rather than questions. Give each interviewer two of the four or five competencies, with the specific questions attached, and tell them not to stray into another's territory. The overlap problem is nearly always a briefing failure rather than a bank problem, and it costs a whole stage when two panelists spend forty minutes on the same competency and nobody covers the third.

Is it worth rotating questions for confidentiality reasons?

For a work-sample task with an answer key, yes, on a slower cycle. Anything with a right answer degrades once it circulates, so plan for versions and change the underlying material rather than the wording. Behavioral questions are the opposite case: there is no key to leak, and the answer still has to be produced by a person with the events to produce it.

References

  1. 1. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three-property definition of a structured interview that rotation breaks: same questions in the same order, a common rating scale, agreement in advance on an acceptable answer.
  2. 2. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the adoption figures for structured interviews, in-person interviews and work sample interviews across executive, middle management and individual contributor hiring.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .33 work sample estimate replacing the older .54 figure, the concurrent-sample limit behind it, and the .42 structured interview comparison.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.