Policy

Fix the Questions and the Anchors; the Follow-Ups Can Differ

The same planned questions and the same written rating anchors go to every candidate; the exact wording of a follow-up does not have to match. Keep the order fixed, record what each person actually said rather than only rating it, and treat probing as the deliberate exception: safe because the ladder was planned in advance, offered to everyone, and aimed at the same thing for everyone. Count who gets the depth.

The takeRead as a ban on deviation, the rule now damages what it was written to protect. When every opening answer arrives equally polished, an interviewer forbidden to go further is collecting delivery and calling it evidence. Identical wording was only ever a proxy for identical opportunity, and it has stopped tracking the thing it stood for. Fix the set, fix the anchors, plan the ladder, and then spend your policing effort on the distribution of depth rather than on the transcript matching word for word.

Where Olive fits

Open a role and see what the work shows

Under the automated-decision rules, a number that stood for a candidate explains nothing on its own. Olive produces no composite at all: a person writes each of the six findings, every one carries the excerpt it rests on, and a released report exports with its rubric, scorer and bank versions attached.

Rank your shortlist

What has to be identical, and what does not?

The question set, the rating anchors, the opportunity to be probed, and what each probe is trying to find out. The US Office of Personnel Management's practical guide defines a structured interview by three properties: every candidate is asked the same questions in the same order, every candidate is evaluated on a common rating scale, and the interviewers agree in advance on what an acceptable answer looks like 1. Those three fix what is asked and how it is judged.

They say nothing about the sentence an interviewer speaks after the answer arrives, because the guide handles that in a step of its own. Step five of its eight-step development process is Create Interview Probes: establish the range of probing before the interview, draft the specific probes each question allows an interviewer to use, and hold the general meaning of a probe constant while its wording is tailored to the response just given, so that candidates are given "the same opportunities to excel" 1. The clause is in the source. The compression to "same questions, same order" is what drops it, leaving a rule that reads as a prohibition on saying anything unplanned.

So the honest boundary has four fixed items and one variable one:

  • Fixed: the planned questions, in the same order, for every candidate for that role.
  • Fixed: the written anchors, describing what strong, adequate and weak answers contain, applied by every interviewer.
  • Fixed: the opportunity to go deeper, meaning the planned follow-up ladder exists in the guide and is available to every candidate.
  • Fixed: the general meaning of each probe, which is the guide's own limit and the one the restatements never reach.
  • Variable: the wording of a probe, which will and should differ, because it responds to what the person in front of you just said.

The OPM guide is federal HR practice guidance from 2008, written for merit-system hiring, not a statute binding a private employer, and it predates every question anyone now has about AI-prepared candidates. It is cited here for what "structured" means and for what it already says about probing, which is the part people are actually arguing about.

Why does the rule get read as a ban on deviation?

Because the compressed version is easier to enforce than the real one, and because the fear underneath it is real. A team that hears "treat every candidate the same" and has no rubric will reach for the one thing it can verify, which is whether the words matched. Identical wording is checkable from a transcript. Equal opportunity to be probed is only checkable if somebody records it, and the standard scorecard has no column for it.

The fear is not misplaced, it is just aimed at the wrong object. Unscored conversation genuinely is where trouble lives, and the federal selection guidelines agree: the Uniform Guidelines define a selection procedure to cover the full range of assessment techniques "through informal or casual interviews and unscored application forms" 2. Swapping a scored assessment for a chat does not move a hiring step outside the rules. That regulation dates to 1978 and names no software, and it imposes a validation burden only where adverse impact appears, so it is a scope definition rather than an audit mandate.

What has changed is the cost of the compressed reading. An opening answer to a well-known question now arrives structured, complete and rehearsed whenever the candidate prepared, so an interviewer holding still after it is recording delivery. The rule was written when the first answer was the candidate's own material. It is now the part they practiced. Whether a structured interview still separates people who have all been AI-coached is the same problem seen from the candidate's side.

Say the corrected rule in one line to your panel: *the questions and the scale are fixed, the probes are planned but phrased in the moment, and everyone gets offered the same depth.*

Check who gets probed rather than which words were used

Count follow-ups per candidate and read the counts before you read the ratings. This is the measurement the consistency rule was reaching for and never specified. If the candidates who present well consistently receive three rungs while the others receive one, the loop has an unequal-treatment problem the rubric cannot see, because a rubric only judges the answers that were allowed to develop.

Structure is still the intervention that helps most, and the instinct it corrects is the belief that an experienced manager's deviation is informed. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by just over 25%, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 3.

Read that carefully, because it is easy to over-claim. It is not a randomized experiment, tenure is a match-quality proxy in a high-turnover service setting rather than a performance rating, the authors observe hires but not offers, and overrides were common rather than pathological at an average exception rate of 22%. It says the average override was worse; it never says the test was right about any individual person, and the instrument was a written job test rather than an interview. What transfers is the direction. The confidence that justifies giving one candidate extra room and another none is the same instinct that was measured there, and it came out negative.

The operational move is small. Put the planned ladder in the guide beside each question, so no interviewer has to invent depth on the spot and no candidate depends on being interesting to receive it. Then add a probe count to the scorecard. How far down to go before the return stops is the discipline that column is measuring.

Write the evidence next to the rating

Require the words the candidate used, in the scorecard, next to the number the interviewer chose. A rating with no excerpt under it cannot be reconstructed a week later and cannot be argued with in a debrief, which means it also cannot be defended if the decision is ever challenged. This is the cheapest half of consistency and the half that gets skipped.

There is evidence that what interviewers write carries information a rating alone does not. Text-mining post-interview notes on 7,650 candidates hired at a large Chinese technology company, researchers found that the number of job-related capabilities an interviewer named in the notes was positively related to later job performance and promotions and negatively related to turnover 4. Only people who passed and joined could be observed, the effect is small, it is one firm in one country over 2016 to 2018, and it measures whether notes name the right capabilities rather than whether notes were taken at all.

The defensibility half is worth stating plainly, and it is public legal fact rather than advice: under 42 U.S.C. 2000e-2(k), added by the Civil Rights Act of 1991, a disparate-impact claim is made out only if the complaining party shows that a particular practice causes a disparate impact and the employer then fails to demonstrate that the challenged practice is job related for the position in question and consistent with business necessity 5. Job-relatedness is a claim about content, and content is what an excerpt records. A scorecard of bare numbers proves that ratings happened. Anything about a US requisition still belongs in front of counsel, since state and local rules add duties the federal text does not carry.

Most of this closes with three lines in the guide. Name the planned probes under each question. Add a column for how many were used. Add a box for one quoted sentence per rating. Getting a panel to judge the same thing the same way is the calibration work those three lines make possible, because until the evidence is written down there is nothing for two interviewers to disagree about except each other.

Read the evidence

Common questions

Is asking a follow-up question legally risky?

The risk lives in who receives depth, not in whether a probe was scripted. An interview is a selection procedure under the 1978 federal Uniform Guidelines whether it is scored or casual, so consistency matters, but consistency means the same questions, the same rating anchors and the same opportunity to be probed by probes that mean the same thing. Plan the ladder in advance, make it available for every candidate, and record how much of it each person actually got. A probe phrased in the moment is normal interviewing. A probe offered only to some candidates is the exposure. Check any specific process with counsel.

Does every candidate have to hear the questions in the same order?

Yes, keep the order fixed. That is the definition most guides use and following it costs nothing. The reason is less about fairness than about comparability: an answer given after twenty minutes of conversation is not the same artifact as the same answer given cold, so varying the order quietly varies the conditions. Order also drifts on its own when interviewers run late and start cutting, which is a bigger practical threat to comparability than deliberate reordering ever is.

Can different panelists ask different questions?

Yes, as long as each panelist asks the same questions of every candidate for that role. Splitting competencies across a loop is normal and useful, and it stops the same ground being covered four times. What breaks is when one interviewer improvises a different set for each candidate, because their ratings are then measuring different things and cannot be compared or combined. Assign the competencies, write each panelist's questions into their own guide, and hold the assignment steady across the whole slate.

What if a candidate volunteers something that changes what I want to ask?

Follow it, then return to the planned set. New information is exactly what a probe is for, and refusing to hear it protects the transcript at the cost of the decision. What to avoid is letting the detour replace a planned question, because that candidate then has a gap in their evidence that others do not. Note in the scorecard what you followed and why. If the same detour keeps arising with multiple candidates, the guide is missing a question and should gain one.

How do I show a challenged hiring decision was consistent?

With the artifacts, not with a recollection. The strongest record is a guide showing the questions and anchors everyone faced, scorecards carrying a quoted excerpt under each rating, and a probe count showing depth was distributed evenly. Ratings alone show that a process ran; excerpts show what it was judging and connect the judgment to the work. Retention rules vary by jurisdiction, so agree with counsel on how long these are kept before the first requisition opens rather than after a challenge arrives.

References

  1. 1. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three properties that define a structured interview (same questions in the same order, a common rating scale, agreement in advance on an acceptable answer) and the guide's own step five, Create Interview Probes, which has the employer set the range of probing in advance, draft the specific probes each question allows, and hold a probe's general meaning constant while its wording is tailored. Together those are the boundary this article draws.
  2. 2. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that an informal or casual interview is a selection procedure under federal guidelines, so replacing a scored stage with a conversation does not move it outside the rules.
  3. 3. Discretion in Hiring National Bureau of Economic Research, Working Paper 21709 (November 2015, revised September 2017); published 2018 in the Quarterly Journal of Economics, 2017. nber.org Supports the claim that managers who override a structured hiring signal produce shorter job tenures on average, with the paper's own limits attached, including an average manager exception rate of 22%.
  4. 4. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company Frontiers in Psychology, Volume 11, Sec. Organizational Psychology (Shanshi Liu, Yuanzheng Chang, Jianwu Jiang, Haigang Ma and Huaikang Zhou), 2021. frontiersin.org Supports the claim that notes naming the job-related capabilities a role calls for carry predictive signal a bare rating does not, in a range-restricted single-firm sample of 7,650 hires.
  5. 5. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases Office of the Law Revision Counsel, United States Code (prelim), 1991. uscode.house.gov Supports the statement of the operative federal test, job related for the position in question and consistent with business necessity, used here to explain why written evidence beats a bare rating.

5 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.