Interviewing
Do Structured Interviews Still Separate AI-Coached Candidates?
A structured interview still separates AI-coached candidates, if you replace the questions. Structure itself holds: even the 2022 re-analysis that cut most validity estimates ranked structured interviews first. What coaching defeats is the public behavioral bank, exactly what mock-interview tools rehearse against. Draw six to eight questions from decisions your team argued about last quarter, ask what happened rather than what would, and probe for something checkable outside the story: a figure, a date, who disagreed. A work sample answers whether someone checks a claim better than any question.
The takeHere is the part nobody says out loud. A universal question bank was always a way to run interviews without deciding what the job is, and coaching only removed the cover. Sit a panel down to write six questions from decisions it argued about last quarter, and watch what happens when it cannot name six. That is a finding, and it is about the team rather than about anyone in the chair. Nobody has measured how often a round fails there instead of in the scoring, but I would put more of it there than most panels would admit.
Where Olive fits
Open a role and see what the work shows
A structured question can capture a candidate describing the source they would open; it cannot capture them opening one. Olive puts that in front of them as work: a 40-to-60-minute occupational assignment, an assistant willing to do all of it, and a human reviewer who writes six findings, each tied to the moment it rests on.
Rank your shortlistDoes AI coaching actually beat a structured interview?
No. It beats the questions, not the structure. In the 2022 re-analysis that pulled most selection-procedure validity estimates down by .10 to .20 points, structured interviews still emerged as the top-ranked procedure 1. Structure is doing real work, and it keeps doing it. What coaching erodes is the assumption underneath a shared question bank: that the answer was composed by the person saying it.
Two findings sit either side of that. Candidates who used generative AI to prepare received higher overall interview performance ratings than unassisted candidates, and what they delivered was often a polished, contextualized response they were repeating rather than composing 3. And faking is not new: in the study that built the interview faking behavior scale, over 90% of undergraduate job candidates faked something during an employment interview, with the subset closer to outright lying running from 28% to 75% 2.
Put those together and the diagnosis is narrower than "the interview is broken." Coaching did not invent embellishment. It removed the effort barrier, and effort was the thing quietly sorting your candidates, separating the ones who had thought hard about the question beforehand from the ones who hadn't. That variance is gone, and it was never the variance you wanted anyway. Every answer arriving polished is a measurement problem, not a character problem.
The structure itself is untouched. Same questions in the same order, a common rating scale, interviewers who agree in advance on what an acceptable answer contains 4: none of that depends on the questions being secret.
Why has the universal behavioral bank stopped separating anyone?
Because it is public, finite, and old. Tell me about a time you disagreed with your manager, describe a failure, walk me through a conflict. That list has been printed in every career guide for decades, which makes it ideal training material. A tool that rehearses a candidate against forty of those items is not guessing at your process. It has already seen it.
The failure is quieter than cheating. A coached answer arrives correctly scoped, in the situation-task-action-result shape your rubric was written to reward, with a plausible number in the result. An anchored rating scale rewards exactly those properties, because those properties were once an honest proxy for having done the thing. Now they are a proxy for having prepared with a tool that produces them on request 3.
So the scores compress. Everyone lands at a 3 or a 4 on a five-point anchor, the panel argues about tie-breaks, and the decision quietly relocates to whoever was most likeable in the room, which is the unstructured interview wearing a rubric.
The tell is easy to check against your own data. Pull the last thirty scored interviews for one role and look at the spread per question. A question where nobody scored below the midpoint is not a hard question. It is a question that has stopped asking anything.
Build the question set from last quarter's real decisions
Take six to eight decisions your team actually made in the last quarter (the ones that were argued about) and write one question per decision. The job analysis that OPM's structured-interview guide puts first is exactly this, done at the level of your own field rather than at the level of a competency label 4. The bank becomes unguessable because it is local.
The mechanics are in that guide and they are worth following literally. Convene six or seven experienced people who do the job, have them write the questions against the competencies the analysis produced, and keep each question aimed at a specific situation, the actions the person took or did not take, and what those actions changed 4. Superlatives do the narrowing: the last one, the worst one, the one you got wrong 4. Four to six competencies is the normal load for a single interview 4.
Written out by field, the questions look like this:
- Claims: "Tell me about the last file you conceded. What in it made conceding right, and who disagreed?"
- Analytics: "Describe the most recent number you published that turned out to be wrong. How did you find out?"
- Marketing: "Walk me through the last channel you killed. What did you stop believing?"
None of those can be answered from the wording alone, and none of them sit in a mock-interview bank, because they came out of a room. The rubric still transfers (what the rows say for an AI-assisted answer is the same work either way), but the questions do not, and that is the cost you are accepting.
Rewrite roughly once a year, or whenever a question stops producing spread.
Which questions survive coaching, past behavior or situational?
Past behavior, on the only direct evidence available. Levashina and Campion's faking experiment found past-behavior questions more resistant to faking than situational ones. It also found, in the half nobody quotes, that follow-up questioning increased faking rather than reducing it 2. Over 90% of the undergraduate candidates in that work faked something during an employment interview 2. Coaching did not create this problem. It industrialized it.
That second finding deserves more attention than it gets, because probing is the standard advice. The MIT Sloan piece on AI-prepared candidates recommends five kinds of follow-up: break down the process, ask why the decision led to the outcome, probe the situational limits, ask what alternatives were weighed, ask what the drawbacks were 3. Those are good questions. They are also more narrative surface, and narrative surface is what an embellished story is made of.
The reconciliation is in what a probe asks for. A follow-up that requests more story invites more story. A follow-up that requests something checkable outside the story does not:
- Not "why did that work?" but "what was the number before and after?"
- Not "what would you do differently?" but "who told you it was wrong, and when?"
- Not "how did you approach it?" but "what was the option you rejected, and what did rejecting it cost?"
Names, dates, figures and the losing option are hard to keep consistent across four minutes of an invented account, and trivially easy to produce when the account is real. How far a second question can get has limits, but that is where the remaining separation lives.
Situational questions are not worthless. They are still fine for judgment on a scenario nobody has met. Just do not carry the weight of the decision on them.
What structure still can't reach
A description of work, however well probed, is still a description. A structured interview can capture a candidate saying they would check the model's most quotable claim against the source; it cannot capture them checking one. That gap existed before AI coaching, and coaching widens it, because fluent description of a process is the specific thing a model produces best.
Three limits worth naming to whoever signs off on the round.
The questions are real work and they go stale. Six to eight local questions cost a room of experienced people an afternoon plus a rating scale, and once thirty candidates have been through them, the good ones leak. Budget for a rewrite, not for a build.
Structure is also the thing carrying your legal position. An interview used to decide who gets hired is a selection procedure, and it has to be job-related and consistent with business necessity 5. Questions traced to a documented analysis of your own job help you there; a bank someone downloaded does not. The standardization (same questions, same order, agreed answers) is what makes the round defensible, so change the source of the questions and leave the process alone 4.
And the format has a ceiling. If what you need to know is whether someone checks a confident claim, a work sample answers that better than a question about it does. If the worry is a candidate reading generated answers during the call rather than before it, that is a different problem with different tells.
None of this means structure is finished. It means the structure was never the part that was secret, and the part that was secret is now public. Move the secret.
Common questions
Should you tell candidates the questions come from your own work?
Yes, in the invite, in the same words for everyone. Say the questions come from decisions the team actually made and that specific examples are what gets rated. It removes the advantage of having memorized a public bank without penalizing anyone who prepared honestly, and it cuts down on candidates arriving with a rehearsed script for a question you are not going to ask.
Does adding follow-up probes fix a coached answer?
Probing alone tends to make it worse. In the experiment that built the interview faking behavior scale, follow-up questioning increased faking rather than reducing it, because more narrative surface gives an embellished story more room to grow. Probes work when they ask for something checkable outside the story: the figure before and after, who disagreed, what the rejected option cost. Those are easy to produce from a real memory and hard to keep consistent from an invented one.
How many local questions do you actually need?
Six to eight, covering four to six competencies, which is the normal load for one structured interview. More than that and the panel runs out of time to probe; fewer and one bad question decides the round. Write them off decisions the team argued about last quarter, and expect to replace two or three a year as they leak or stop producing any spread in the scores.
What if a candidate prepared with AI and still answers well?
Then they answered well. Preparation was never the problem. A candidate who used a tool to think harder about your question before arriving has done the thing you would want an employee to do. The problem is only that a polished answer no longer distinguishes preparation from capability. That is what the checkable-detail probe is for: it separates the two without penalizing anyone for having prepared.
Do unique questions weaken the legal defensibility of a structured round?
No, provided they stay standardized and job-related. Defensibility comes from every candidate getting the same questions in the same order against the same rating scale, with agreement in advance on what an acceptable answer contains, not from the questions being generic. An interview used to make a hiring decision is a selection procedure and must be job-related and consistent with business necessity, and questions traced to an analysis of your own job are easier to defend on that ground, not harder.
References
- 1. Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range ✓ europepmc.org Revised validity estimates cut most high-ranked selection procedures by .10 to .20 points, and structured interviews emerged as the top-ranked selection procedure.
- 2. Measuring faking in the employment interview: development and validation of an interview faking behavior scale ✓ europepmc.org Past behavior questions were more resistant to faking than situational questions, follow-up questioning increased faking, and over 90% of undergraduate job candidates faked during employment interviews (28% to 75% for behavior closer to lying).
- 3. When Candidates Use Generative AI for the Interview ✓ sloanreview.mit.edu Candidates using generative AI received higher overall interview performance ratings, often delivering polished contextualized responses they repeated rather than composed; the piece proposes five follow-up question types.
- 4. Structured Interviews: A Practical Guide ✓ opm.gov Structure means the same questions in the same order, a common rating scale and agreement on acceptable answers; questions derive from a job analysis, are written by six or seven subject matter experts using superlatives, and typically cover four to six competencies.
- 5. Employment Tests and Selection Procedures ✓ eeoc.gov A procedure used to make an employment decision must be job-related and consistent with business necessity.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.