Interviewing

Give Each Panel Seat One AI Question Nobody Else Asks

Four interviewers on a panel will each reach for the same AI question unless the seats are split by stage of the work rather than by competency. One interviewer asks what the candidate chose to hand to a model and why. One asks how they checked the output. One asks what they did when it came back wrong against a deadline. One asks what they refuse to delegate. Each seat writes evidence for its own question only, and the debrief reads four parts of the job in sequence.

The takeRedundancy in a panel used to be free insurance: four readings of one answer, averaged. That trade stopped paying the moment the answer became preparable. Asking the same AI question four times now measures how well the candidate rehearsed it, four times over, and produces four confident debrief opinions about the same ninety seconds of speech. A panel is the most expensive hour a hiring process buys. Spend it on coverage.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing a check they once ran; it cannot capture them running one. Olive puts the work in front of them instead: a role-grounded assignment with an assistant that will overreach, read afterwards by a person who writes what happened and when.

Rank your shortlist

Why does the whole panel end up asking the same question?

Because the panel was split by competency, and AI is not one. Give four interviewers a list that says technical depth, collaboration, ownership and culture, then add one new topic nobody owns, and all four reach for it. Each gets a version of how do you use AI in your work, each hears the same rehearsed ninety seconds, and the debrief turns into an argument about tone.

Redundancy used to be the point. Two interviewers hearing one answer and rating it separately is a reliability check, and it is cheap. It stopped being a check when the answer became something a candidate can draft, refine and practise the night before, because four samples of a rehearsed paragraph are still one sample of the person.

The selection literature is clear about where combining evidence pays. A composite of genuinely different predictors reaches about .61, an estimate that assumes those predictors are combined mechanically with sensible weights, which is not what a team gets from stacking four unscored conversations 1. A panel is one method however its seats are arranged, so no split of it buys that number. What a split buys is coverage: the same hour comes back carrying four parts of the job.

The second cost is the one nobody budgets for. A freeform conversation can push out information the process already has. Undergraduates who interviewed a classmate predicted that classmate's GPA at r = .31, worse than the r = .65 they managed from prior GPA alone, and in a follow-up study 96 of 169 participants chose an interview in which the answers were generated at random over conducting no interview at all 2. Students predicting grades are not managers predicting job performance, and the size of that gap does not transfer. The mechanism does: people build a coherent story out of whatever they are handed, including noise.

Give each seat one stage of the work

Four stages, four seats, in the order the work happens. Seat one takes the delegation decision: what went to the model and what stayed. Seat two takes verification: how the output got checked, and against what. Seat three takes recovery: the day it came back wrong with the deadline already gone. Seat four takes the refusal: the piece this person will not hand over, and the reason.

1. Which parts of that project did you keep, and which did you hand over? Follow up on the boundary itself. A useful answer names a specific task and a reason drawn from the work. Speed on its own is the answer to push on. 2. How did you know the output was right? The strong answers name a source outside the model: a filing, a test suite, a payer policy, a person who would know. The weak ones describe reading it over. 3. Tell me about a time it was confidently wrong and the deadline was that afternoon. This is the seat that finds out whether verification is a habit or a slogan. In one field experiment, on a task deliberately chosen to sit outside the model's capability, consultants working with GPT-4 reached the correct answer 60% and 70% of the time against 84.5% in the control group, and the condition given a prompt-engineering overview did worse than the one given none 3. One task, one sample, a 2023 model. What it illustrates is the shape of the failure: nobody could tell which side of the line the task was on. 4. What will you not put through a model, at this company, whatever the deadline? Listen for a line drawn around consequence rather than around policy.

Seat three and seat four are the hardest to rehearse, because both require an actual episode with a cost in it. If the loop only has three interviewers, cut seat one: what somebody delegates is also the easiest thing to read off a work sample later.

Write the scorecard before the loop opens

Each seat needs three things on paper before anyone meets the candidate: its own question, the evidence it is listening for, and what an acceptable answer contains. That third item is the one loops skip. A federal guide to structured interviewing defines the format as the same questions asked in the same order, evaluated on a common rating scale, with interviewers who agreed in advance what an acceptable answer looks like 4.

Write the acceptable answer as a description of content, never as a rating. For the verification seat: names at least one source outside the model, and says how the check would have caught a specific error. For the recovery seat: describes what they did in the hours after, including what they told whoever was waiting. An interviewer who has that sentence in front of them stops grading confidence.

One rule holds the split together: each seat records evidence only for its own question. If the delegation seat hears a good verification story, it goes in the notes as context and never in the rating. Without that rule the split collapses, because every interviewer wants to answer the whole question. Getting a panel to judge AI use the same way is the briefing problem underneath this one, and running a structured interview about AI use is the mechanics of asking a single question well.

Set the level bar in the same document. An acceptable verification answer from a new graduate looks different from one at staff level, and the three bars at junior, mid and senior belong in the scorecard. Left to the debrief, they turn into a negotiation.

What does the debrief do with four different answers?

It assembles them rather than averaging them. Nobody rates the candidate on AI. Each seat reports what its own question produced, in the order the work happens, and the room reads the four together: what this person hands over, how they check it, what they do when the check fails, and where they stop. Treat disagreement between seats as information about the candidate.

A candidate who delegates cleanly and verifies thinly is a real profile, and a coachable one. A candidate who verifies obsessively and delegates nothing needs a different conversation. Averaging four ratings would have erased both, and reading the seats in sequence puts both in front of the room.

Hold one thing back from the room until every seat has read out. A summary delivered first hands every other seat a frame to read its own notes through. Read in sequence, then discuss.

The honest limit: none of this converts a rehearsed paragraph into a demonstration. Four questions get four accounts, and an account is what a conversation can produce. When the decision genuinely turns on catching a wrong output, the loop needs a piece of work in it, and scoring an answer the candidate produced with AI is the narrower version of that problem. Panels that also want the hiring managers on them to read AI-assisted work consistently should start with the managers who do not use these tools themselves.

See how it works

Common questions

What if the panel only has three interviewers?

Keep verification, recovery and refusal, and drop the delegation seat. What a candidate chooses to hand over is the stage most visible elsewhere: a work sample, a portfolio walkthrough or a reference conversation all surface it. Verification and recovery are the two that need a live question and a follow-up, because both turn on a specific episode the candidate has to reconstruct. Refusal is the cheapest of the four to ask and the most revealing about judgment, so it stays even in a short loop.

Should every interviewer ask an AI question at all?

Only if the role genuinely involves the work. For a job where a model touches the daily task, four seats of coverage is proportionate. For a role where it does not, one seat is enough and the other three should spend the time on something the job requires. Adding an AI question to every seat because the topic feels current is how a loop ends up assessing a skill nobody will use, at the cost of the ones the role actually needs.

How do I stop interviewers drifting back to the same question?

Give each seat its question in writing, in the calendar invite, and require notes filed against that question only. Drift comes from an interviewer who does not know what the other seats hold, so publish the whole split to the panel rather than sending each person their own line. A five-minute briefing before the first loop of a new role costs less than one debrief spent untangling four overlapping accounts of the same story.

Does this work for a peer interview seat?

It works well there, and the recovery question is usually the right one to give a peer. A peer can follow up on specifics a manager would not know to ask, and can hear whether the account of that afternoon matches how the work actually runs. Give the peer seat one question, the same acceptable-answer description as everyone else, and the same instruction to record evidence only for that question.

What if two seats hear contradictory answers?

Treat the contradiction as the finding and go back to the candidate rather than resolving it in the room. Someone who describes a rigorous verification habit to one interviewer and cannot name a single check to another has told you something specific, and a short follow-up call usually settles which account holds. Resolving it by vote produces a rating nobody can explain later, which is the worst outcome available if the decision is ever questioned.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300, doi 10.1017/iop.2023.24 (Cambridge University Press), 2023. cambridge.org Supports the claim that a composite of different predictors reaches about .61, and that the estimate assumes mechanically combined predictors rather than several similar conversations.
  2. 2. Belief in the unstructured interview: The persistence of an illusion Judgment and Decision Making, 8(5), 512-520 (Society for Judgment and Decision Making), 2013. sjdm.org Supports the r = .31 versus r = .65 dilution finding and the 96-of-169 result that participants preferred an interview with randomly generated answers to no interview.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the outside-the-frontier result quoted for the recovery seat: 84.5% correct in the control group against 60% and 70% in the two AI conditions.
  4. 4. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three-part definition of a structured interview used for the scorecard: same questions in the same order, a common rating scale, and an agreed acceptable answer.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.