Teams
How Do You Get an Interview Panel to Judge AI Use the Same Way?
An interview panel judges a candidate's AI use the same way only when the standard replaces each panelist's private baseline. Put three acts on one page, written by the hiring manager and a senior practitioner in the job's own vocabulary: what the candidate framed before generating anything, one claim they checked and where, and one direction they refused. Calibrate for an hour on two recorded answers, then take a written rating from every panelist before the debrief opens, because the first voice in the room otherwise sets everyone's score.
The takeMost of what a panel calls a disagreement about AI is a disagreement about who gets to decide, wearing a costume. The sheet matters. But the rule that every panelist writes a rating before anyone talks is the part that moves a hire, and it is the first thing a loop drops when the calendar tightens. Anchors are easy to admire and cheap to skip. I suspect three plain anchors plus a strict written-first rule beat a beautiful rubric argued aloud, though nobody in hiring has measured that. The order of the room is the intervention.
Where Olive fits
Open a role and see what the work shows
Three acts on a panel's sheet are three of the six dimensions Olive reads from a real occupational session (problem framing, evidence sourcing and output rejection), each written up by a human reviewer with the moment it rests on attached, rather than judged from a candidate's account of how they usually work. The candidate is granted the same report the employer reads.
Rank your shortlistWhy does the panel split along function?
Because each interviewer is scoring against their own working norms rather than the job's. Daily AI use is ordinary in some occupations and rare in others, so the engineer on the loop hears a candidate describe generating a first draft and thinks nothing of it, while the finance lead hears the same sentence as a shortcut. Both are reading accurately, against different baselines. The standard has to replace the baseline.
The gap is measurable. In Stack Overflow's 2025 developer survey, 51% of professional developers reported using AI tools daily, and 46% said they distrust the accuracy of what those tools produce 1. Across the wider US workforce, Bick, Blandin and Deming found that 23% of employed respondents had used generative AI for work at least once in the previous week, and 9% used it every work day 2. Different instruments and different populations, so the two figures are not a like-for-like comparison. But a panelist drawn from the first world and a panelist drawn from the second walk into the debrief with incompatible ideas of what is normal.
The split is not only about tolerance. It is about which step counts as the checkable one. An engineer wants to know what the candidate ran. A finance lead wants to know what the candidate recomputed. A marketer wants to know which source got opened. Three vocabularies, one question, and none of the three recognizes the others' version of a right answer.
This is also why the fix is not an AI course for the panel. A manager who has never opened the tool can score this fine once someone writes down what a check looks like in their field, and judging AI-assisted work without using AI yourself turns out to be an answer-key problem rather than a training problem. A panel with no sheet at all defaults to whoever argues hardest after the candidate leaves.
Write the standard in the role's vocabulary
One page, written by the hiring manager and one senior person who does the work, not by HR and not from a general definition of good AI use. The sheet describes what someone doing this job would have done: which artifact an assistant produces fluently here, and which check only a practitioner would think to run. Everything on it names an act, not a quality.
That is how federal structured-interview rating scales are built, and the process transfers directly. Subject-matter experts individually write how an employee at each proficiency level would actually answer, the group discusses those drafts, and the examples they agree on become the anchors interviewers rate against 3. Half a day of that work is the reason two strangers can score one answer the same way.
The same act, written three ways:
- Software engineering. Strong: names a test they wrote themselves or an input they tried, and what it turned up. Weak: "it compiled and the tests passed," when the tests were generated too.
- Financial analysis. Strong: a figure recomputed against the filing, and the recommendation that moved because of it. Weak: a model whose logic is described confidently and reconciles to nothing.
- Marketing. Strong: the statistic traced back to the primary source, and what happened when the source turned out to say something narrower. Weak: "the copy was on-brand and I fixed the tone."
Write anchors as behavior, never as adjectives. "Thoughtful about AI" is a rating five people will give five different answers to; "named the claim, opened the source, said what it did not support" is one that two people can agree either happened or didn't. Getting two reviewers to score the same rubric the same way is mostly that substitution, repeated until nothing on the page is an opinion about a person.
What three acts should every panelist score?
Framing, one check, one refusal. What the candidate asked for first, before anything was generated. Which specific claim they went and checked, and where they checked it. What they threw out, and on what grounds. Each is an act with a time and a place attached, so a panelist can score it without sharing the candidate's tooling, and each survives translation into every function on the loop.
1. The first move. Strong: the opening ask states the problem, a constraint, or what would make an answer wrong. Weak: the opening ask is the deliverable. "Write the Q3 forecast" hands over the framing before the framing exists. 2. The check. Strong: a named claim, a location, and a result, as in "the churn figure, I opened the filing, it was annual rather than quarterly." Weak: "I reviewed all of it." A real check has a place, and the candidate can name it. 3. The refusal. Strong: a direction rejected with the reason stated. Weak: everything kept and edited. Editing is tidying. Refusing a framing is judgment, and it is the act panels most often forget to ask about.
Three things stay off the sheet, and saying so out loud in the briefing prevents most of the drift. Volume of AI use is not a virtue. A candidate who decided the model was the wrong instrument for a step and did it by hand has demonstrated the thing being measured. Prompt vocabulary is a month of exposure rather than a year of judgment. And fluency describing AI is precisely what a coached candidate arrives with, which is why interview questions about AI collaboration have to ask for a specific past act rather than a philosophy.
If your loop already runs a values or behavioral round, these three replace a question rather than adding one. Twelve minutes, same three, same order, every candidate, and what good AI use looks like stops being a matter of taste the moment it is three lines on a page.
How do you calibrate the panel in one hour?
Book one hour before the first candidate. Read two recorded answers to the same question aloud, have every panelist score both alone and write the reason, then compare. Where two of them land two levels apart, the anchor is at fault, so rewrite it in the room that day. Interviewer training raises the accuracy of the ratings that follow, and this is the cheapest version of it 3.
Then fix the order of the debrief, because that is where a calibrated panel still loses. OPM's panel procedure has each member individually observe, record and evaluate a candidate's responses, and only then discuss those individual ratings 3. Run it that way: every panelist submits a written rating with its reason before anyone speaks. The debrief resolves disagreements; it does not produce the rating. Without that rule the first person to talk sets the anchor for everyone after them, and the hire is decided by seniority and volume.
Two more mechanics are worth the minute they cost. Use the same interviewers across all candidates for a role wherever the calendar allows, which is the standing recommendation for panel consistency 3. And seed the calibration hour with a genuine disagreement (the answer your engineer and your finance lead already read differently) rather than with an easy pass and an easy fail, since the anchors only earn their keep in the middle of the scale.
One thing calibration cannot settle by itself is the bar. Two panelists can apply the same anchors perfectly and still disagree about what counts as strong for a first-year hire, so state the level in the briefing: what a junior working with AI is being compared against is a different question from whether the acts happened, and it belongs in the competency framework rather than the interview sheet.
What can't a briefed panel tell you?
Whether the candidate would actually do any of it. Every answer on the sheet is a description of a check, given by someone with an obvious interest in the description, and self-report about your own AI work is unreliable well outside an interview. Sixteen experienced developers in a randomized trial took 19% longer to finish issues with AI tools allowed, and still believed afterwards that the tools had sped them up by about 20% 4.
Two closures are cheap. Extend one round by fifteen minutes and hand the candidate a short piece of AI-assisted work with one error planted in it (your error, in your field, so no amount of preparation reaches it), then score what they do rather than what they say. Or move the question to a work sample, where an assistant is available and the confident answer is wrong in a way only checking reveals.
Keep the paperwork either way. An interview used as a basis for a hiring decision is a selection procedure under the Uniform Guidelines, which reach "informal or casual interviews" by name 5, and a procedure that screens out a protected group has to be shown job-related and consistent with business necessity 6. "The panel felt she wasn't AI-native" cannot be shown. A dated sheet, three anchors, and one written reason per candidate can, and it costs nothing extra once the panel is already writing reasons before the debrief.
The record has a second use, which is the one panels underrate. Whatever a panelist writes about a candidate's AI use is a sentence someone may eventually have to read back to that candidate or to a lawyer, so it should be worth reading in both rooms, which is the same standard that governs what a candidate-facing report can safely say.
Common questions
What if two interviewers still disagree after calibration?
Look at the anchor before the interviewers. Persistent disagreement usually means the level descriptions are adjectives rather than acts, or that the two are scoring different things: one on whether the check happened, the other on whether it was the check they would have run. Ask each to point at the sentence in their notes that produced their rating. If they point at the same sentence and still differ, the anchor needs rewriting that day. If they point at different sentences, the question needs a follow-up that forces the specific act into the answer.
Should candidates be told the panel is judging their AI use?
Yes, in the invitation, along with whether using AI is expected or restricted for the exercise. An unstated rule gets guessed at, and the guessing measures interview coaching rather than judgment. Some candidates will hide ordinary tool use they think you disapprove of, and others will perform enthusiasm they don't have. One sentence is enough: this round covers how you work with AI, and using it is expected. It also removes the most common source of panel disagreement, which is two interviewers assuming opposite defaults about what the candidate was allowed to do.
Who should write the standard, HR or the hiring manager?
The hiring manager plus one senior person who does the work, because the anchors are field-specific and nobody else can write them. HR owns what surrounds them: that every candidate for the role meets the same three questions, that ratings are recorded before the debrief, that the sheet is dated and kept, and that the same panel runs the loop where the calendar allows. Split it the other way and you get a document that is procedurally sound and unscoreable, which is how most AI-use rubrics die.
Does a panelist need to use AI themselves to score this?
No. They need the answer key. All three acts are things the panelist already evaluates in their own discipline: did this person understand the problem before producing output, did they check the claim that mattered, did they refuse anything. An interviewer who has never opened an assistant can score a candidate who names the figure they recomputed and what it changed. The panelist who cannot score it is the one working without written anchors, whatever their own tool habits are.
How many people should be on the panel?
Two or three interviewers is the standard recommendation, with the same people used across all candidates for the role wherever scheduling permits 3. Past three, the marginal opinion adds calendar cost and debrief noise rather than accuracy, and the reason for a panel is documentation and independent observation rather than headcount. If more people want input, give them the written ratings to read instead of a seat in the room.
How often does the standard need rewriting?
When the anchors stop discriminating. If nearly every candidate lands at the same level, the sheet has become a formality and the anchors are describing a bar the market has already cleared. Re-read it whenever a role reopens, and rewrite the examples when the work itself changes: a new tool in the team's daily path, or a check that used to be manual and now isn't. Keep the old version with its date rather than overwriting it, so a decision made last quarter can still be read against the standard that was actually in force.
References
- 1. 2025 Stack Overflow Developer Survey: AI survey.stackoverflow.co 51% of professional developers use AI tools daily; 46% distrust the accuracy of AI tool output.
- 2. The Rapid Adoption of Generative AI (NBER Working Paper 32966) nber.org 23% of employed respondents used generative AI for work at least once in the previous week; 9% used it every work day.
- 3. Structured Interviews: A Practical Guide opm.gov Subject-matter experts write behavioral examples that anchor each rating level; interviewer training increases interview accuracy; panel members individually observe, record and evaluate before discussing ratings; two or three interviewers, the same ones across candidates where feasible.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity metr.org Randomized trial of 16 experienced developers: 19% slower with AI tools allowed, while participants estimated a 20% speed-up.
- 5. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.16 (Definitions) ecfr.gov The definition of a selection procedure reaches "informal or casual interviews and unscored application forms."
- 6. Employment Tests and Selection Procedures eeoc.gov A selection procedure with disparate impact must be shown job-related and consistent with business necessity.
6 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.