Assessment design

How Do You Score an Interview Answer the Candidate Produced With AI?

When a candidate produces an interview answer with AI, rate four acts separately: what they framed before generating, what evidence they demanded, what they refused from the model's output, and what they checked outside the conversation. Polish is the assistant's contribution and earns nothing. Write each anchor in your field's evidence standard: a source opened for a consultant, a failing test for an engineer. A fluent, unverified answer takes the middle band on framing at best and the bottom band elsewhere, which buys a second conversation, not a rejection.

The takeThe four items are the copyable half. They fit on a slide, they look like rigor, and any vendor can ship them. The anchor does not copy, because it is local knowledge: what a competent colleague in your job asks to see before signing. Nobody has measured whether anchored rubrics predict performance better than unanchored ones, and I would not bet a loop on it either way. But an unanchored scale still produces numbers, and a number outlives the doubt that made it. A borrowed rubric hands you that trade without mentioning it.

Where Olive fits

Open a role and see what the work shows

An interview rubric rates a candidate's account of what they opened, recomputed and cut; the anchors only bite when those acts are in a record. Olive runs the work instead (a 40-to-60-minute occupational assignment with an AI assistant available), and a human reviewer writes six separately-evidenced findings, each tied to a timestamped moment rather than to a band on a scale, with the candidate granted the identical report.

Rank your shortlist

What Do You Score When the Model Wrote the Draft?

Score four acts the candidate had to perform themselves: what they framed before generating, what evidence they demanded for the claim the answer rests on, what they refused from the model's output, and what they checked outside the conversation. Rate each separately, on its own anchored scale. Fluency, structure and length are the assistant's contribution, and rating those tells you about the tool.

The reason to split them is not tidiness. Federal selection rules put the ratable material in a specific place: a job analysis should focus on observable work behaviors and, where a behavior is not observable, on the aspects that can be observed and on the work products; a skill has to be operationally defined in terms of observable aspects of work behavior; and a selection procedure resting on inferences about mental processes cannot be supported solely or primarily on content validity 1. So "judgment," "critical thinking" and "AI fluency" are not rubric items. Opening a source is. Recomputing a figure is. Cutting a section is.

The four items also match where the failure actually lands. In the 2025 Stack Overflow developer survey the single most common frustration with AI tools, at 66%, was output that is almost right but not quite; 46% of developers said they distrust the accuracy of what the tools produce against 33% who trust it 2. Almost-right passes every proxy an interviewer used to read (structure, confidence, vocabulary, a clean numbered plan) and fails on the one claim the recommendation hangs from.

And the confident-wrong claim is not rare in professional tools built for exactly this. A Stanford RegLab and HAI study of purpose-built legal research systems found Lexis+ AI and Thomson Reuters' Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time, against vendor claims of hallucination-free retrieval 3. A candidate who takes a cited-looking claim at face value is not being careless by the standards of the marketing; they are being careless by the standards of the job. That distinction is what the rubric has to encode, and it is why the scale gets written before the questions do, which is the ordering argument in Structured Interview for AI Use.

How Do You Anchor "Demanded Evidence" in Your Own Field?

Name what a source is in the work. "Demanded evidence" is empty until it says: for a consultant, a named source opened and the figure traced to the page; for an analyst, a reconciliation that ties to the system of record; for an engineer, a test that fails before the change and passes after. Write the anchor as the act the job would accept, then rate against it.

This is the half most rubrics skip. The federal definition of a structured interview has two clauses, not one: every candidate gets the same predetermined questions in the same order, and every response is evaluated using the same rating scale and standards for acceptable answers 4. Teams ship the scale and skip the standard, then wonder why the numbers do not travel between interviewers. A five-point scale is fine if all five points are labeled with an act; unlabeled points are where taste enters.

DisciplineThe load-bearing claimTop band: demanded evidenceBottom band
Strategy and management consultingA market size, a growth rate, a benchmarkNames the source, opens it, locates the figure, and compares its definition against the question being askedAttaches a firm's name to a number with no page behind it
Financial analysisA figure that feeds the recommendationTies the figure to the filing or the system of record and explains the varianceRestates the number the assistant produced
Software engineeringHow the code behaves in a case nobody wrote downA test that fails before the change and passes after, named in the answerSays the change "was tested" or that the code "looks correct"
Data and analyticsThe direction or the significance of a resultRe-runs the query against the raw table, states row counts and what was excludedAccepts the assistant's reading of the output
Journalism and communicationsAn attributed statementOpens the primary document, or gets a second independent sourceCites a secondary summary of a primary source
Legal operationsA citation, a clause position, a precedentOpens the case or the executed contract and quotes the clauseRepeats a case name the assistant supplied
Underwriting and claimsA fact about the riskReads the survey against the application and names the disagreementWorks from the application alone

The pattern across the rows is one sentence: evidence is whatever a competent colleague in that job would ask to see before signing. Write that sentence for your role, put it in the top band, and the middle and bottom bands mostly write themselves: middle is asked for and not opened, bottom is neither asked for nor opened. Setting the height of that top band for a specific role is its own decision, worked through in How to Set a Defensible AI Proficiency Bar by Role.

Score One Answer: A Filled-In Rubric

Here is one answer, rated. The question went to a senior analyst candidate: top-of-funnel conversion fell from 3.1% to 2.4% over a quarter, walk through the first two days. The reply came back in 600 fluent words, five numbered steps, one benchmark claim and one recommendation. It takes the top band on nothing, the middle on framing, and the bottom on the other three.

The shape below is the one that arrives most often.

> Step 1: segment the drop by channel, device and campaign. Step 2: check for tracking changes; a broken tag explains a 0.7-point move more often than demand does. Step 3: compare against the industry benchmark; B2B SaaS top-of-funnel conversion runs around 2.8%, so 2.4% is below but not alarming. Step 4: interview two account executives. Step 5: ship a landing-page test.

  • Framing: middle band. Top band asks for the decision the answer serves and one condition that would make the answer wrong. This names the diagnostic and never the decision (whether this is a spend question or a product question changes every step), and never a falsifier. It earns the middle band on step 2 alone, which raises an alternative explanation before committing to analysis. That is one real framing act, and it is the only one.
  • Demanded evidence: bottom band. The load-bearing claim is "around 2.8%." Top band asks for the source named, opened, the figure located, and its definition compared against ours. No source appears. Asked in the follow-up, the candidate produced a vendor benchmark report and had not opened it; the report counted free-trial signups, which this funnel does not. Bottom band, and the follow-up is the evidence for the rating.
  • Output rejection: bottom band. Top band asks for at least one direction refused on stated grounds. All five steps survive into the answer. A five-step plan with nothing dropped is the assistant's default reply with the candidate's name on it, and the give-away is that the steps are not ordered by cost: a broken tag takes ten minutes and sits second, behind segmentation.
  • Verification: bottom band. Top band asks for one claim tested against something outside the conversation, where the result changed the answer. The candidate had the funnel export. Nobody recomputed 2.4% from it, so the number the entire reply rests on was never confirmed against the file sitting next to it.

One middle and three bottom is not a rejection. It is a scored, specific second conversation: open the benchmark with them, hand them the export, ask which step they would cut. What the rubric bought is that the next interviewer starts from four rated items rather than from "strong communicator." Expect an empty top band often, and resist grading the batch on a curve, which is the same discipline that keeps a polished submission from carrying a weak decision through a whole rubric, covered in Grading Polished Take-Homes.

Why Do Two Interviewers Score the Same Answer Differently?

Because the anchor is written in adjectives. "Strong evidence," "good rigor" and "shows judgment" each mean whatever the reader already believed, so two interviewers rating one reply disagree about the words rather than about the answer. Rewrite every band as an act with a trace: what was opened, what was recomputed, what was cut. Then double-score the first five candidates and compare item by item.

Two habits do most of the work after that.

  • Rate the trace, never the account. A candidate's description of their own process is not evidence of it. METR randomized 246 real issues from experienced developers' own repositories to allow or forbid AI tools: with the tools allowed the work took 19% longer, and the same developers, having finished, still believed the tools had sped them up by 20% 5. People are unreliable narrators of their own AI-assisted work, and that holds for the sincere ones.
  • Rate item by item, across candidates. Score every candidate's framing, then every candidate's demanded evidence, then rejection, then verification. Reading one reply end to end is how a fluent opening paragraph sets the tone for four ratings that were supposed to be independent.

Where the two raters disagree in the first five, the anchor is ambiguous rather than the candidate borderline. Fix the wording, log what changed, and hold the fixed version for the rest of the loop. The disagreement is cheap information at candidate three and expensive at candidate thirty; the mechanics of keeping a collaboration rubric consistent across a panel are in Can Two Interviewers Score AI Collaboration the Same Way?.

There is one bias specific to this rubric worth naming. A survey of 319 knowledge workers across 936 first-hand examples found that higher confidence in the AI tool tracked with less critical thinking, while higher confidence in one's own ability tracked with more; the effort that remains shifts toward verifying information, integrating the response and stewarding the task 6. An interviewer who trusts these tools reads an unsourced benchmark as fine. One who distrusts them reads it as disqualifying. Neither is rating the candidate, which is what a written standard for an acceptable answer is for.

What This Rubric Cannot See

An interview answer is an account of work, not the work. The rubric rates what a candidate says they opened, recomputed and cut, and someone who did none of it can describe all of it fluently. Follow-up questions shrink the gap (borrowed specifics survive one round and rarely two), but they do not close it. Only a record of the acts closes it.

Two ways to get closer, both cheap. Ask for the artifact rather than the summary: the source file instead of the exported chart, the transcript of the exchange instead of the conclusion, the commit instead of the description of the commit. And put the material in front of the candidate live, so the check has to happen in the room: hand over the funnel export and ask for the number, hand over the benchmark PDF and ask whether it counts what we count. Both convert a claim about an act into an act. The question design that makes those follow-ups land is worked through in Follow-Up Questions That Expose Understanding.

It is also worth being clear about what a rubric of this kind does buy, because it is not a prediction of performance. It buys comparability across candidates who were asked the same questions and rated on the same standards for acceptable answers 4. It buys a defensible record, because the rating rests on observable behavior and work products rather than on an inference about someone's mental processes 1. And it buys a specific next step, which an overall impression never does.

The last limit is the one to say out loud in the debrief. A bottom band on verification means the candidate did not verify anything in this answer, under these conditions, with this much time. It does not mean they cannot. Say that plainly when the panel meets, because a four-item rubric read as a verdict on a person is worse than no rubric at all: it lends arithmetic to a judgment that has not earned it.

See what gets scored

Common questions

Should a candidate lose points for using AI at all?

No. If the job uses the tool, an answer produced with it is a work sample rather than a violation, and volume of use is not the thing being rated. What earns a low band is an unsourced claim, an unexamined plan, or a number that was never checked against the file next to it. And a candidate who decided the tool was the wrong instrument for a step and did that part by hand can score at the top on all four items. Rate the acts, not the tool count.

How many bands should the rating scale have?

At least three, and every band labeled with an act. Three is enough for most loops: the evidence was demanded and opened, it was asked for and not opened, or it was neither. Five points work too, if all five carry a written description of what an answer at that point contains. The failure is not the number of points; it is unlabeled points, which is where one interviewer's four meets another's two and nobody can say why.

What score does a polished but unverified answer get?

Middle band on framing at best, bottom on demanded evidence, output rejection and verification. Polish is real work product and it is not zero, but it carries no information about the three items that need a trace. In practice that answer is not a rejection. It is a scored, specific second conversation where the candidate opens the source in front of you. Candidates who move two bands in that conversation are common, which is itself worth knowing before a decision.

Can one rubric cover every role we hire for?

The four items travel; the anchors do not. Framing, demanded evidence, output rejection and verification describe capable AI work in any occupation, so the rubric's skeleton can be shared across a company. The top-band text cannot be, because a reconciliation that ties, a failing test and a primary document opened are different acts with different costs. Write the anchor per role with someone who does the job, and re-use everything above it.

Should candidates be told the rubric in advance?

Give it to them in the invitation, in four lines. Undisclosed criteria add noise: half the candidates guess that polish is being rated and spend their preparation there, so the ratings partly measure the guess. Naming the four items also changes behavior in the direction you want: candidates arrive having opened their sources. That is not gaming it. A candidate who opened the source because you said you would ask has done the act the item rates.

References

  1. 1. 29 CFR 1607.14 - Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) U.S. Equal Employment Opportunity Commission, via eCFR, 1978. ecfr.gov Content validity requires a job analysis focused on observable work behaviors and work products, skills operationally defined in terms of observable aspects of work behavior, and states that a selection procedure based on inferences about mental processes cannot be supported solely or primarily on content validity.
  2. 2. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  3. 3. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Magesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford RegLab and Institute for Human-Centered AI (arXiv:2405.20362), 2024. arxiv.org Purpose-built legal research tools from LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research, Ask Practical Law AI) each hallucinate between 17% and 33% of the time, against vendor claims of hallucination-free retrieval.
  4. 4. Assessment and Selection: Structured Interviews U.S. Office of Personnel Management, 2024. opm.gov In a structured interview all candidates are asked the same predetermined questions in the same order, and all responses are evaluated using the same rating scale and standards for acceptable answers.
  5. 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 246 real issues from experienced developers' own repositories, randomized to allow or forbid AI tools: the work took 19% longer with the tools allowed, and participants still believed afterwards that the tools had sped them up by 20%.
  6. 6. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks and Wilson (Microsoft Research and Carnegie Mellon), ACM CHI 2025, 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in the tool is associated with less critical thinking, higher confidence in one's own ability with more, and remaining effort shifts toward information verification, response integration and task stewardship.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.