Interviewing
Four Lines Worth Scoring on AI Use, and Two to Delete
An interview scorecard for AI use holds four lines, each a quote or an artifact reference instead of a rating: what the candidate handed to the model and the reason they gave, how they checked the part that mattered, what they caught or missed, and what they would do differently. Take two lines off the form: the overall AI rating, because nobody can say what a 4 means, and anything asking whether the answer sounded AI-generated, since untrained readers guess at chance and detectors misfire on second-language writers.
The takeThe one-to-five column with weights is the part of a scorecard that does not survive being questioned. It looks like measurement because it produces arithmetic, and there is nothing under the arithmetic once somebody asks what the rating rests on. Keep the column for things you can anchor in advance and drop it for this one. A line quoting what the candidate actually said costs no more than the rating did and answers the question the number cannot.
Where Olive fits
Open a role and see what the work shows
Olive applies the same discipline to a full session: six findings written by a person, each stated as demonstrated, partly demonstrated or not demonstrated with the timestamped excerpt it rests on. No composite, and no number standing for a candidate.
Rank your shortlistWhat belongs on the scorecard?
Four lines, all describing something the candidate did or said rather than something the interviewer concluded. What they handed to the model and the reason they gave for it. How they checked the part that mattered. What they caught, or what they walked past. What they would do differently next time. Every line carries a quote or a reference to a specific artifact underneath it.
Written out as they appear on the form:
1. Delegation and the reason. What went to the assistant, what stayed with the person, and the reason in the candidate's own words. 2. The check. The specific thing they did to test the part that mattered, including what they compared it against. 3. The catch. What they found, or the moment they moved past something they should have found. Both are evidence. 4. The revision. What they would do differently, and whether the answer names a change to the process or only to the effort.
The third property of a structured interview is where AI scorecards usually break. The US Office of Personnel Management's practical guide defines one by three things: every candidate is asked the same questions in the same order, every candidate is evaluated on a common rating scale, and the interviewers agree in advance on what an acceptable answer looks like 1. That guide is federal HR practice from 2008 rather than statute, and it predates every question here by more than a decade, but the third property is the one teams skip and the one that decides whether the form collects anything.
The form can only collect what the round made askable, so the script comes first. Saying the AI rule out loud in the first two minutes is what puts a real artifact in the answer.
Write each line as a quote or an artifact reference
Capture the candidate's own words as they say them. A line reading "dropped the third paragraph because the case it cited does not exist" still means the same thing in a debrief next quarter. A line reading "good verification instincts" means whatever the reader brings to it. The two cost the same thirty seconds at the time, and only one of them is still evidence later.
Three mechanics make this practical rather than aspirational. Keep the quote short, six to fifteen words, since a full transcript never gets written and a paraphrase drifts. Name the artifact when there is one, so "the pricing claim in the drafted brief" beats "the example he gave". And write during the answer rather than after the round, because the reconstruction at the end of the day is where the specifics disappear and the impression takes over.
The reason this matters more than the choice of method is visible in the validity evidence. In the 2022 re-analysis of the selection literature, structured interviews came out top ranked at .42, ahead of job knowledge tests at .40, empirically keyed biodata at .38, work samples at .33 and unstructured interviews at .19, with most methods keeping their old rank order but with mean validity estimates reduced by .10 to .20 points 2. Those are corrected correlations with supervisor performance ratings pooled across many jobs rather than accuracy rates, the .42 carries an 80% credibility interval running from .18 to .66, and nothing in the paper concerns AI. What the ranking rewards is administration: same questions, same scale, recorded the same way.
Two interviewers writing quotes will still disagree about what a quote showed, which is a productive disagreement and a solvable one. Writing a rubric for AI use that two reviewers score the same way is the calibration half of this.
Drop the overall AI rating
The composite is the line that cannot be defended and it is usually printed first. Nobody in a debrief can say what separates a 3 from a 4 on AI fluency, so the number ends up recording how much the interviewer enjoyed the conversation. Take the column off the form, keep the four evidence lines, and let the debrief argue about what a specific moment showed.
Weighting makes it worse. Multiplying an unanchored number by 0.3 produces a more precise-looking figure resting on exactly the same absence, and the precision is what makes it hard to challenge internally. The person who wants to overrule it now has to argue with arithmetic instead of with a judgment, which is how a weak signal wins a debrief it should have lost.
What replaces it is a bar rather than a scale. For this role, does the evidence collected clear what the work requires: yes, not yet, or no. Three outcomes, stated against a standard written before the round, with the four evidence lines underneath as the reason. A hiring manager can explain that to a candidate, to a panel and to counsel, and each of those readers gets the same answer.
One exception worth keeping in view. A scale is defensible when the anchors were written first and describe observable answers, which is real work and occasionally worth doing for a role you hire twenty times a year. Scoring an interview answer the candidate produced with AI sets out what those anchors have to contain before a number under them means anything.
Why can't you score whether an answer sounded AI-generated?
Because it is a guess about a person that no instrument supports, and the error does not land evenly. Untrained readers do not separate machine-written from human-written text at better than chance, and the tools built for the job misfire hardest on people writing in a second language. A line asking an interviewer for that judgment collects an impression and files it as evidence, which is the worst thing a form can do.
The human version was measured first. In an ACL 2021 study, non-expert evaluators asked to tell GPT-3 text from human writing across stories, news articles and recipes performed at random chance, and three quick training methods lifted accuracy only to about 55%, not consistently across the three domains 3. That is 2021 output judged by crowdworkers on short passages, so it does not describe an experienced reviewer reading work in their own field. Models have improved since, which pushes unaided judgment down rather than up, though that is an inference and not something the paper measured.
The tooling version is worse in a specific direction. Seven widely used detectors run over 91 human-written TOEFL essays by non-native English speakers produced an average false-positive rate of 61.3%, with 97.8% of those essays flagged by at least one detector, while the same tools were near-perfect on 88 essays written by US eighth-graders 4. Ninety-one essays from one language background, tool versions as they stood in 2023, academic essays rather than interview answers. The number is not a constant. The direction of the error is the durable finding, and it points at exactly the candidates a hiring process is most often asked to treat carefully.
So when an answer feels rehearsed, write down the thing that was actually missing rather than the impression: no artifact named, no consequence described, no source given. Then ask the follow-up, which is the only move that resolves it. Stopping hiring managers rejecting candidates for sounding like AI is the same problem one level up, at the debrief table.
Common questions
Where do the four lines sit on an existing scorecard?
Inside the competency the AI question was attached to, not as a separate section. A standalone AI block invites a standalone AI rating, which is the thing being removed, and it also implies the topic is a category of its own rather than part of how the work gets done. If the question was asked in the technical round, the lines belong under that competency. The form grows by four rows and loses one column, which most teams find is a net simplification.
How do you compare two candidates without a number?
Put the evidence lines side by side against the standard written before the round, one dimension at a time. Comparison by quote is slower for two candidates and considerably faster for a disagreement, because the panel argues about what a specific answer showed rather than about whose 4 was correct. When two candidates genuinely tie on the evidence, the tie is real information: the round did not separate them, and inventing a decimal to break it manufactures a distinction the interview never observed.
What should the scorecard say if the candidate did not use AI at all?
Record what they did instead and how they checked it, under the same four lines. Someone working under a tool ban, in a regulated setting or on air-gapped systems can have strong verification habits and no model anywhere in the story, and the lines about the check, the catch and the revision all still apply. Write down the constraint they were working under so the debrief reads the answer in context rather than as an absence.
Who should fill in the scorecard, the interviewer or a notetaker?
The interviewer, during the round, with a notetaker capturing longer quotes if one is available. Handing the form to somebody else separates the judgment from the person who heard the answer and produces notes with no follow-up behind them. If an assistant is transcribing, treat the transcript as raw material for the quotes rather than as the record itself, and check any consent requirements for recording, which vary by state and by whether the interview is analysed afterwards.
How long should scorecards for AI questions be kept?
As long as the rest of the hiring file for that requisition, under whatever retention rule already covers your interview records. Splitting AI notes into a separate store creates two problems: the notes go missing when someone asks how a decision was reached, and the split itself suggests the material is different in kind, which it is not. One file per candidate, one retention rule, and the evidence lines readable by somebody who was not in the room.
References
- 1. Structured Interviews: A Practical Guide opm.gov Supports the three properties of a structured interview, including agreeing in advance what an acceptable answer contains.
- 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range gwern.net Supports the corrected validity ranking and the point that consistent administration and recording is what the estimates reward.
- 3. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that untrained readers separate machine-written from human-written text at chance, and that brief training barely helps.
- 4. GPT detectors are biased against non-native English writers pmc.ncbi.nlm.nih.gov Supports the claim that the error in judging text authorship lands unevenly on second-language writers, so the line cannot go on a scorecard.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.