Screening

How Do You Reference-Check for AI Judgment When the Reference Says 'They Were Great'?

To reference-check for AI judgment, ask about one incident, not the person. Open with whether the reference ever watched the candidate work with an assistant, and take a plain no as a real answer. If it's yes: tell me about a time an AI tool gave them something confidently wrong, and what did they do in the next hour. Grade the reply on three things: an act done outside the chat, a refusal with a reason attached, and work the candidate kept rather than handed over. Same questions, every reference.

The takeThe vague 'they were great' isn't the reference being careful. It's what you get for asking about a whole person when the thing you want is one week. And the people who can answer the narrow version are getting rarer: assistant work happens in a conversation the manager never sat in, so the reference with something to report is drifting from the boss who approved the output to the peer who reviewed it. Nobody has measured that shift yet. If it holds, seniority on a reference list stops being weight, and a reference check becomes corroboration rather than evidence.

Where Olive fits

Open a role and see what the work shows

A reference can tell you what somebody remembers about a decision; it cannot put the decision in front of you. Olive is an employer-purchased assignment built for the occupation, 40 to 60 minutes with an AI assistant, returned as six findings a human reviewer writes, each anchored to the moment in the session it came from, and the candidate is granted the same report.

Rank your shortlist

Why does every reference say 'they were great'?

Because the question had no wrong answer, and the person answering had a reason to stay vague. Many employers refuse to give negative information about a former employee for fear of a defamation suit 1, and where a policy bars disclosure, most will still provide at least start and end dates and position titles 2. The generic endorsement is not evasion. It is what an open question about a whole person produces.

Your reference may never have watched the candidate work with an assistant at all. In NBER surveys of US workers fielded in August and November 2024, 23% of employed respondents had used generative AI for work in the previous week and 9% used it every workday 3. A manager two jobs back may have nothing to report because there was nothing to see, and that is worth ruling out before you push harder.

So open with the prior question (did you see them use AI tools on this work, and on what), and take a plain no as a real answer. If it is no, the call is still worth twenty minutes on everything else, and the AI question moves to a round where somebody can watch it happen. If it is yes, you are one narrow question away from something specific. The resume claim itself is a different check with its own method: how to verify 'AI-proficient' on a resume.

What question gets a specific answer?

One that names an event, a time box, and a decision. "Tell me about a time an AI tool gave them something confidently wrong" works where "how are they with AI" fails, because the reference either has that memory or does not, and both replies tell you something. Then ask the question carrying the weight: what did they do in the next hour?

The shape is three parts, and the third is where an answer stops being a character reference.

  • The event. Something with an outcome the reference had to deal with: a rollback, a correction sent to a client, a figure pulled out of a board deck. If the reference has to invent one, you asked about a person instead of about a week.
  • The act. What the candidate did next, in verbs. Opened the source. Reran the query. Rewrote it by hand. Called the one person who would know. An act performed outside the conversation with the assistant is the signal; a better-sounding explanation is not one.
  • The consequence. What changed: a number, a recommendation, a stated limit, a thing that did not ship. An answer that ends without a consequence is usually a reconstruction rather than a memory.

Keep the questions open-ended and about behavior the reference is likely to have observed 2, and build the set from what the job actually requires 1. Two follow-ups do most of the remaining work: how did they know it was wrong, and what did they check. Follow-up questions that expose real understanding behave the same way here as in an interview: the first answer is prepared, the second is not.

Ask a different question in each field

The question that gets a real answer changes with the work. Ask an engineering manager about a rollback, a finance lead about a number that turned out wrong, a consulting principal about a claim that did not survive the client. Each names an event that field genuinely has, which is what stops a reference reaching back for 'they were great.'

The referenceAsk thisA real answer contains
Engineering manager"Tell me about a change of theirs that got rolled back. How much of it had an assistant written, and what did they do the next morning?"An act outside the editor: a test written, a log read, the diff walked line by line before the second attempt
Finance lead"Tell me about a number of theirs that turned out to be wrong. Where did it come from, and how did it get caught?"The figure re-derived by hand or against the source document, and a limit stated in the next memo
Consulting principal"Tell me about a claim of theirs a client pushed back on. What happened to the claim?"The source opened in the room, and the recommendation changed rather than the wording
Data or analytics lead"Tell me about a result an assistant explained convincingly and got wrong. Who noticed, and how?"A check against the data itself (a recount, a null check, a query rerun), not a more fluent explanation
Product lead"Tell me about a spec where the requirement was genuinely vague. What did they do with the vagueness?"The question taken back to a person, rather than an assumption filled in by the assistant and shipped as fact
Marketing lead"Tell me about a statistic that made it into a draft and then came out. Why did it come out?"The source opened and found not to say it, before the piece ran

Two rules hold across all six. Ask about work the reference personally handled, not work they heard about, because secondhand praise is where the vagueness starts. And do not ask how much AI the candidate used. Volume answers nothing, and someone who decided the model was the wrong instrument for a step and did it by hand has demonstrated exactly the thing being asked about. What good AI use looks like is a short list, and frequency is not on it.

When a reference stalls, narrow again instead of repeating yourself. "Was there a time it was wrong and nobody caught it until later" gets an answer more often than the same question asked louder, because it hands the reference one specific week to look at. The same probe run on the candidate directly is the natural next round.

How do you grade the answer?

A real answer has a shape, and it is worth settling before you dial. Three things: an act outside the chat window, a refusal with a reason attached, and a piece of work the candidate kept rather than handed over. An answer carrying all three is evidence. An answer carrying none is a story about a pleasant colleague, which you already had.

  • Strong. "The model told him the vendor contract auto-renewed. He pulled the signed PDF, found a ninety-day window instead, and we cancelled two weeks before it closed." An act, a source, a consequence.
  • Thin. "She's careful, she always double-checks things." A disposition with no event under it. Ask once more for the last time it happened; if nothing comes back, record it as unanswered rather than as a negative.
  • Nothing. "He used it for everything and shipped fast." Speed with no refusal in it is the answer you were looking for, and it is not a good one.

Ask the same questions of every reference for every finalist, and write the answers into the same form. That is OPM's own recipe for structure, and structure is what makes a reference check worth doing at all: questions built from the job, an identical set for each contact, and standardized recording 1. Keep the questions job-related too, because a selection procedure that screens out a protected group has to be job-related and consistent with business necessity 4.

Budget the time honestly. A structured reference call runs about twenty minutes, and a minimum of three contacts is recommended 1, so each finalist costs roughly an hour of somebody senior. That is affordable for two finalists and absurd for ten, which is an argument for settling who did the thinking before you start dialing.

One caution. Some managers narrate other people's work badly and some narrate it beautifully, and neither fact is about the candidate. Hold the answer to whether it names an act, never to how well it was told.

Where does a reference check stop?

At secondhand memory. The best answer you will get is a manager recalling a moment they half-watched months ago, filtered through whether they liked the person. Structure raises what a reference check is worth 1, and it still cannot show you the candidate deciding, only somebody's account of a decision, with no record left behind it.

That ceiling bites harder on AI judgment than on most competencies, because the behavior happens inside a conversation nobody else was in. The manager saw the pull request. They did not see the four prompts before it, what came back, which paragraph got refused, or whether anything was checked against the world before the branch went up. The record that would settle it belonged to the candidate and their assistant, and nobody kept it.

So use the reference check for what it does well (corroborating that an incident happened and that this person was the one who caught it), and put the remaining question in front of work rather than in front of another conversation. For a small team that usually means one sample, graded against something written down beforehand; the smallest defensible process is a screen, a sample, and a decision. It is also the cheapest insurance against the pattern where somebody interviews and references beautifully and then stalls in the first quarter.

See a sample report

Common questions

What if the reference never saw them use AI?

Take it as an answer and move on. In NBER surveys of US workers fielded in August and November 2024, 23% of employed respondents had used generative AI for work in the previous week 3, so a manager two jobs back may simply have had nothing to watch. Spend the call on the rest of the job, and record the AI question as unanswered rather than as answered badly. It moves to the round where somebody can see the work happen.

Can you ask a reference about a candidate's AI use?

Yes, if you ask about observed work rather than about the person. Tie every question to something the job requires, ask the identical set of each reference for each finalist, and write the answers down the same way 1. A selection procedure that screens out a protected group has to be job-related and consistent with business necessity 4. What that rules out is drift: a question about how an assistant's output got checked is fine, a question about how somebody spends their evenings is not.

What if the reference will only confirm dates and title?

It happens, and it says nothing about the candidate. Where a policy bars disclosure, most former employers will still provide at least start and end dates and position titles 2. Thank them, ask the candidate for a second reference who sat closer to the work (a tech lead, a client-side counterpart, a peer analyst), and put the incident question there. If nothing shakes loose, the reference check has told you what it can, and the open question belongs to a work sample.

How many references does this take?

Three contacts at roughly twenty minutes each is the standard recommendation 1, and it is enough. One reference gives you an opinion; three give you a pattern, which is the only thing worth acting on. Ask two of them the field-specific incident question and the third about something else entirely, so a rehearsed story has somewhere to come apart. Do not add a fourth call to break a tie between the first three. The tie is the finding.

Does this work for a candidate with no work references?

Partly. A professor, a client, an open-source maintainer or a bootcamp instructor can all answer the incident question if they watched the work happen. What they usually cannot answer is the consequence half (what changed downstream), because they were not there for it. Weight the answer accordingly, and do not let a thin reference set stand in for a work sample. Early-career hiring is exactly the case where the sample carries almost all the load.

How does Olive relate to a reference check?

It answers the part a reference cannot. Olive is an employer-purchased assessment: the candidate spends 40 to 60 minutes on a task built for their occupation with an AI assistant available, and a human reviewer writes six findings, each attached to a timestamped moment in the session. There is no composite number and no hiring recommendation; the report is an input to your decision. The candidate is granted the identical document, free. Ten attempts a month are included at no cost, so a shortlist of three is a pilot rather than a purchase.

References

  1. 1. Reference Checking (Assessment & Selection: Other Assessment Methods) U.S. Office of Personnel Management, 2008. opm.gov Adding structure greatly enhances the validity of reference checking: questions based on job analysis, the same set asked each time, standardized recording. Structured phone checks run about 20 minutes with a minimum of three contacts. Many employers refuse to give negative information for fear of a defamation suit.
  2. 2. Reference Checking Guide U.S. Office of Personnel Management, 2008. opm.gov Questions should be open-ended and based on behavior the reference is likely to have observed; where policy bars disclosure, at a minimum most will provide start and end dates and position titles.
  3. 3. The Rapid Adoption of Generative AI National Bureau of Economic Research (Bick, Blandin and Deming), 2024. nber.org Combined results of nationally representative surveys fielded in August and November 2024: 23 percent of employed respondents had used generative AI for work in the previous week; 9 percent used it every workday.
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A selection procedure that screens out members of a protected group must be job-related and consistent with business necessity under Title VII.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.