Assessment design

Ask a Reference for One Episode, Not an Overall Impression

A reference call produces evidence instead of a character summary when it asks for one bounded episode with a date on it, the same way for every reference on the requisition. "Tell me about the last thing they shipped that you had to send back" gets answered even by a former manager under instructions to confirm employment only, because recounting what happened is not a rating and carries less perceived exposure. Rating questions ask the reference to do your evaluating, and that is the request they are trained to decline.

The takeDelete the rehire question. Published reference-check question lists recommend it more than anything else, and it is also the item most likely to be answered by an HR policy rather than by the person you called, so it reliably spends your best slot on a sentence somebody else wrote. Keep it only where there is a specific reason: a role a return is genuinely plausible for, or a reference who has already volunteered that they would. Otherwise it is twenty seconds you could have spent on something a person is willing to describe in detail.

Where Olive fits

Open a role and see what the work shows

Writing the question set is the easy half; the answer key and the evidence trail behind each finding are the hard one. Olive ships authored cases grounded in twelve occupations and returns six separately evidenced findings, each anchored to a timestamped moment in the session.

Rank your shortlist

What turns a reference call into evidence?

An episode. A question naming a bounded piece of work with a rough date on it asks the reference to recount something, and recounting is a different act from evaluating: it does not feel like a judgment the former employer could be held to, and it produces detail a second reader can weigh. Rating questions ask the reference to do the evaluating, which is the one thing a reference does least reliably.

The published question lists are almost entirely rating questions. Strengths and weaknesses, would you rehire, how would you rate them on a scale of one to ten. The competing genre answers a different worry, what an employer may lawfully ask, and stops at dates, title and eligibility for rehire. Neither genre contains a question shaped like an episode, which is why so many calls end with a warm sentence and nothing to write down.

The policy sentence is the predictable result. The researchers who reprinted the old reference-check validity figure in 2016 said in the same paper that it may no longer hold, because in the United States many former employers now release only dates of employment and job titles 1. That is the environment your script has to work in. A question that asks for a judgment runs straight into that policy. Ask instead what happened on a Tuesday in March, and a former manager will usually answer, because describing a project does not feel like passing judgment on a person.

Structure is the other half, and it is the half that makes two candidates comparable. Federal guidance is plain about the shape: a structured interview format, conducted by phone, with written requests producing low response rates and less useful information 2. Same questions, same requisition, same number of calls for each finalist, which is the same argument as running fewer reference calls and structuring them. Without that, you have two conversations that happened to be about two people.

Write three episode questions for the role you have open

Start from the two or three things this role will actually be judged on in its first year, then write one question per thing that names a moment instead of a trait. The test of a good one is whether a reference who has not thought about the candidate in eighteen months could still answer it with a story.

Three that travel well, with the version they replace:

1. "What was the last piece of work of theirs you sent back, and what did you send it back for?" Replaces: what are their weaknesses. It gets a specific answer because it presupposes something ordinary and accuses nobody of anything, and "honestly, I never sent anything back" is itself informative. 2. "Tell me about a call they made without you in the room that you would have made differently." Replaces: how are their judgment and decision-making. It surfaces the boundary the person actually drew around their own authority. 3. "What were they working on in their last two months, and how did that land?" Replaces: why did they leave. It gets you the end of the tenure without asking a question the reference has been told to route to HR.

Then ask one follow-up per answer, and make it the same follow-up: what happened next. Most of the information in a reference call sits one layer under the first answer, which is the same reason a second question exposes what a first answer only claimed in an interview. Stop at three plus follow-ups. A twenty-minute call with three questions produces more than a forty-minute call with twelve.

Ask how the work actually got made

Ask what the person produced themselves and what they assembled, framed as a description of process rather than as a suspicion. "Walk me through how that analysis came together, start to finish" is a question a reference will answer at length. A former colleague who watched the work names the parts that were drafted, the parts that were reused and the parts the candidate built from nothing, because that is simply what happened.

The standard question lists do not carry it, and it is now the question the rest of your process cannot answer. A resume, a cover letter and a take-home can all be produced with an assistant, so the finished artifact does not settle who did the thinking. A person who sat beside the candidate for two years does know, and will say so in ordinary language if the question does not sound like an accusation.

The wording matters more here than anywhere else in the script:

  • Ask: "How did that piece of work come together?" and "What part of it did they do that nobody would guess from reading it?"
  • Do not ask: "Did they use AI for that?" It is a yes-or-no about a tool, the answer is yes often enough to carry no information, and it teaches the reference that you are looking for something to hold against the candidate.

What comes back is a description of how the work got made, and it carries no more weight than the rest of the evidence you have. When two finalists have produced work of similar quality, the reference call is one of the places to look for who did the thinking behind the output, and it is the weakest of them because it is recollection at second hand.

Keep the words, not the rating

Keep the sentences the reference actually said and drop the five-point scale you were going to convert them into. A rating hides what was said, and it is the step where a call stops being comparable to anything: two people who both wrote 4 out of 5 may have heard opposite stories. A quote survives a debrief with its context attached.

There is direct evidence that the content of an interviewer's notes tracks what happens after the hire. Text-mining post-interview notes on 7,650 hires at one large Chinese technology company, Liu and colleagues found that the number of job-related capabilities an interviewer named in the notes was positively related to later performance and promotions and negatively related to turnover, with a one standard deviation rise in how well the notes matched the job analysis corresponding to roughly a 2 percent rise in performance 3. That is one firm, one country, only people who were hired, and a small effect, and the study never tested the notes against the interviewer's own rating. It still points the same way as everything else here: name the capability, quote the evidence, skip the number.

Free text carries one risk worth stating plainly. Narrative references are not neutral just because they are detailed. Coding 624 authentic letters of recommendation for academic posts, Madera and colleagues found 54% of letters written for women contained at least one doubt raiser against 51% of letters for men, with a wider gap on letters carrying two or more 4. Academic letters are not employment references, the gap in those headline percentages is small, and the paper's own test runs on coded doubt-raiser scores rather than on those two percentages. The mechanism is still worth carrying: what a writer chooses to volunteer differs by group before any reader gets to it. Asking every reference the same three questions is a partial answer to that, because it puts the same request in front of every writer.

See what gets scored

Common questions

Should you ask whether they would rehire the candidate?

Skip it in most cases. It is the most common item on published lists and the one most likely to come back as policy, because eligibility for rehire is the field large employers route to HR with the answer already decided. When it does get answered, a yes tells you almost nothing and a no tells you something you then cannot ask about. Spend the slot on an episode question instead. Keep the rehire question for roles where a return really is plausible, or when a reference raises it themselves.

How long should a reference call take?

About twenty minutes for three questions and their follow-ups. Longer calls do not produce more evidence; they produce more conversation. Book twenty-five minutes, spend the first two saying what the role is and what you are going to ask, and leave the last two for anything the reference wants to add unprompted. That last stretch is where the most useful sentence often arrives, because it is the only unstructured part of the call.

Should you send the questions to the reference in advance?

Send the shape, not the script. Tell them you will ask about two or three specific projects and roughly which period, so they arrive with something to say. Sending the exact wording invites a prepared answer, and prepared answers now often mean drafted answers, which is the failure mode written references already have. A one-line email naming the role and the window is enough.

What if the reference is a peer or a friend rather than a manager?

Ask them the same questions and read the answers for what they saw. A peer who sat next to the candidate often has better detail on how work actually got made than a manager who saw the finished version, which makes them more useful on process questions and less useful on scope and impact. What matters is that you know which you are talking to and note it in the file. A friend who did not work with the candidate cannot answer episode questions at all, and that is worth recording as the outcome of the call.

Can you ask a reference about a specific concern from the interview?

Yes, if you turn it into an episode question and ask it of every reference for that role. "They seemed to struggle when the requirements changed mid-project, does that match what you saw" invites agreement with your premise. "Tell me about a project of theirs where the requirements changed partway through" gets you the same ground without planting the answer. Asking a pointed question of one candidate's references and not the other's is the version of this that turns a check into a search for a reason.

References

  1. 1. The Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 100 Years of Research Findings (working paper) Frank L. Schmidt, In-Sue Oh and Jonathan A. Shaffer (unpublished working paper; copy hosted by the University of Baltimore), 2016. home.ubalt.edu Supports the claim that many former employers now release only dates of employment and job titles, which is the environment a reference script has to work in.
  2. 2. Assessment and Selection: Other Assessment Methods - Reference Checking U.S. Office of Personnel Management (page is undated; first archived February 2013), 2013. opm.gov Supports the description of a defensible reference check as a structured telephone interview, and the point that written requests produce low response rates and less useful information.
  3. 3. Predictive Validity of Interviewer Post-interview Notes on Candidates' Job Outcomes: Evidence Using Text Data From a Leading Chinese IT Company Frontiers in Psychology, Volume 11, Sec. Organizational Psychology (Shanshi Liu, Yuanzheng Chang, Jianwu Jiang, Haigang Ma and Huaikang Zhou), 2021. frontiersin.org Supports the claim that notes naming job-related capabilities carried signal on later performance, promotions and turnover, at roughly 2 percent of performance per standard deviation, among 7,650 hired candidates at one Chinese firm.
  4. 4. Raising Doubt in Letters of Recommendation for Academia: Gender Differences and Their Impact Journal of Business and Psychology, 34, 287-303 (Madera, Hebl, Dial, Martin and Valian), 2019. s3.wp.wsu.edu Supports the caution that free-text references carry systematic differences in what gets volunteered: 54% of letters for women contained at least one doubt raiser against 51% for men, across 624 letters for academic posts.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.