Interviewing

STAR Survives If You Score the Result and the Discarded Option

The STAR method is worth keeping even now that candidates draft their answers with AI; what changes is where the scoring weight goes. Situation, Task and Action can be produced from a job description in seconds, so they now confirm preparation rather than experience. Result was always the expensive letter: a real one carries a number, a denominator and somebody who could dispute it. Add the slot the template never had, the option the candidate discarded and why, and score those two.

The takeSTAR is being blamed for something it never claimed to do. The complaint is that every answer now arrives in the same shape, and the shape was never the evidence. It was a request for the parts of a story interviewers otherwise forget to ask for. Dropping it usually means falling back to a freeform conversation, which is the worst-evidenced format available and the one most exposed to whoever happens to be in the room that day. Keep the container. Change what has to go inside it.

Where Olive fits

Open a role and see what the work shows

A discarded option is easy to describe and hard to invent once the follow-ups keep going. Olive puts the choice inside an assignment the candidate is working on rather than one they are recalling, and a human reviewer writes down what they kept, what they handed to the assistant, and what they threw away.

Rank your shortlist

Should you keep the STAR format?

Keep it. STAR is a prompt for the interviewer as much as for the candidate: it stops a round drifting into hypotheticals and it makes two answers comparable, which is the property the evidence actually rewards. What changed is that the frame no longer costs anything to fill. Treat a well-formed STAR answer as the start of the conversation rather than a result in it.

Comparability is what the numbers are about. In the 2022 correction of the selection literature, structured interviews came out top ranked at .42, ahead of job knowledge tests at .40, work samples at .33 and cognitive ability at .31, with unstructured interviews at .19 1. Every one of those is a corrected correlation with supervisor ratings of job performance, not an accuracy rate, and every one of them was revised down from the 1998 estimate still circulating in training decks. Read the list for rank order and treat the magnitudes as soft: what separates its top from its bottom is whether the same thing was asked and rated the same way.

The drafted answers are good because short, self-contained professional writing is exactly where the measured gains sit. In a pre-registered experiment, 444 college-educated professionals did occupation-specific writing tasks, and the half given ChatGPT finished 37% faster than a control group averaging 27 minutes while scoring 0.45 standard deviations higher, with the largest gains going to the weakest writers 2. That was mid-level business writing done once, online, for pay, with no revision cycle and no consequences. A STAR answer is very close to that task, which is why the shape arrives polished and the substance does not necessarily arrive with it.

Dropping the format does not drop the obligation. The 1978 Uniform Guidelines on Employee Selection Procedures, the federal rules that govern this in the United States, define a selection procedure to include the full range of assessment techniques through informal or casual interviews and unscored application forms 3, so swapping a scored format for a chat does not move the step outside federal selection law. It only removes the record of what was asked. How that lands on a specific process is a question for counsel.

Which of the four letters was ever expensive?

Result, and only barely. Situation and Task describe context a job posting already implies. Action describes what a competent person would do, which is close to the most-written text on the internet for most roles. Result is the only letter that requires the candidate to have been present when a number came back, and most answers still fill it with an adjective.

So interrogate the R and let the rest go past quickly. Four asks do it:

  • What did the result measure, over what period, against what baseline? An answer with no baseline is a claim about direction only, and direction is free.
  • Who else could have produced it? Shared credit is normal and specific. Sole credit for a team outcome is worth one more question, because the other names are either ready or they are not.
  • What would have happened if nobody had done anything? This separates a result from a trend somebody stood next to.
  • What is the number now? Improvements that survived are different from improvements that were announced.

A fabricated result answers the first ask with a percentage and the other three with generalities, because a percentage is easy to generate and a denominator is not. That asymmetry is the whole diagnostic, and it does not require anyone to guess how the answer was written. Why every candidate suddenly gives the same polished STAR answer covers what to do when the whole slate arrives in the same shape.

Add two slots the template does not have

The discarded option and the disagreement. Ask which approach came second and what would have broken if they had taken it, then ask who argued for something else and what their best point was. Neither slot exists in the template, so neither has a rehearsed answer waiting. Both are easy for somebody who was there and awkward for somebody who was not.

What a real discarded option sounds like: it is close. Somebody who genuinely weighed two approaches picks the second one with a note of regret, names a condition under which it would have won, and can say what it would have cost to switch six weeks in. An invented one is symmetric and safe, the obviously worse choice presented as a serious alternative, with no cost attached to either side and no residue of the argument.

The disagreement slot works the same way and is harder to fake, because it requires a second person with a position. Ask for the strongest version of what they said, not a summary of why they were wrong. Candidates who lived through it can state the other case well and often still concede a point. Candidates reconstructing it produce a straw opponent, because a straw opponent is what a plausible narrative needs.

Both slots reward going one layer further rather than one question wider. What follow-up questions expose whether someone actually understands the answer they just gave is the mechanics of that chain.

How do you score a STAR answer now?

On two things: how specific the Result gets when pushed, and whether the discarded option holds up. Write the anchors before the round. A weak answer names an outcome with no measure. A middling one has a number with no baseline. A strong one has a number, a comparison, and an honest account of what it cost, including the part that did not work.

Give structure a weight of zero and say so on the form. If an interviewer can score an answer high because it was well organised, they will, and that line now measures preparation. The three lines worth keeping are evidence quality, tradeoff quality and consistency under follow-up, each rated on the same scale by every interviewer and each backed by a quoted sentence from the candidate rather than an impression of them.

One line does not belong on the scorecard at all: sounded rehearsed. It is not a rating, it cannot be evidenced, and it will be applied unevenly to people who prepare because preparing is the only way they can compete. If a panel is reaching for it, the honest reading is that the follow-ups stopped one layer too early.

One level up from STAR sits the question the format serves: do behavioral interview questions still predict anything when candidates rehearse them with AI.

See how it works

Common questions

Is STAR still worth teaching to interviewers?

Yes, as a note-taking frame rather than a scoring frame. It gives an interviewer four places to put what they heard, so two answers can be compared in a debrief and a round stops drifting into hypotheticals. The training change is small: tell interviewers that a complete STAR arc earns no points by itself, and that the points live in what the follow-ups turn up.

What about SOAR, CAR and the other variants?

The letters do not matter. Every variant asks for context, action and outcome in some order, and every variant is equally easy to fill from a job description. Pick one, use it consistently across the panel so the notes line up, and spend the energy you saved on the two questions none of the acronyms include: what was given up, and who disagreed.

Should candidates be told not to use AI to prepare?

It is unenforceable and it selects for compliance rather than skill. A rule you cannot check is a rule that penalises the candidates who follow it, and preparing with a model is now closer to reading the company blog than to cheating. If preparation genuinely damages what a question measures, the question is the thing to change. Say what you will be asking about and what you will be rating, and let people get ready.

How many STAR questions belong in one round?

Three or four in a 45-minute round, with real follow-ups on each. Six questions at one layer produces six clean arcs and nothing to compare, because the first pass is where every answer now looks similar. Three questions taken two or three layers down produce specifics a debrief can actually argue about, and they fit the same clock.

What does a strong Result sound like?

It has a number, a baseline and a limit. Something like: the queue went from about eleven days to six over one quarter, measured on the same report finance already ran, and it drifted back to eight once the temporary contractor left. That answer volunteers the part that did not hold. Somebody who watched a result over time usually knows its shelf life, so ask what the number looks like now and let the answer carry itself.

References

  1. 1. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the rank order and revised validity estimates quoted for structured interviews, job knowledge tests, work samples, cognitive ability and unstructured interviews.
  2. 2. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) MIT Department of Economics, 2023. economics.mit.edu Supports the claim that short professional writing tasks are where measured AI gains are largest: 37% faster, 0.45 standard deviations higher, largest gains to the weakest writers.
  3. 3. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that an informal or casual interview is itself a selection procedure, so replacing a scored format with a chat does not move the step outside the Uniform Guidelines.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.