Interviewing

Why Does Every Candidate Give the Same Polished STAR Answer?

Every candidate gives the same polished STAR answer because STAR is a fixed four-beat template an assistant fills from the job posting in one prompt. Everyone applying pastes the same posting, so everyone gets the same story back. Coaching has always raised interview scores by organizing answers better [2]; the coach is now free and awake at 2am. None of it is dishonesty, and more follow-ups only give a rehearsed story more room. Ask instead for the option they rejected, the person who disagreed, the baseline under the number.

The takeThe uniform answer is a verdict on the question, not on the person giving it. Behavioral interviewing spent decades rewarding whoever could narrate in four beats, which quietly rewarded whoever could afford the coaching, and that advantage is now free. Losing it makes the round fairer. The reflex worth resisting is the hunt for tells: I suspect a panel that starts scoring who sounds rehearsed is mostly scoring nerves, second languages and practice time. Keep scoring polish and the round hires the best-rehearsed candidate in the room, which is a decision you never meant to make.

Where Olive fits

Open a role and see what the work shows

A STAR answer is a description of work, and the description is the part that now generates; Olive is the work itself: a 40-to-60-minute occupational assignment with an AI assistant that will happily do all of it, read afterward by a human reviewer who writes six findings on what the candidate framed, demanded, kept, refused and tested. The candidate is granted the same report the employer reads.

Rank your shortlist

Why do the answers all sound the same?

Because STAR is a public, fixed, four-beat template, and a fixed template is the easiest thing in the world to generate. Paste the job description into any assistant and ask for behavioral answers; what comes back is a Situation, a Task, an Action and a quantified Result, in that order, in competent prose. Every candidate for your role pastes the same posting, so every candidate gets the same story.

Coaching is not new. Only the price is. Voluntary interview coaching was already raising situational-interview scores in police and fire promotional processes twenty-five years ago, controlling for job knowledge, motivation to do well, race and sex. The study traced the effect to a specific mechanism: coached candidates organized their answers better, and that organization predicted the score 2.

So the format has always rewarded rehearsal. What changed is that the coaching session used to be a scheduled event a fraction of candidates attended, and is now free and available at 2am to all of them. The distribution didn't improve; it compressed. Your strongest candidate and your weakest one sound alike in round one, because the thing they used to have in different amounts, preparation, stopped being scarce.

State that plainly before doing anything else, because it sets the whole diagnosis: this is not a wave of dishonesty. It is an instrument that measured preparation, still measuring preparation, at the moment preparation became free. The answers aren't fake. They're no longer diagnostic. Structured interviews versus AI coaching works through which parts of the format survive that and which parts don't.

What converges, and where?

The convergence point is occupational, not universal. Marketing candidates land on the same campaign-lift narrative; data analytics candidates land on the same stakeholder-alignment story; engineers land on the outage they debugged under pressure. An assistant reaches for the most-documented task in an occupation, and the most-documented tasks are the ones with a clean arc: a problem, an intervention, a number that moved.

So build the breaking question from the occupation's real failure mode rather than from the competency you meant to test. Two worked examples, drawn from what these jobs actually do:

  • Marketing, SOC 13-1161. O*NET lists the work as measuring the effectiveness of marketing, advertising and communications programs, and forecasting and tracking marketing and sales trends 3. The generated answer always attributes a lift to the campaign. The question that breaks it: what else was running that quarter, and how did you separate it from your effect? Attribution is the actual failure mode of the occupation, and no template carries it.
  • Data and analytics, SOC 15-2051. O*NET lists comparing models using statistical performance metrics such as loss functions or proportion of explained variance, and delivering results to management or other end users 4. The generated answer is always about aligning stakeholders on a definition. The question that breaks it: which model did you compare against, and what did the losing one get right? A rehearsed narrative names the meeting. The work names the metric.

The pattern generalizes. A competency such as "influencing without authority" or "driving results in ambiguity" has a story attached, and the story is generatable because thousands of versions of it are already written down. A failure mode (attribution, leakage, a metric that improved while the thing it stood for got worse, a denial appealed that should have been conceded) has one specific answer per candidate, and specificity is what a script cannot carry into a room.

Does asking more follow-ups fix it?

The evidence runs the other way. In the study that built the standard measure of interview faking, past-behavior questions were more resistant to faking than situational ones, and follow-up questioning increased faking rather than reducing it 1. More turns give a prepared candidate more room to embellish, unless each turn asks for something that can be checked.

The same research put the base rate where it is uncomfortable. Over 90% of the job candidates studied faked during employment interviews, with the behaviors semantically closest to lying (inventing an experience, borrowing someone else's) running between 28% and 75% depending on the specific behavior 1. That was measured long before any of this was automated. Treat a polished answer as normal, not as evidence of anything.

A follow-up earns its place when it makes the candidate operate on their own answer instead of continuing it. "Tell me more about the stakeholder piece" continues. "You said conversion rose eleven points. What was the baseline period, and what does that number become against the prior quarter?" operates. The second one has an answer that is either there or it isn't, and the candidate finds out which in front of you.

Follow-up questions that expose understanding covers the mechanics. The short version: depth of one is the tell. A rehearsed narrative survives the first probe and dies on the second, because the preparation covered the story rather than the decisions inside it.

Ask for what the story left out

Every STAR answer is a selection: it names the action taken and silently drops the ones rejected. The discarded branch is where the judgment lives, and it is the part no template produces, because a generated story never considered anything it then dropped. Ask for it directly, and ask what the choice cost. Four probes do most of the work, and none of them need a tool.

  • What did you decide against, and what did it cost? A real project has a rejected option with a price attached. A generated one has a straight line from problem to result.
  • Who disagreed, and were they right? Named disagreement is specific, checkable at reference stage, and almost never in the prepared version.
  • What would have made you wrong? This asks for the falsifier, which is the same thing a good analyst asks of their own work before shipping it.
  • Describe the version before the good one. Every real deliverable has a bad draft. Ask what changed between them and why.

The stronger version stops asking about past work and puts material in front of the candidate. Hand them a short, confidently wrong AI output from their own occupation and ask what they'd do with it. The answer is produced on the spot, against something you wrote, and no amount of preparation reaches it. Testing whether a candidate catches AI errors covers how to build one that reads as work rather than as a trick. See how Olive measures this.

Change what the round measures

Keep the structure and change the evidence. Structured interviewing is still worth having (same questions, same order, a rubric written before the first candidate), but the rubric has to score something other than narrative quality. Score whether a specific number was defended, whether a rejected option was named, whether a claim was checked against something outside the story the candidate came in with.

Three changes carry most of the load:

  • Write the rubric first, and score behaviors rather than stories. "Named a constraint before proposing an approach" is observable. "Strong communicator" is a description of polish, which is the thing that stopped varying.
  • Let the assistant into the room on purpose. A round that bans AI measures a version of the job nobody has. A round where the model is open, and the candidate decides what to hand it and what to keep, produces evidence you can read and score.
  • Move one round from description to work. Twenty minutes of real occupational material beats an hour of narrative. The candidate cannot rehearse a task they see for the first time in front of you.

Polish stops carrying information the moment the round stops rewarding it. The same trade appears one stage later, where every take-home comes back clean and well-formatted: grading polished take-homes is that problem, and the fix has the same shape: score the decisions the work reveals, not the finish on it.

See how it works

Common questions

Is a polished STAR answer a red flag?

No. Preparation is a reasonable response to a high-stakes conversation, and candidates were coached long before assistants existed. Coaching was just rationed by money and access. Penalizing polish selects for candidates who didn't prepare, which is not a trait worth hiring for. Judge what the answer survives instead: a probe for the number's baseline, for the option that was rejected, for the person who disagreed. That test works the same on a coached answer and an uncoached one.

Do behavioral or situational questions hold up better?

Past-behavior questions hold up somewhat better. The research that developed the interview faking behavior scale ran an experiment on exactly this and found past-behavior questions more resistant to faking than situational ones. Neither is immune. A past-behavior question still asks for a story, and stories generate well, so the resistance comes from the follow-ups a real event can support and an invented one cannot: dates, names, the draft that got thrown away, a number nobody would have guessed.

Should you ban AI from interview preparation?

You can't enforce it, and the rule would select for candidates who follow instructions nobody checks. Say the opposite. Tell candidates preparation is expected and that the round asks about decisions rather than narratives, which changes what they prepare. A candidate who spends prep time reconstructing what they actually decided, and why, arrives better than one who memorized four stories, and the conversation improves either way.

How many candidates actually embellish?

Most of them, at some level, and it predates AI by decades. The study that built the standard faking measure reported that over 90% of the job candidates it examined faked during employment interviews. The behaviors closest to lying (inventing an experience, borrowing someone else's) ran lower, between 28% and 75% depending on the specific behavior. The useful conclusion isn't suspicion. It's to stop treating a smooth narrative as evidence and start asking for things that are either true or not.

Does Olive check whether an interview answer was prepared with AI?

It does not. Olive is an employer-purchased assessment of how a person works with AI: the candidate does a 40-to-60-minute task from their own occupation with an assistant available, and a human reviewer writes six findings on what happened: how the problem was framed, what evidence was demanded, what was kept, what was built in between, what was refused, and what was tested against something outside the conversation. There is no composite number and no ranking, and the candidate is granted the same report the employer reads.

References

  1. 1. Measuring faking in the employment interview: development and validation of an interview faking behavior scale Journal of Applied Psychology, 2007. pubmed.ncbi.nlm.nih.gov Six studies, N = 1,346. Over 90% of undergraduate job candidates faked during employment interviews; behaviors closer to lying ranged 28-75%. The Study 6 experiment found past-behavior questions more resistant to faking than situational questions, and found follow-up questioning increased faking.
  2. 2. Interviewee coaching, preparation strategies, and response strategies in relation to performance in situational employment interviews Journal of Applied Psychology, 2001. pubmed.ncbi.nlm.nih.gov 213 candidates for police and fire promotions: coaching attendance predicted situational-interview performance controlling for job knowledge, motivation, race and sex, and the effect ran through better-organized answers, which themselves predicted the interview score.
  3. 3. 13-1161.00 - Market Research Analysts and Marketing Specialists O*NET OnLine (U.S. Department of Labor), 2026. onetonline.org Occupational task list, including measuring the effectiveness of marketing, advertising and communications programs, and forecasting and tracking marketing and sales trends.
  4. 4. 15-2051.00 - Data Scientists O*NET OnLine (U.S. Department of Labor), 2026. onetonline.org Occupational task list, including comparing models using statistical performance metrics such as loss functions or proportion of explained variance, and delivering results to management or other end users.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.