Interviewing

What a Good AI Answer Sounds Like at Junior, Mid and Senior

Ask junior, mid and senior candidates the same question about how they work with AI, and grade it against three bars. From a junior, expect verification: can they tell an output is wrong and say what they did next. From a mid-level candidate, expect delegation: which parts of a task they hand over, and why those parts. From a senior, expect the second-order call: what they sign off without reading closely, and what that cost when it went wrong. Write the three bars into the scorecard before the loop opens.

The takeA single fluency bar for every level produces the two worst debriefs available. A junior gets marked down for lacking strategy nobody has ever let them practise, and a senior passes on vocabulary picked up in a webinar last month. Level is not a volume dial on AI use. It is a statement about what a person is accountable for when the output is wrong, and that changes completely between someone whose work gets reviewed and someone who decides what gets reviewed.

Where Olive fits

Open a role and see what the work shows

Level shows in what somebody does when the output is wrong, and a conversation only reaches what they remember doing. Olive returns six findings from one occupational session, each stated as demonstrated, partly demonstrated or not demonstrated and anchored to a timestamped excerpt.

Rank your shortlist

What should a junior candidate be able to show?

Verification, in one concrete instance. A junior does not yet own a delegation strategy, and asking for one produces a memorised answer. What they can honestly show is a moment when the output was wrong, how they noticed, and what they did next. If they cannot name a single instance, that is the finding, and it does not depend on how long they have worked.

The reason to set the bar there rather than at output is that the tool does most of its measured lifting at this level. Across the staggered rollout of a conversational assistant to 5,179 customer support agents at one software firm, issues resolved per hour rose 14% on average, with a 34% improvement for novice and low-skilled agents and almost nothing for experienced ones 1. Pooling three company-run randomized trials across 4,867 developers, completed tasks rose 26.08% with a standard error of 10.3%, and less experienced developers both adopted the tool more and gained more from it 2.

Read those two carefully before using them. The first is one firm, one occupation, a pre-ChatGPT assistant trained on that firm's own transcripts, measured on problems with correct answers. The second measures weekly completed pull requests, which says nothing about defect rates or review burden, and its standard error puts the honest range somewhere between roughly 6% and 46%. Neither is a general law about juniors.

What they do establish is the shape of the hiring risk. If a tool raises a junior's output substantially, output stops separating candidates at that level, and the thing that still separates them is whether they can tell a good output from a confident wrong one. Hiring for verification rather than production is that argument at full length, and it matters most for new graduates who never learned to work without these tools.

Why is delegation the mid-level bar?

Because a mid-level engineer, analyst or marketer owns a whole task and has to divide it. The judgment being hired for is where the line falls: which parts go to the model, which parts stay, and what makes a part belong on one side. A candidate at this level who describes handing over everything or nothing has not been asked to make the call yet.

The follow-up is what does the grading. Ask why that part and not the next one. A weak answer is about speed or about what the tool is good at. A strong answer is about consequence: this part gets read by a client, this part feeds a model downstream, this part is where I am the only person who would notice the error. The candidate is describing a map of where damage lives in their work, and that map is the thing that transfers to a new company.

A second useful probe at this level is what they changed their mind about. Somebody who has been dividing tasks for two years has moved the line at least once, usually after something went out wrong. The answer names the incident, the new rule, and what the new rule costs them in time. Nobody rehearses that shape, because it requires a real loss.

A mid-level candidate who lacks a team-wide policy view has not failed this bar. That view belongs to the level above, and to a different question. Whether to hire the junior who is fast with AI over the senior who is not is the trade this bar exists to make legible.

What does a senior owe that a mid-level does not?

An answer about work they did not do. A senior signs off on other people's output, so the bar is what they read line by line, what they accept on trust, and what they escalate. That is a standing rule with a cost attached, and the strong version names the cost: what got through the last time the rule sat too loose, and what changed after.

Treat a claimed speed improvement as the least reliable input in the set. A reported speedup describes what somebody believes about their own work, and belief and measurement can come apart even for experts. Sixteen experienced open-source developers, working on repositories they had known for about five years, forecast that AI tools would make them 24% faster, believed afterwards they had been 20% faster, and were measured 19% slower 3. Sixteen people in one setting, on mature codebases, with early-2025 tooling. The magnitude does not transfer. What travels is that these developers had the direction wrong, not just the size.

So grade the rule. A senior who says the model saved the team a day a week has told you what they believe. A senior who says code touching billing gets read twice regardless of who wrote it, because of an incident in March, has told you what they do.

The people-management version of this question is genuinely separate and should be asked separately: what a manager lets their team ship, how they found out somebody was over-relying on a model, and what they did about it. Four questions for a manager whose team ships with AI covers that seat rather than this one.

Calibrate the bar to the level, not to the candidate

Put the three bars in the scorecard before the first interview, written as content rather than as a rating. Junior: names a specific wrong output and what happened next. Mid: draws a line through a task and defends where it sits. Senior: states a review rule and what it cost when it was wrong. An interviewer holding that page stops grading fluency.

The failure this prevents is the debrief where levels get compared. Four interviewers who each heard a different level of candidate, with no written bar, converge on whoever sounded most current, and current vocabulary is the cheapest signal in the room. Writing the bar down before the loop is the only cheap fix, and it is the same discipline that makes a defensible bar for good enough at AI in a specific role hold up if anyone questions it later. Written in decisions rather than in AI answers, the same page answers the general question of what level a candidate really is when every title in the stack says senior.

Never score a self-rating on any of the three. In a study of 288 Taiwanese teachers who took both a self-report and a knowledge-based test of AI literacy built on the same framework, correlations between the objective and self-reported factors ran from 0.07 to 0.24, and the six profiles the authors found included people who underestimated themselves as well as people who overestimated 4. Teachers in Taiwan are not job candidates, and a weak correlation is not proof that anyone flatters themselves. It is proof that a claimed level and a demonstrated one are two different measurements.

If the same question is being split across a panel, put the level bar in every seat's copy of the scorecard rather than in one interviewer's head. Giving each panel seat its own AI question divides the work; the level bar is what every seat grades against.

See how it works

Common questions

What if a junior has only ever worked with AI available?

That is now the normal case and it is not disqualifying. Ask the verification question anyway: a candidate who has always had the tool has still had outputs come back wrong, and the useful answer names one. What to watch for is a person who cannot describe the work happening any other way, because they will struggle the first time the tool is unavailable, restricted or wrong in a domain nobody on the team knows well. That is a training plan rather than a rejection.

Should a senior be expected to use AI more than a junior?

No, and expecting it produces bad hires. Measured gains concentrated among the less experienced in a 5,179-agent customer support study and a 4,867-developer coding-assistant trial, and a randomized trial of sixteen experienced developers found them slower with the tools while believing they were faster. A senior whose considered position is that the tool helps with three specific things and gets in the way elsewhere has given a better answer than one who claims it does everything. Grade the reasoning, never the volume.

How do I compare two candidates applying at different levels?

You do not compare them to each other, you compare each to the bar for the level being filled. If the requisition is for one role and the two candidates sit at different levels, the real question is which level the job needs, and that should be settled before either interview. Loops that skip that step end up hiring the more articulate candidate and then discovering the role required the other one's accountability.

What counts as evidence at the senior bar if they have never managed?

Review decisions, which most senior individual contributors make constantly without calling it management. What they approve without close reading, what they always check, what they push back on, and what standing rule they set for their own work all qualify. A staff engineer who requires a test before a generated change lands has a rule with a cost attached, which is exactly the shape the senior bar is looking for. Formal authority is not the requirement; accountability for other people's output is.

Does this change by function, or only by level?

Both, and level is the axis most loops get wrong first. The three bars hold across functions because verification, delegation and review are general shapes, but what a good answer contains is entirely local: a source outside the model means a filing to an analyst, a test suite to an engineer, a payer policy to a revenue-cycle specialist. Write the level bar once, then write the acceptable answer separately for each role the team hires.

References

  1. 1. Generative AI at Work (NBER Working Paper 31161) National Bureau of Economic Research, 2023. nber.org Supports the claim that measured gains concentrate in less-experienced workers: 14% on average, 34% for novices, minimal for experienced agents.
  2. 2. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers MIT Department of Economics (working paper; later Management Science), 2025. economics.mit.edu Supports the 26.08% increase across 4,867 developers, its 10.3% standard error, and the finding that less experienced developers gained more.
  3. 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that senior self-report is unreliable: 24% forecast, 20% believed, 19% slower measured.
  4. 4. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the instruction not to score a self-rating: correlations of 0.07 to 0.24 between self-reported and objective AI literacy across 288 teachers.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.