Interviewing

Description Is Trainable; Discernment Is What You Hire For

The scarce AI competency is discernment, noticing that a fluent, confident answer is wrong. Description, specifying the task well, is the one you can train after the hire: it is a habit, and a new hire picks it up watching colleagues do it for a fortnight. Discernment runs on domain knowledge that accumulates over quarters of doing the work. In an interview, the deciding evidence is a candidate catching something wrong in material you handed them and naming what they checked it against.

The takeA framework built for development gets weighted evenly across four competencies because a course has to teach all four. A requisition is not a course. Spreading the bar evenly means a candidate who briefs a model beautifully and checks nothing clears it, and that is precisely the hire who looks strong for a quarter and then puts a wrong number in front of a client. If the bar has to come down to one thing, make it the wrong answer they caught.

Where Olive fits

Open a role and see what the work shows

An interview stops where the account stops. Olive hands the candidate the task instead, with an assistant available and a human reviewer who writes six findings, each carrying the moment in the session it rests on.

Rank your shortlist

Which AI competency is actually scarce?

Discernment: judging what came back, including noticing that a confident, well-formatted answer is wrong. It is scarce because it is not really an AI skill. It runs on knowing a domain well enough that a wrong number feels wrong before anyone can articulate why, and that knowledge accumulates over quarters of doing the work rather than over an onboarding week.

The named version of the four competencies is delegation, description, discernment and diligence, and what each one looks like in real work is worth reading before weighting them, because the framework itself is deliberately neutral about which matters most. It was written to teach. A requisition has to choose.

The cost of missing discernment has been measured once with unusual clarity. In a field experiment with 758 consultants at a large firm, on a single task deliberately placed outside the model's capability, the group using GPT-4 was 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right, against 60% and 70% in the two AI conditions 1. One task, one sample, a 2023 model, and the capability line moves with every release. What travels is the shape of the failure. Those consultants did not lack a fact. They could not tell which side of the line the task was sitting on, and the model gave them no clue, because a model outside its competence sounds exactly like a model inside it.

Why is description trainable and discernment not?

Because one improves by repetition and the other by accumulation. Specifying a task is a habit: name the constraint, name the audience, hand over the source material, say what a wrong answer would look like. A new hire picks it up from watching two colleagues do it for a fortnight. Discernment cannot be practised in the abstract, because what looks wrong depends entirely on what a right answer usually looks like in this field.

The productivity literature backs the first half of that, if it is read as a claim about which part of the work the tool substitutes for rather than as a claim about output. In a staggered rollout of a conversational assistant to 5,179 customer support agents at one software firm, issues resolved per hour rose 14% on average, with a 34% improvement for novice and low-skilled agents and close to nothing for experienced ones 2. Pooling three company-run randomized trials across 4,867 developers, an AI coding assistant raised completed tasks by 26.08%, with a standard error of 10.3%, and less experienced staff both adopted it more and gained more 3.

The standard error on that second figure is wide enough that the range, not the point estimate, is the honest reading. Both studies also measure volume on tasks with correct answers rather than judgment on open work, in different firms and different eras. Taken together they still show one consistent thing: the tool closes the gap fastest exactly where the work is production, and it closes nothing where the work is knowing the produced thing is wrong. That last inference is a reading of the pattern, not a finding either paper reports.

The weighting follows from that pattern. If a tool levels up the specification-and-production half of the job for junior staff within weeks, paying a premium at hire for that half is paying for something you were about to get anyway.

Ask for the wrong answer they caught

One question decides more of this round than the rest of the loop. Hand the candidate material out of your own work with something wrong in it, then ask what they would do with it. Prompt quality, tool names and speed all move around that question. What you are listening for is whether the error surfaced at all, and what got checked in order to surface it.

Two ways to run it, in ascending cost:

1. The recall question, ten minutes. "Show me something an assistant gave you that you did not use, and how you knew." It works in a phone screen and it returns an account rather than an act. The follow-up is the whole question: what did you open, run, recompute or call to find out. 2. The seeded exercise, twenty minutes plus an afternoon to build. Real material from the role with one confidently wrong thing in it, an assistant available, and a rubric written before the error was planted. How to plant an error a candidate's field would catch is the design problem worked through.

What counts as catching it is naming the thing checked against. "It felt off" is taste, which is weak evidence and unevenly distributed by how confidently somebody talks. Opening the filing, re-running the query, reading the payer policy or calling the customer is verification, and it is the same move in every function even though its object changes.

Diligence belongs next to discernment for the same reason: it is a disposition rather than a technique. The question there is whether the person withdraws a claim they cannot stand behind, which nobody learns during onboarding either. Hiring for verification rather than production is the version of this argument aimed at the job description.

Where does a weighted bar go wrong?

Three ways, and the first is the one nobody notices. Weighting discernment means testing your domain, so the exercise has to be built out of your material or it measures nothing at all. The second is dropping description to zero, which is not what a floor means. The third is grading the catch rather than the check behind it.

  • It is domain-bound, so it does not travel. A candidate strong in revenue-cycle work will miss a planted error in a marketing brief, and that is the instrument working correctly rather than the candidate failing. Never carry a discernment result across functions, and never let one panel's exercise become the company's standard exercise.
  • A floor is still a bar. Somebody who cannot brief a model at all will burn a week of a senior's time in month one. Test description as a pass or a fail on one short task, not as a scale that a strong specifier can use to outrank a strong checker.
  • Grade the check, not the catch. A candidate who missed the planted error but described exactly what they would have opened to verify the claim has shown more than one who spotted it by luck and cannot say how. Write that into the rubric before the first interview, because it is the distinction panels lose first.
  • And the ceiling on all of it. An interview gets you a description of a check, never a check. That gap is why the exercise exists, and it is also why the word on the posting matters less than the behavior underneath it: what fluency means and who gets to set the bar.

See how it works

Common questions

Isn't discernment just domain expertise under a new name?

Largely, and knowing that is useful. The AI-specific part is applying that expertise to output arriving fluent, complete and confident, which is a harder condition than reading a colleague's rough draft where the seams are visible. Someone who already checks a peer's work carefully will usually check a model's, once they stop assuming the model has done the checking. What is new is the assumption, not the skill, and the assumption is what an exercise has to break.

Can discernment be trained at all?

Slowly, and mostly by making it somebody's job to be wrong in public. Structured review of AI-assisted work, post-mortems on errors that shipped, and a norm that a withdrawn claim is a good outcome all move it. What no onboarding programme can compress is the domain knowledge underneath, which is why the training version of this takes quarters and the hiring version takes an afternoon. Budget for both and expect different timelines.

How does this apply to a new graduate with no domain yet?

It changes what you ask for rather than removing the requirement. A new graduate has no store of right answers to check against, so the observable version is whether they ask. Look for the person who names which colleague, which document or which source they would go to before shipping a claim they could not verify, and treat that as the early-career form of the same disposition. Somebody who ships confidently at that stage is not fast, they are unsupervised.

Where does delegation fit in the weighting?

It splits, so weight the halves separately. Part of it is a habit, learnable quickly, and part of it is a judgment about what should never be handed to a model at all, which is closer to discernment. It is also the competency most shaped by what your own policy allows, so a candidate's delegation instincts from a previous employer may not transfer. Ask what they deliberately kept and why, then judge the reason rather than the split.

Should the job description use the word discernment?

No. Write the behavior: catches errors in AI-drafted material in this domain, and says what the claim was checked against. Framework vocabulary in a posting invites resumes tuned to the vocabulary and tells a candidate nothing about what the work is. Keep the four competencies as the panel's internal language for what the exercise is looking for, and keep the posting in the language of the job.

References

  1. 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the expensive failure is not knowing which side of the capability line a task sits on: 19 percentage points worse outside it.
  2. 2. Generative AI at Work (NBER Working Paper 31161) National Bureau of Economic Research, 2023. nber.org Supports the claim that measured gains concentrate in less-experienced staff: 14% on average, 34% for novices, near zero for experienced agents.
  3. 3. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers MIT Department of Economics (working paper; later Management Science), 2025. economics.mit.edu Supports the claim that less experienced staff adopt an assistant more and gain more: 26.08% more completed tasks, standard error 10.3%, across 4,867 developers.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.