Interviewing

Can the AI Fluency 4Ds Work as an Interview Rubric?

The four Ds of AI fluency, Delegation, Description, Discernment and Diligence, work as interview rubric columns but not as a scoring scale. They come from a teaching course built to hold in every field [1], so each D says where to look, nothing about what counts as good in the role you're filling. Write each D's evidence standard off that occupation's task list, run the same four columns on every candidate, and mark them separately with no total. Discernment and Diligence are acts; watch those two in a work sample.

The takeFrameworks travel because they are empty, and the four Ds travel further than most. That is the appeal and the whole problem: a hiring rubric earns its keep by being useless at the company down the road. The anchors you write are the asset here, and the four names are filing. I'd expect the two Ds a candidate can talk through, Delegation and Description, to lose discriminating power fastest as coaching tools rehearse the shapes, which leaves the two you have to watch someone perform. A panel unwilling to write the anchors has not built a rubric. It has built a vocabulary.

Where Olive fits

Open a role and see what the work shows

A four-D rubric can rate how well a candidate describes checking a confident claim; it cannot rate the check. Olive puts those moments in front of the candidate as work: an occupational assignment with an AI assistant, read by a human reviewer who writes six findings and attaches the moment each one rests on.

Rank your shortlist

Can the four Ds work as an interview rubric?

As a structure, yes. As a scale, not until you add something. Delegation, Description, Discernment and Diligence come from a course Anthropic built with two academics to teach practical skills for effective, efficient, ethical and safe AI interaction 1. It names four behaviors worth watching. It says nothing about what a good one looks like in the job you are hiring for.

The names do most of the work, which is why the framework travels so well. Delegation is the decision about what gets handed over. Description is how the request is made. Discernment is the read on what came back. Diligence is what happens before the work leaves the candidate's hands. Four moments, in the order they occur in a real piece of work.

A curriculum and a rubric have different jobs, though. A curriculum has to hold in every discipline, so it stays deliberately domain-neutral. A rubric has to separate two candidates who both sound competent, and then survive somebody asking why one of them was passed over. Domain-neutrality is the exact property that stops it doing either.

So the conversion is not a formatting exercise. Keep the four Ds as your columns and write the thing the framework was never going to supply: what counts as evidence for each D in this occupation. Everything below is that step. For the behavioral picture underneath it, start with what being good at using AI actually looks like in an interview.

What does each D look like as an interview question?

Each D becomes one question about a specific piece of work the candidate already shipped, plus one follow-up asking for the thing nobody can rehearse: the artifact, the number, the moment something changed. The question names the moment and the follow-up gets the evidence. Asked in order, the four track one real project from brief to delivery in about fifteen minutes.

  • Delegation. "On the last thing you shipped with a model, what did you decide to do yourself?" Follow-up: what would have gone wrong if the model had done that part. A strong answer names a specific consequence. A weak one describes a preference.
  • Description. Skip prompt technique entirely. "Reconstruct the brief you gave it, then tell me what you added after the first answer came back wrong." You are listening for whether the constraint was in the brief or discovered in the output.
  • Discernment. Hand over a confident, wrong artifact from their own field and ask what is wrong with it. Strong candidates name the claim they would check first and say why that one. Weak ones proofread. This is the D worth building a real exercise around rather than a question: how to test whether a candidate notices when AI gets something wrong.
  • Diligence. "What did you check before it went out, and against what?" Then: "What changed as a result?" Nothing changed is a fair answer once. Four times in a row is the finding.

Two rules keep the round comparable. Ask the same four questions in the same order of every candidate, and score each answer before the next question is asked. A rating written after the fourth answer is a rating of how much you enjoyed the conversation. Unstructured interviewing is not a safe default to fall back on: hiring managers rate it their most effective tool, and dozens of studies put it among the worst predictors of on-the-job performance 4.

How do you write the evidence standard for each D?

Take the occupation's own task list and convert each D into a verb from it. O*NET publishes those tasks per SOC code, which makes it a serviceable starting point when nobody on the panel agrees on what the job requires. The test of a finished anchor is blunt: a competent person in that field would recognize it, and a competent person in a different field would not.

Financial and investment analysts (SOC 13-2051) are expected to evaluate and compare the relative quality of securities in an industry, and to interpret data on price, yield and future investment risk 5. Diligence in that role means re-deriving a figure from the filing it came out of. A candidate who says they asked the model to double-check its own number has not done it: they asked one source the same question twice.

News analysts, reporters and journalists (SOC 27-3023) are expected to check reference materials such as books, news files or public records to obtain relevant facts 6. Diligence there is a second independent source, named. Same D, same word, two standards that could not stand in for each other. A copywriter's version is different again, because the claim that has to survive is the one printed in the ad.

Delegation splits the same way. The boundary a paralegal draws around a filing is not the boundary a marketer draws around a campaign email, and a rubric that says "delegates appropriately" has quietly asked each interviewer to supply their own. Write one anchor sentence per D per role family, before the first interview, and put it in the same document as the questions. The version of this problem outside engineering is worked through in screening for AI judgment in a finance or marketing role.

Where does a four-D interview rubric stop?

At the difference between an account and an act. Each of the four Ds is a behavior, and an interview can only collect the candidate's description of a behavior that happened somewhere else. Structure raises what that description is worth, but it does not change what it is: a report, given by the person being evaluated, about their own work.

That gap has widened. Coaching tools now rehearse candidates on the exact question shapes a structured loop uses, so four well-formed answers will start arriving from everybody. Whether a structured interview still separates AI-coached candidates turns on how much of your rubric can be satisfied by a description alone, and on the four Ds as questions, most of it can.

Procedure is the second limit. The moment a rubric produces a rating that decides who advances, it is a selection procedure, and the EEOC's guidance on tests and selection procedures applies: a procedure that screens out a protected group has to be shown job-related and consistent with business necessity 3. Under the Uniform Guidelines, a content validity argument requires a job analysis of the important work behaviors, and the behavior demonstrated in the procedure has to be a representative sample of the behavior of the job 2.

That is the anchoring step arriving from the other direction. A generic four-D rating has no job analysis behind it and represents no particular job. The same four Ds, anchored to the tasks of a named occupation and applied identically to everyone in the slate, have one, and those written anchors are the record you hand over when a rejected candidate or a regulator asks what the rating meant.

Score each D separately, and never total them

Four marks, no sum. A total buys comparability you do not have and hides the shape that matters: strong Description with weak Discernment is a candidate who gets fluent output and ships it, which is a different hire from weak Description with strong Discernment. Averaging the two produces the same middling number and tells you nothing about either one.

The mechanics that make four separate marks hold up:

  • Three levels, not five. Demonstrated, partly demonstrated, not demonstrated. The two extra levels on a five-point scale are almost never distinguishable in a transcript, and they invite the middle.
  • Write the anchors first. One sentence per level per D, fixed before the first candidate. An anchor written after the third interview is a description of the third candidate.
  • Two scorers, independently, then reconcile on evidence. Each cites the moment behind the mark; the conversation is about the moments, not the impressions. Getting two people to the same mark is its own build: see writing a rubric for AI use that two reviewers score the same way.
  • Weight by the role, not by the level. Seniority is a poor proxy. A senior analyst seat can be almost entirely Discernment and Diligence, while a high-volume junior seat hinges on Delegation. Decide the weighting before the loop opens and leave it alone once interviews start.
  • Record the evidence, not the rating. One line per D quoting what the candidate actually said. Ratings age badly; the quote is what survives a debrief three weeks later.

None of that closes the account-versus-act gap. If the decision needs evidence of the four moments rather than testimony about them, the rubric belongs on a work sample where the moments happen in front of you, with the same four columns and the same anchors attached.

See how it works

Common questions

What are the four Ds of AI fluency?

Delegation, Description, Discernment and Diligence. They come from the AI Fluency course Anthropic built with Prof. Joseph Feller of University College Cork and Prof. Rick Dakan of Ringling College, which teaches practical skills for effective, efficient, ethical and safe AI interaction. Delegation is what gets handed over, Description is how the request is made, Discernment is the read on what comes back, Diligence is what happens before the work goes out. It is a teaching framework, so none of the four carries a standard for what good looks like in a particular job.

Can you assess the four Ds with questions alone?

Partly. Delegation and Description survive a question, because the candidate is describing a decision they made and can be pushed for its specifics. Discernment and Diligence do not: both are acts, and asking about an act returns a story about one. If the round has to produce evidence rather than an account, put a wrong artifact in front of the candidate, or run a short work sample with the assistant open and watch the four moments happen.

Should the four Ds be weighted differently for senior candidates?

Weight by what the role needs, not by the level on the req. Seniority is a poor proxy: a senior analyst seat can rest almost entirely on Discernment and Diligence, while a high-volume junior seat hinges on Delegation because the throughput forces the choice every day. Write the weighting into the rubric with one sentence saying why, apply it to everyone interviewing for that role, and do not change it partway through a slate.

How many rating levels should each D have?

Three, each with a written anchor sentence. Five-point scales invite the middle, and the two extra levels are almost never distinguishable from an interview transcript. Demonstrated, partly demonstrated and not demonstrated is enough to sort a slate, and it forces the interviewer to point at a moment instead of splitting hairs. Write the anchors before the first candidate, and keep the evidence line beside the mark.

Does Olive use the four Ds?

No. Olive assesses six named dimensions (problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification) from a 40-to-60-minute occupational assignment done with an AI assistant available. A human reviewer writes all six findings and attaches the moment in the session each one rests on. There is no composite number and no hiring recommendation, and the candidate is granted the same report the employer reads.

References

  1. 1. AI Fluency: Framework & Foundations Anthropic, 2025. anthropic.skilljar.com Names the four Ds (Delegation, Description, Discernment, Diligence), states that the course teaches practical skills for effective, efficient, ethical and safe AI interaction, and credits Prof. Joseph Feller (University College Cork) and Prof. Rick Dakan (Ringling College).
  2. 2. 29 CFR § 1607.14 — Technical standards for validity studies Uniform Guidelines on Employee Selection Procedures (Cornell Law School LII), 1978. law.cornell.edu Content validity requires a job analysis of the important work behaviors (14(C)(2)) and that the behavior demonstrated in the selection procedure be a representative sample of the behavior of the job (14(C)(4)).
  3. 3. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A selection procedure that screens out a protected group must be shown job-related for the position and consistent with business necessity.
  4. 4. How to Take the Bias Out of Interviews Harvard Business Review, 2016. hbr.org Unstructured interviews receive the highest perceived-effectiveness ratings from hiring managers, while dozens of studies find them among the worst predictors of on-the-job performance.
  5. 5. Financial and Investment Analysts (13-2051.00) O*NET OnLine, 2026. onetonline.org Tasks include evaluating and comparing the relative quality of securities in an industry and interpreting data on price, yield, stability and future investment-risk trends. Accessed 24 August 2026.
  6. 6. News Analysts, Reporters, and Journalists (27-3023.00) O*NET OnLine, 2026. onetonline.org Tasks include checking reference materials such as books, news files or public records to obtain relevant facts. Accessed 24 August 2026.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.