Assessment design

Delegation, Description, Discernment, Diligence: What Each One Looks Like at Work

The four Ds of the AI Fluency Framework: Delegation is deciding what the model does and what stays yours. Description is saying what you want well enough to get it. Discernment is judging what came back, including noticing that a fluent answer is wrong. Diligence is owning the output and being straight about how it was made. Only Description leaves a readable trace in a finished document, which is why hiring against these four means watching the work rather than reading it.

Where Olive fits

Open a role and see what the work shows

Three of the four are only visible while the work is happening, which is the part that is hard to build in-house. Olive's reviewer writes six findings from a single 50-to-70-minute session, each carrying the timestamped excerpt it rests on.

Rank your shortlist

What are the four Ds?

Four competencies named by the AI Fluency Framework: Delegation, Description, Discernment and Diligence 1. The framework defines fluency as the ability to work "effectively, efficiently, ethically, and safely" with AI, and treats the four as the parts that make that possible. Underneath each sits a set of named sub-competencies, and the whole thing is set against three modalities of human-AI interaction rather than one 1.

  • Delegation covers goal and task awareness, platform awareness, and the act of splitting the work 1.
  • Description covers describing the product wanted, the process to follow, and the performance expected 1.
  • Discernment covers judging the product, the process and the performance of what came back 1.
  • Diligence covers creation, transparency and deployment, the last of which is verifying and vouching for AI-assisted output 1.

The three modalities matter more than they look. Automation is the model performing a task independently on direct instruction; augmentation is a person and a model co-defining and co-executing a task in a loop; agency is a person configuring a model to run future tasks on its own 1. The same four competencies apply in all three, but what Delegation means changes completely between them, which is why a single interview question about "how you use AI" collapses three different jobs into one answer.

One caveat before this becomes a rubric. The framework publishes no levels, no scoring anchors and no proficiency thresholds, and it describes its own sub-competency list as a current version that has already been revised. "Assessed against the four Ds" cannot mean a number.

What does each one look like in one real task?

Take something a people team actually runs: a quarterly attrition analysis for a leadership meeting, with an exit-survey export, a headcount file and four working days. The four Ds are four decisions inside that task, and they arrive in order. Watching where a person's version goes wrong tells you more than any definition of the terms does.

1. Delegation. The split gets decided before anything is generated. The model does the pivot, the chart drafts and a first pass at grouping free-text exit comments. The person keeps the question being answered, the definition of regretted attrition, and the recommendation. The failure shape is pasting the whole brief in and asking for the deck. 2. Description. Handing over the definition of regretted attrition this company uses, the comparison quarter, who is in the room and what decision the slide has to support. The failure shape is "analyze this attrition data", followed by twenty turns of correction that cost more than doing the pivot by hand. 3. Discernment. The model reports a jump in one function. The question is whether that function reorganised last quarter and whether the headcount denominator moved with it. Fluent output and a wrong denominator look identical on the slide. 4. Diligence. The comments were grouped by a model, so the person reads a sample of them, says so in the notes, and withdraws the one finding the sample does not support. The failure shape is a number in a leadership deck that nobody re-derived and everybody now believes.

Read top to bottom, only step two produced anything a reviewer could see afterwards. The other three produced a deck that happens to be right.

Which of the four leaves a trace you can read?

Description, and only Description. How well a task was specified shows through in what came back, so a sharp deliverable is real but partial evidence. Delegation, Discernment and Diligence are choices made and discarded during the work. The file records the outcome of those choices and never the choices themselves, and a wrong answer is formatted exactly like a right one.

CompetencyVisible in the finished file?Where it does show
DelegationNoWhat the person kept, and whether they can say why
DescriptionPartlyThe fit between the brief and what came back
DiscernmentNoWhat got thrown out, and what it was checked against
DiligenceNoWhether a claim was withdrawn rather than shipped

What the measurement shows is uncomfortable. In a field experiment with 758 consultants, on one task deliberately placed outside the model's capability, the group using GPT-4 was 19 percentage points less likely to reach the correct answer, 84.5% correct in the control group against 60% and 70% in the two AI conditions 2. One task, one sample, a 2023 model. The part that travels is that the people doing the work could not tell which side of the model's capability the task sat on.

The cheapest partial fix is asking for the process alongside the file. A chat log records what was handed over at the start, what got pushed back on, and what never made it into the file, which is three of the four competencies in a form somebody can actually read. It is also an account rather than an observation, and it can be tidied before submission. What to look for in a candidate's AI chat log covers where that reading holds and where it stops.

Don't copy a teaching framework into a resume checklist

The four Ds were built to teach four things in parallel, which is the right shape for a course and the wrong shape for a screen. A course can develop all four over weeks. A screen has to observe them in an afternoon, and three of the four cannot be observed in an application at all. Copying the list into a checklist keeps the vocabulary and loses the evidence.

What survives the copy is Description, inferred from a written artifact. So a checklist built this way weights hardest the one competency a finished deliverable already shows, and scores zero on the three that cost money to get wrong.

A self-rating does not rescue it. In a randomized trial, 16 experienced open-source developers working on repositories they had known for about five years forecast that AI tools would make them 24% faster, believed afterwards that they had been 20% faster, and were measured 19% slower 3. Sixteen developers in one setting is not a general productivity result, and the magnitude does not transfer. What transfers is that practitioners can be wrong about the direction of their own performance, which is the whole premise of a five-point self-assessment on Diligence.

So use the four Ds for what they are good for: naming the thing you want, in language a hiring panel can hold in its head, before designing an exercise that puts three of them in front of a person. The four Ds used as an interview rubric is the closest working version, and the vocabulary question underneath all of it is what the phrase means and who sets the bar.

See what gets scored

Common questions

What is the diligence competency actually about?

Taking responsibility for the finished product: verifying it, being transparent about how it was made, and vouching for it before it goes out. The framework splits it into creation, transparency and deployment, and deployment diligence is the one an employer cares about most, since it covers fact-checking, testing for accuracy and validating claims. It is also the item most often dropped when the four Ds get compressed into a slide, which is unfortunate, because it is the one that shows up as a cost when it is missing.

Are the four Ds a scoring rubric?

No. They are published as a taxonomy with no levels, no scoring anchors, no thresholds and no norms, and no validated instrument has been released alongside them. Any number attached to the four Ds was added by whoever is showing you the number, and the scale they invented is the thing to ask about. Used as a shared vocabulary for what an exercise is looking for, the four work well. Used as a scale, they are borrowing authority the framework never claimed.

Do the four happen in a fixed order?

Roughly, on any single task, and then they loop. A person delegates, describes, judges what came back, and either accepts it or goes round again with a better description. Diligence sits at the end because it is about what ships, but the decision to withdraw a claim usually happens mid-loop. Treating the order as rigid causes one specific mistake: assuming a person who described the task badly cannot have judged the output well. Those are different competencies and they come apart often.

Can a take-home assignment show more than one of the four?

Only if the process comes with the file. A submitted deliverable shows the fit between brief and output. Adding a short written account of what was handed to the model, what came back wrong and what was checked gets you a self-reported version of the other three, which is better than nothing and worse than watching. A live or recorded session is the only format where delegation, discernment and diligence are observed rather than described, and it costs more to run.

What if a candidate has never heard of the four Ds?

That tells you nothing about whether they have the competencies. The vocabulary comes from a taught course, so recognising the words is a poor proxy for doing the work. Plenty of strong practitioners frame the same decisions as scoping, briefing, checking and owning. Ask about the decisions, in your own domain's language, and let the candidate use whatever words they use.

References

  1. 1. Framework for AI Fluency Ringling College of Art and Design (Rick Dakan and Joseph Feller), 2025. ringling.libguides.com The four competencies, their named sub-competencies and the three modalities of human-AI interaction quoted throughout this article.
  2. 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the 19-percentage-point accuracy drop on the one task placed outside the model's capability, and the finding that the consultants could not tell which side of that line the task was on.
  3. 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that a self-rating is unreliable evidence: forecast 24% faster, believed 20% faster, measured 19% slower.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.