Assessment design
What Does 'AI Fluency' Mean on a Job Description, and How Do You Test It?
'AI fluency' on a job description means nothing you can test until the occupation is named. Written well, the requirement says which tasks the person does with an assistant, what the assistant may do inside them, and what they must defend without one. Test it with one real deliverable from the job, an assistant that will do all of it, and source material that only settles when someone opens it. Then read what the candidate framed, demanded evidence for, refused and checked, not how fast they finished.
The takeThe word is doing public relations work, and most of the hiring loop knows it. A requirement nobody can picture being demonstrated screens for confidence, and confidence is the one trait an assistant inflates for free. Which is why the strongest candidate is often the one who used the thing least: framed the problem, saw the model was the wrong instrument for step three, and did step three by hand. A rubric that counts prompts marks that person down. On what is public so far, nobody has shown heavy usage predicts anything worth hiring for. Grade the refusals.
Where Olive fits
Open a role and see what the work shows
The per-function part is the expensive part: an assignment and an answer key for each occupation, plus a record showing what the candidate actually did. Olive publishes twelve authored cases per occupation and returns six findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer against a timestamped moment in the session, and given to the candidate as well.
Rank your shortlistWhat does 'AI fluency' mean on a job description?
On its own, nothing you can test. In practice the phrase carries one of three meanings: an assistant will be available, the person is expected to use one on named tasks, or the person can be trusted with what it produces. The third is the actual skill, and it turns into a standard only when the job's own work is named.
The frameworks on offer stop one step short of that, by design. Anthropic's AI Fluency course, built with two academics, names four behaviors (Delegation, Description, Discernment and Diligence) and teaches practical skills for effective, efficient, ethical and safe AI interaction 1. It travels across every discipline because it is a curriculum, and a curriculum cannot say what strong delegation looks like for an underwriter. Four good columns; no rows.
So a requirement written at framework altitude reproduces the problem it was meant to solve. "Fluent with AI tools" tells a candidate nothing about what to prepare and tells an interviewer nothing about what to mark. It also ages badly: a tool named in a requirement this quarter is a different product in two. The fix is to write the line from the tasks, which is worked through in how to write AI skills into a job requirement.
The search results for this phrase make the gap visible. Most of what ranks is course material (module summaries, quiz answers, certificate curricula), which answers what fluency is called and never what it looks like in the job you are filling.
What does it mean in six different functions?
Six different things, because the artifact that has to survive is different in each. In marketing the claim in a deck has to hold; in finance the number that moves a recommendation has to be re-derived; in software the change has to pass a test somebody ran. Each occupation's O*NET task list already names that work, which makes it the shortest route to a definition you can defend.
| Function | What the occupation has to produce | What fluent means there |
|---|---|---|
| Marketing (SOC 13-1161) | Forecasts and trend reports that translate collected data into written findings 2 | The figure that reaches the deck is the one whose source was opened |
| Financial analysis (13-2051) | Judgments on the relative quality of securities, read off price, yield and future investment-risk data 3 | Any number that moves the recommendation is re-derived from the filing, not restated from the assistant |
| Software engineering (15-1252) | Modifications to existing software, and the testing or validation procedures that cover them 4 | The engineer owns the test that proves the generated change, and runs it |
| Data and analytics (15-2051) | Model comparisons on stated performance metrics, presented to non-specialists 5 | A fluent explanation of a result is checked against the data before it is repeated |
| Legal operations (23-2011) | Investigation of the facts and law of a matter against public records and other sources 6 | The cited authority is opened, and what it did not support is said out loud |
| Revenue cycle (29-2072) | Patient data abstracted and coded against standard classification systems and manuals 7 | The code the assistant proposed is checked against the manual before the appeal goes out |
Read down the last column and the same shape appears six times: something the assistant asserted was tested against something outside the assistant. What does not appear is volume. A candidate who decided the model was the wrong instrument for a step and did that step by hand has demonstrated the thing being measured, and a rubric that rewards heavy usage marks them down for it.
The six also differ in what a wrong answer costs, which should set how hard the exercise pushes. A coding error surfaces in a test run; a miscoded claim surfaces in a payer denial weeks later. For the functions where nobody on the panel writes code, the translation is worked through in screening for AI judgment in a finance or marketing role.
Write the requirement so it can be tested
Three parts, one sentence each: the task the person will do with an assistant, the part of that task the assistant may do, and the thing they must be able to defend without one. Written that way, the requirement doubles as the test specification: anything you cannot picture a candidate doing in front of you does not belong in it.
A financial analyst opening reads: "You will draft diligence memos with an assistant available. It can summarize the packet. You are accountable for every figure in the recommendation, re-derived from the filing it came from." A revenue cycle opening reads: "You will work a denial queue with an assistant drafting appeals. You decide which denials to concede, and the code on every appeal is yours."
Four things to cut from the version you have now:
- Tool names as requirements. Name the task and let the tool be the candidate's business. A named product in a requirement screens for the last employer's licensing, not for judgment.
- Years of experience with a two-year-old category. "Three years of generative AI" excludes people who started when the tool did.
- "Proficiency," "fluency," "savvy" with nothing after them. Each of these is a rating with no object. Put the object in.
- Anything you will not assess. A requirement nobody tests is a filter that runs on how confidently a candidate claims things, which is the trait AI-assisted applications inflate first.
One requirement per function, not one for the company. The same sentence cannot cover a paralegal and a data scientist without becoming untestable, and whether one assessment can cover every role is the version of this question that decides your build.
How do you test for it?
Hand the candidate a real deliverable from the job, an assistant that will do all of it if nobody stops it, and source material whose claims only settle when someone opens them. Then read the record of the work rather than the finished artifact. The output of a strong candidate and a passive one look similar; the paths do not.
What to build, in order:
1. Take the task from the occupation, not from a framework. Pick one deliverable the person would produce in their first month, and take its shape from the job's own task list rather than from a generic prompt exercise. 2. Plant something that only fails on inspection. A figure the packet does not support, a precedent that says the opposite of the summary, a passing test that tests nothing. Fluency is only visible when there is something to catch. 3. Write the answer key before the first candidate. Name the two or three moments where a competent person diverges from a passive one, and what each looks like in the record. 4. Score behaviors separately, and never total them. Framing, evidence demanded, what was kept, what was refused, what was verified. A single number buys comparability you do not have and hides the shape that matters. 5. Two scorers, independently, reconciled on evidence. Each cites the moment behind the mark. Getting two people to the same mark is its own build. See writing a rubric two reviewers score the same way.
The legal frame points the same direction as the design one. A procedure that decides who advances is a selection procedure, and one that screens out a protected group has to be shown job-related for the position and consistent with business necessity 8. Under the Uniform Guidelines, a content validity argument rests on a job analysis of the important work behaviors, and the behavior demonstrated in the procedure has to be a representative sample of the behavior of the job 9. A generic AI-fluency quiz samples no particular job. One task from this occupation, with the assistant open, samples exactly one.
Keep it to the length people finish. An exercise built around one deliverable runs 45 to 70 minutes; past that, completion falls and the sample skews toward candidates with unclaimed evenings. The redesign path from a take-home you already run is in rebuilding a work sample test for AI-assisted work.
Don't test speed, and don't trust the self-report
Both fail on the same evidence. In a randomized trial, 16 experienced open-source developers took 19% longer to complete real issues when AI tools were allowed, and still believed afterwards that the tools had sped them up by 20% 10. If practitioners cannot report the effect on their own work correctly, an interview answer about how well someone works with AI is not a measurement.
That rules out most of what gets used as an AI fluency test today. A multiple-choice literacy quiz measures recall of tool vocabulary, which is the part that goes stale fastest. A certificate records attendance at a course. A four-D questionnaire returns an account of behavior rather than the behavior. Discernment and Diligence are acts, and asking about an act gets you a story about one.
Speed is the other trap, because it is the easiest thing to record. Time-to-completion rewards the candidate who accepted the first draft over the one who opened the source and found it did not say what the summary claimed. If the exercise is timed at all, cap it and ignore the time inside the cap.
What survives all three problems is an act somebody can point at afterward: a source opened, a figure recomputed, a direction refused with the reason stated. Build the exercise so those acts are possible and so their absence is visible, which is the design behind testing whether a candidate catches what the AI got wrong.
Common questions
Is 'AI fluency' a real skill or a buzzword?
Both, depending on what follows it. As a standalone requirement it is a rating with no object, and it screens for confidence rather than judgment. Attached to an occupation it names something real and observable: whether a person frames a problem before generating, demands a source for the claim that matters, keeps the judgment they should not hand over, and tests an assertion against something outside the conversation. The difference is not the term. It is whether the job's own tasks are in the sentence.
What should an AI fluency requirement say on a job description?
Three things: the task the person will do with an assistant, what the assistant may do inside it, and what they must be able to defend without one. "You will draft diligence memos with an assistant available; it can summarize the packet; you are accountable for every figure in the recommendation, re-derived from the filing" is testable. "Fluent with AI tools" is not. Leave tool names out (they name the last employer's licensing, not the candidate's judgment), and leave out years of experience in a category that is two years old.
Can a multiple-choice AI literacy test measure fluency?
It measures recall of tool vocabulary, which is the part that ages fastest and the part a candidate can revise the night before. What it cannot reach is the behavior: nothing in a quiz shows whether someone opened the source behind a claim, refused a fluent draft, or re-derived a number. Quizzes are cheap to run at volume and fine as a knowledge check. Treat the result as evidence about study habits, not about how the person will work with an assistant on Tuesday.
How long should an AI fluency work sample take?
About 45 to 70 minutes, built around one deliverable the person would actually produce. That is long enough for a plan, a wrong turn and a check, and short enough that completion does not skew your pool toward candidates with free evenings. Anything past a couple of hours starts selecting for availability rather than judgment. Make it async and pausable if you can, and tell candidates the cap up front so nobody spends a weekend on it.
Does an AI fluency test have to be validated?
If it decides who advances, it is a selection procedure, and one that screens out a protected group must be shown job-related for the position and consistent with business necessity. Under the Uniform Guidelines, a content validity argument requires a job analysis of the important work behaviors and a procedure whose demonstrated behavior is a representative sample of the job. In practice that means writing down the task list you built the exercise from, keeping the answer key, and applying the same case and the same rubric to everyone in the slate.
References
- 1. AI Fluency: Framework & Foundations ✓ anthropic.skilljar.com Names the four Ds (Delegation, Description, Discernment, Diligence) and states the course teaches practical skills for effective, efficient, ethical and safe AI interaction. Built with Prof. Joseph Feller (University College Cork) and Prof. Rick Dakan (Ringling College).
- 2. 13-1161.00 - Market Research Analysts and Marketing Specialists ✓ onetonline.org Tasks include forecasting and tracking marketing and sales trends by analyzing collected data, and preparing reports of findings that translate complex findings into written text. Accessed 24 August 2026.
- 3. 13-2051.00 - Financial and Investment Analysts ✓ onetonline.org Tasks include evaluating and comparing the relative quality of securities in an industry and interpreting data on price, yield, stability and future investment-risk trends. Accessed 24 August 2026.
- 4. 15-1252.00 - Software Developers ✓ onetonline.org Tasks include modifying existing software to correct errors or improve performance, and developing or directing software system testing or validation procedures. Accessed 24 August 2026.
- 5. 15-2051.00 - Data Scientists ✓ onetonline.org Tasks include comparing models using statistical performance metrics such as loss functions or proportion of explained variance, and delivering oral or written presentations of results to management or other end users. Accessed 24 August 2026.
- 6. 23-2011.00 - Paralegals and Legal Assistants ✓ onetonline.org Tasks include investigating facts and law of cases and searching pertinent sources such as public records and internet sources, and preparing, editing or reviewing legal documents. Accessed 24 August 2026.
- 7. 29-2072.00 - Medical Records Specialists ✓ onetonline.org Tasks include identifying, compiling, abstracting and coding patient data using standard classification systems, and consulting classification manuals to locate information about disease processes. Accessed 24 August 2026.
- 8. Employment Tests and Selection Procedures ✓ eeoc.gov A selection procedure that screens out a protected group must be shown job-related for the position and consistent with business necessity.
- 9. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 ✓ ecfr.gov Content validity requires a job analysis of the important work behaviors (14(C)(2)) and that the behavior demonstrated in the selection procedure be a representative sample of the behavior of the job (14(C)(4)).
- 10. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org Randomized controlled trial: 16 experienced open-source developers on 246 real issues took 19% longer with AI tools allowed, while believing afterwards that the tools had sped them up by 20%.
10 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.