Assessment design

A Multiple-Choice AI Fluency Test Measures Recall, Not Fluency

AI fluency can be tested, if the test is a task. A quiz measures AI literacy, which is recall, and a self-rating measures what someone believes about themselves, which tracks demonstrated skill weakly. Fluency shows up in the decisions someone makes while doing the work, with a wrong answer sitting in front of them, so measuring it takes an occupational assignment, an assistant, a deadline, and something in the supplied material that is not true. What comes back is evidence a person can quote, not a number.

The takeThe scoring is the tell. If an instrument returns a number the moment the candidate clicks submit, nobody read the work, and nobody read the work because there was no work to read. That is a defensible thing to sell and a weak thing to buy, since anything gradeable in milliseconds is testing something the candidate could have looked up, and looking things up is the part that got cheap this decade. Pay for the reading, or do the reading yourself.

Where Olive fits

Open a role and see what the work shows

Anyone building this in-house pays twice: once for the answer key, once for the evidence trail behind every finding. Olive runs a forty-to-sixty-minute assignment on the candidate's own clock and returns six findings written by a person, stated as demonstrated, partly demonstrated or not demonstrated.

Rank your shortlist

Can a test measure AI fluency?

Part of it. A written test measures knowledge about AI cheaply and consistently: what a model does, where it fails, what a hallucination is, which data should never go into a prompt. That is worth having and it is not fluency. The competencies people mean by the word are decisions made while the work is happening, and a question with four options presents none of the conditions under which those decisions get made.

One research group has put a version of this on record. Researchers building an AI literacy assessment for a US Navy robotics training programme reported that a scenario task simulating AI use on the job outperformed both the tests they had adopted from earlier research and the ones they wrote themselves, and argued that prevailing AI literacy assessments emphasise foundational technical knowledge over practical knowledge such as interpreting model outputs, selecting tools and identifying ethical concerns 1. One programme, one occupational context, and no published margin a writer can quote: the abstract reports no sample size and no effect size. It supports a design argument rather than a benchmark, and it is worth noticing that it is the academic version of a claim assessment vendors have an interest in making.

So keep the quiz if it is doing a job. It sets a floor, it administers identically to everyone, and it is easy to defend on content. It will not tell two competent-looking candidates apart, which is usually the decision actually in front of you. What an AI skills assessment should measure before you pay for one is the buying-side version of the same split.

Why doesn't a self-rating fill the gap?

Because a self-rating and a demonstration measure different things. In a study of 288 teachers who took both a self-report and a knowledge-based test of AI literacy built on the same underlying framework, correlations between the objective and self-reported factors ran from 0.07 to 0.24 2. A weak correlation is not proof that people flatter themselves. It says the two instruments do not substitute for each other.

That study found six profiles, including people who overestimated themselves and people who underestimated themselves. Its subjects were teachers in Taiwan, so it says nothing about a hiring population. Read it narrowly: it separates claimed skill from demonstrated skill without telling you which direction any individual leans.

The sharper version comes from a setting where the ground truth was measured. Sixteen experienced open-source developers, working on repositories they had known for about five years, forecast that AI tools would make them 24% faster, believed afterwards that they had been 20% faster, and were measured 19% slower 3. Sixteen people in one setting is not a general result, and the magnitude does not transfer. The sign travels, though: skilled practitioners had the direction of their own performance backwards, on code they knew well.

Practical consequence for the form on your careers page. Keep the self-rating if it helps candidates self-select, and never score it. A five-point scale on "rate your AI proficiency" collects what a person believes about their own skill, and the two studies above are the reason not to treat that as a measure of the skill.

Build the test out of a task the role actually does

Four ingredients, and each earns its place. Real occupational material the role touches. An assistant available rather than banned. A time box short enough to force a decision about what to do yourself. And one thing in the supplied material that is wrong in a way this field would catch. Remove any of the four and the exercise stops measuring what it was built for.

1. Occupational material. Generic puzzles measure puzzle skill. A recruiter grades a drafted requirement against a real role. A financial analyst ties a figure back to a filing. A revenue-cycle analyst reads a denial against the payer's own policy. The material is what makes discernment possible, because judging an answer requires knowing what a normal answer looks like here. 2. The assistant available. Banning it measures work without AI, which is a different job than the one being hired for. It also turns the exercise into a compliance check. 3. A time box. Forty to sixty minutes of work. Long enough to frame the problem and check one thing, short enough that the person has to choose what to keep and what to hand over. Anything longer is unpaid labour. 4. A planted error. The exercise turns on this. Something confidently wrong that the material contains and the model will happily build on. How to plant an error a candidate's field would actually catch is the design problem in full.

Grade the decisions rather than the polish, using the four competencies as the vocabulary and your own material as the standard. Write the rubric before you write the error, or the rubric will be written to fit the first submission you read.

How much does the format buy you?

Less than a vendor slide claims and enough to matter. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42 against .19 for unstructured ones, and work samples at .33 4. Read that pairing carefully: the distance between a structured and an unstructured interview is larger than the distance between an interview and a work sample, so how the exercise is run matters more than which exercise was picked.

The .33 for work samples replaced a figure of .54 that circulated for decades and traces back to a 1974 narrative review, and 53 of the 54 studies behind the newer estimate tested people already doing the job rather than applicants 4. That is a real limit on how much the number can be asked to carry, and it is also why .33 against .42 should not be read as work samples losing. The paper argues the pattern is coherent, and the credibility intervals overlap.

The honest limit sits underneath all of it: none of that validity evidence is about AI-open assessment, because none exists yet. Nobody has published a study showing that scores on an AI-fluency exercise predict performance in an AI-heavy role. What is available is a design argument, a set of adjacent validity estimates, and the discipline of running the same exercise the same way for every candidate and writing down what the work contained rather than how it felt. Detector, structured interview, or work sample sets out what each of the three can actually defend.

See what gets scored

Common questions

Is an AI literacy test worth running at all?

Yes, as a floor rather than a bar. It costs almost nothing, every candidate gets the same one, and a blank result tells you something real. Run it early and weight it lightly. The failure mode to watch is the test becoming the whole screen: a written test cannot separate two candidates who both know the vocabulary, and that is almost always the decision in front of a hiring team. Whoever clears the floor still has to do the work.

Can you assess AI fluency without proctoring?

Yes, and proctoring is largely beside the point once the assistant is allowed. If the assignment turns on catching an error planted in the material, outside help does not remove the decision, it just adds another voice that has to be judged. Screen capture of the work itself, with the candidate told in advance and free to decline, gives a reviewer the process without watching the person. Watching the person is a different thing and it measures composure.

How long should the exercise be?

Forty to sixty minutes of actual work, with the candidate choosing when to do it. That is long enough for framing, one real check and a short written account, and short enough that a working candidate can fit it in. Past about ninety minutes the exercise starts selecting for who has a free evening, which has nothing to do with the job. Pay for anything longer, and say so up front.

What if a candidate declines to use AI in the exercise?

Declining is a supported outcome, and an informative one. Ask what they did instead and what they checked, then read the work on its own terms. Somebody who works under a tool ban, in a regulated environment or on air-gapped systems may have excellent verification habits and no model in the story. A refusal only becomes a problem if the role genuinely requires the tool daily, and that is a conversation rather than a score.

Should the result be a number?

No. Use findings with the evidence attached. A number is easy to compare and impossible to explain: the moment a candidate or a regulator asks why one person advanced, a composite gives you nothing to point at, while a finding quoting the moment it rests on answers the question directly. Findings also survive disagreement better, because two reviewers can argue about what a specific moment showed, where two reviewers arguing about a composite are arguing about arithmetic nobody can inspect.

References

  1. 1. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arXiv (Bogart, Warrier, Agarwal, Higashi, Zhang, Flot, Savelka, Burte, Sakr), 2025. arxiv.org Supports the claim that a scenario task simulating AI use on the job outperformed the knowledge tests the same researchers adopted or wrote.
  2. 2. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the claim that self-reported and demonstrated AI literacy do not substitute: correlations of 0.07 to 0.24 across 288 teachers.
  3. 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that practitioners can be wrong about the direction of their own performance: forecast 24% faster, measured 19% slower.
  4. 4. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the validity figures quoted for structured interviews, unstructured interviews and work samples, and the correction of the older .54 work-sample estimate.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.