Assessment design

How to Run an AI Skills Assessment on the Team You Have

To assess the AI skills of the team you already have, give everyone the same short piece of real occupational work with an AI assistant available, then read what they did rather than what they say they can do. One hour, one task drawn from the job, four behaviors read off the record: was the problem framed before anything was generated, was a source opened, was any output rejected, was a claim tested against something outside the chat. A survey measures confidence. A work sample measures practice.

The takeSelf-rating is the wrong instrument here, and honesty has nothing to do with it. A scale asks people to judge their own judgment, which is the thing the assessment exists to find out, so whatever comes back has to be verified by doing the work anyway. The survey is not a cheaper version of the exercise. It is the exercise's question, asked of the one person who cannot answer it, and then filed as though it had been answered.

Where Olive fits

Open a role and see what the work shows

If you build this in-house, the expensive parts are the answer key and the evidence trail. Olive ships authored cases grounded in one occupation and returns six separately evidenced findings, each anchored to a moment in the session rather than to a number.

Rank your shortlist

Why not just ask people to rate themselves?

Because self-rating and demonstrated ability turn out to be barely the same measurement. Among 288 teachers who took both a self-report and a knowledge-based test of AI literacy built on the same framework, correlations between the two ran from 0.07 to 0.24, and the profiles the researchers found included people who overrated themselves and people who underrated themselves 1. Confidence is not a proxy you can correct for.

A Taiwanese teacher sample is not your finance team, and one study is not a law. What travels is narrower and better evidenced: people are unreliable narrators of their own performance with these tools, sometimes about the direction and not just the size. Sixteen experienced developers in a randomized trial forecast that AI tooling would cut their task time by 24%, estimated afterwards that it had cut it by 20%, and were measured 19% slower 2. Sixteen people working on codebases they had known for years is a small, specific setting. It is still the cleanest demonstration available that the belief and the measurement can disagree about the sign.

Even a scrupulously honest answer misleads you here. The person who accepts the first draft, ships it and hears no complaints has a genuinely good experience of the tool and no evidence that anything went wrong. Asking that person to rate their AI skill collects an accurate report of a good experience.

So collect a record, and make it cheap enough to collect again next year.

What should the exercise ask people to do?

One piece of work the team genuinely produces, done in under an hour with an assistant available and nothing about the assistant prescribed. A quarterly client summary, a vendor comparison, a bug triage, a policy exception memo. The brief carries a real constraint, a real deadline, and at least one claim that would be expensive to get wrong.

Grounding matters more than difficulty. Researchers building an AI literacy assessment for a US Navy robotics training program reported that a scenario task simulating AI use on the job outperformed the tests they had adopted from earlier research or written themselves, though they published no sample size or effect size behind that comparison 3. The general version is unsurprising and still widely ignored: a question about what a language model is tells you nothing about whether someone will check its output at four in the afternoon.

Three things belong in every brief.

  • A constraint the assistant cannot see. An internal policy, a client preference, a number that lives only in your systems. This is what separates specification from transcription.
  • One load-bearing claim. Something the finished piece rests on that a careful person would go and confirm.
  • A plausible trap. A sub-task adjacent to the ones these tools handle well, where the confident answer is wrong.

Keep the brief identical inside a function. Across functions, change the case rather than the format, so the underwriter and the analyst are read on the same four behaviors in their own occupational language. If it is not yet settled which functions need this at all, work out which roles actually need AI skills before building the exercise.

Which behaviors are worth reading off the record?

Four, and they show up in the trail rather than in the finished document. Whether the person wrote down what they wanted before generating anything. Whether they opened a source when a claim carried the decision. Whether they rejected something the assistant produced and said why. Whether they checked an output against something outside the conversation.

Each is a yes-or-no read with an excerpt attached. No scale, no weighting, no total. A reviewer marking the first is looking for a sentence written before the first prompt: the constraint, the audience, what a wrong answer would look like. A reviewer marking the fourth is looking for a moment where the person left the chat window and came back with something the chat window could not have given them.

What stays out of the read matters as much. Prose quality, prompt syntax, tool trivia, how fast the work came in, and how much of the assistant's output survived are all observable and none of them are the skill. Someone who used the assistant for nine tenths of a draft and caught its one wrong claim did better work than someone who typed every word and left the same claim standing. If the instinct is to put a stopwatch on this, read the case on speed against judgment first.

The trail is the whole instrument, which means the exercise has to produce one. Ask for the working record alongside the deliverable: the transcript, the notes, the discarded version. People need to be told that at the start, in writing, along with what happens to it afterwards.

How do you run it without turning it into a rating?

Say what the results will and will not be used for before anyone opens the brief, then hold to it. A capability read is not a performance rating, and the fastest route to honest work is a written promise that no one's answer lands in a compensation file. It carries more weight signed by whoever owns the budget than by whoever is running the exercise.

That promise costs less than it sounds. What a training program needs from this is a picture of the function: how many people framed before generating, how many opened a source, where the practice already sits. Individual records go back to the individual. The aggregate goes to whoever is designing the training.

Three guards worth writing down.

  • Announce the format, the length and the four behaviors a week out. Surprise adds nothing here and costs trust.
  • Offer the accommodation before anyone asks. Extra time, a different input method, a written rather than spoken commentary.
  • Let people decline. A refusal is feedback on the program, and coercing participation contaminates every record collected.

Put two readers on a sample. Have them mark the same ten records independently and compare before either sees the other's marks; where they disagree, the rubric is ambiguous, and fixing that now is cheaper than defending it later. Then fix the format so next year is comparable: same length, same four behaviors, same kind of artifact, a fresh case.

The bar being cleared is low. In SHRM's benchmarking survey, 12% of member organizations reported using work sample interviews for executive candidates, 11% for middle management and 9% for individual contributors, against 79%, 78% and 76% for in-person interviews, all self-reported and collected in 2021 4. Those are hiring figures rather than internal ones, and they are the closest available read on how common a graded work sample is anywhere. An hour of real work with a written rubric behind it is rarer than it sounds. Then use it: build the upskilling program around what the records showed, budget for the reading as carefully as for the courses because training and verification are two different line items, and if the real question is whether one person can move into an AI-heavy job, internal mobility is a separate decision with separate rules.

See what gets scored

Common questions

Can we just add a question to the annual engagement survey?

It will produce a number, and the number will not tell you what people can do. In a study that gave the same people both a self-report and a knowledge test of AI literacy, the two barely tracked each other, and the group included people who overrated themselves and people who underrated themselves, so there is no correction factor to apply afterwards. A survey question is still useful for something narrower: which tools people have access to, which they are blocked from, and where they think the friction is. Keep it for inventory and stop asking it for capability.

How long should the exercise be?

Under an hour, including the written commentary. The constraint is not attention span, it is repeatability: an exercise that takes half a day gets run once, and one run tells you nothing about whether anything changed. An hour is long enough to contain framing, a source check, a rejection and a verification, which are the only four things being read. If the occupational task genuinely cannot be compressed, cut the deliverable and keep the trail.

What if someone refuses to use AI for the task?

Accept it and record it as a result, not a failure. A refusal on a specific task can be the correct professional call, and the reason is worth more than the compliance would have been. Ask for the reason in writing. Blanket refusal across every task is a different signal, and it belongs in a conversation with a manager. Coercing participation gives you a record that nobody trusts, including you.

Should managers grade their own reports?

Not on the first run. A manager reading their own team introduces the exact contamination the exercise is trying to avoid, and it turns a capability read into an appraisal in everyone's mind whatever the memo said. Use a reader from a neighboring function who knows the work, or pair two managers and have each read the other's team. Once the format has run twice and the promise has visibly held, managers reading their own people is fine.

Does this replace the skills matrix the HRIS already produces?

The matrix and the exercise answer different questions. An inferred skills profile built from job titles, course records and project tags is a reasonable map of what people have been exposed to. It cannot tell you what someone does when an assistant produces a confident wrong answer, because nothing in the input data records that. Keep the matrix for planning and coverage. Use the exercise for the four behaviors, which is the part the matrix cannot see.

How often should this run?

Once a year is enough for most functions, with a fresh case each time and the format held constant. More often than that and you are measuring familiarity with the exercise rather than capability. Less often and the models underneath have moved too far for the comparison to mean much. Run it out of cycle only when something specific changed: a new tool rolled out, a function restructured, a policy that altered what people are allowed to do.

References

  1. 1. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the claim that self-rated AI skill and demonstrated AI skill are weakly related, with error running in both directions.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that people's estimates of their own AI-assisted speed can be wrong about the direction, not only the magnitude.
  3. 3. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arXiv (Bogart, Warrier, Agarwal, Higashi, Zhang, Flot, Savelka, Burte, Sakr), 2025. arxiv.org Supports the design argument that a realistic occupational scenario reads applied AI skill better than an abstract knowledge test.
  4. 4. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the claim that work-sample exercises remain a minority practice compared with interviews.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.