Assessment design

One AI-Skills Assessment for Every Department, or Per-Role?

An AI-skills assessment is one rubric everywhere and one case per role. The behaviors you score (what got framed before anything was generated, what got checked outside the chat) hold in every department, but two results compare only when rows, outcome words and seniority match. The case cannot travel: an assessment is only as job-related as the work it samples. The exception: two roles a job analysis shows do substantially the same work on the same material can share a case. The rubric still won't give you a company-wide number.

The takeThe pull toward one instrument is a procurement instinct, and it is right about exactly one layer. Standardizing the language buys a bar; standardizing the task buys a number that means nothing, which is often what the mandate wanted, because a percentage is usually the only thing that survives the trip to a board. These programs die of maintenance far more often than of law, I suspect: nobody budgets for the answer key that ages, and a case nobody has rewritten in a year measures who has seen it before. The rubric is the cheap half. The case is the job.

Where Olive fits

Open a role and see what the work shows

If you build the two layers yourself, the recurring cost is an answer key and an evidence trail for every occupation you add. Olive ships twelve authored cases per occupation against one fixed set of six dimensions (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), with every finding written by a human reviewer, anchored to a moment in the session, and granted to the candidate in the same document.

Rank your shortlist

What has to be identical, and what has to change?

The rubric is identical; the task is not. Score the same behaviors in every department: what got framed before anything was generated, which claim got a source demanded, what the candidate kept for themselves, what existed between the brief and the deliverable, what was refused and on what grounds, and what was checked against something outside the chat. Draw the work those behaviors happen inside from each role's own material.

That split is not a compromise between two positions. It is the only arrangement where both halves of the question get a true answer: the shared rows are what let a marketing result and an engineering result talk to each other, and the role-drawn case is what makes either result mean anything about that job.

The two failure modes are easy to name once the split is clear. One instrument with one generic task (write a prompt, spot the hallucination, summarize this article) produces perfectly comparable results about nothing anyone does at work. Six departments each writing their own rubric produces six defensible assessments and no bar: the finance rubric says rigorous, the design rubric says iterative, and nobody can say what a good result means across the company.

Executives usually reach this question through procurement rather than psychometrics: one contract, one rollout, one number for the board. Keep that instinct for the rubric, the outcome words, the disclosure, the time cap and the review protocol. Spend the per-role money where it actually buys validity, which is the case itself. What an AI assessment should measure is the same list in every department; what it should ask people to do is not.

Why can't one task cover every department?

Because a selection procedure holds up only as far as it samples the content of the job it is used for. The Uniform Guidelines on Employee Selection Procedures say a procedure can be supported by content validity "to the extent that it is a representative sample of the content of the job," and that where a test purports to sample a work behavior, its manner, setting, level and complexity "should closely approximate the work situation" 1.

A generic AI exercise approximates nobody's work situation. Ask a revenue cycle specialist and a backend engineer to summarize a press release with an assistant open, and you have measured neither job. You have measured reading speed and prose taste on a task with no correct answer either of them would recognize.

The Guidelines go further than most buyers expect. Where the observed work behaviors and the observed work products in two jobs are not the same, the federal enforcement agencies "will presume that the work behavior(s) in each job are different" 1. Difference is the default; sameness is the thing you have to show. And to carry validity evidence from one job to another at all, incumbents in both must perform "substantially the same major work behaviors, as shown by appropriate job analyses" on each job 2.

That is a design constraint, not a legal footnote. A step that decides who advances is a selection procedure, and it has to be job-related and consistent with business necessity. Work samples and simulations are named examples of exactly that 3. A single company-wide AI quiz used as a gate is the one design with no job analysis behind it, no work sample in it, and no answer to which job it was validated for. The choice between a detector, an interview and a work sample turns on the same point: only one of the three has content you can trace back to a job.

How do you keep results comparable across departments?

Fix three things and the comparison holds: the same rubric rows, the same outcome words, and cases at a comparable level of complexity. The rows are cross-occupational by design. The O*NET content model is built on the same idea, pairing descriptors that apply across many jobs and industries with tasks specific to one occupation 4. What changes per role is the anchor beneath each row: the sentence saying what satisfying it looks like in that work.

Anchors are where most shared rubrics quietly become six different rubrics. "Demanded evidence for the claim that mattered" is one row everywhere. In financial analysis it is satisfied by opening the filing behind the number the memo rests on; in software engineering, by running the failing test before accepting a patch; in legal operations, by pulling the signed precedent instead of the playbook's summary of it. Same row, same outcome words, different observable. Write the anchors with someone who does the job, and write them before the first candidate rather than after the third.

Complexity is the second thing that breaks a comparison, and the one nobody checks. The Guidelines allow a group of jobs to be treated together only where they share critical work behaviors "at a comparable level of complexity" 1. A 50-minute case for a first-year analyst and a 50-minute case for a director are not the same instrument even with identical rows, so hold seniority constant inside a comparison or stop calling it one.

Third, keep the outcome language coarse and verbal: demonstrated, partly demonstrated, not demonstrated, per row. A number invites averaging, and an average across six behaviors is a claim about a person that none of the six rows supports. Two reviewers reading the same session should land on the same three words; when they don't, the rubric's reliability is the problem rather than the candidate. That agreement is also what a defensible proficiency bar rests on, because the bar is a statement about rows, not a cut on a total.

Build the shared rubric once, the case per role

Split the build in two and the cost stops scaling with headcount. Written once, for the whole company: the six rows, the anchor format, the outcome words, the time cap, the disclosure a candidate reads before starting, the accommodation path, and the protocol two reviewers follow. Written per role: the case, its source packet, the answer key, and the anchors under each row.

The case is where the local knowledge goes, and it needs one thing the model cannot have: a constraint that lives in your material rather than in the brief. That device is the same one that makes an AI-open interview produce signal: the assistant answers fluently and confidently, and the gap between that answer and the constraint is the assessment.

RoleThe caseWhat the assistant gets confidently wrongThe act that catches it
Financial analysisA diligence memo over a source packetRepeats a growth figure the filing's footnote contradictsThe footnote is opened and the number changes
Software engineeringOne failing integration test in a real branchPatches the code under test rather than the shared fixtureThe suite is run before the diff is written
MarketingA positioning brief from vendor researchQuotes the most on-message statistic in the packetThe statistic's source is opened, and it is dropped
Revenue cycleA queue of denials to workDrafts a persuasive appeal for the claim that should be concededThe payer policy is read against the code

Cost, honestly: about a day per case including the answer key, plus a dry run against two people already doing that job, plus a rewrite after the first three candidates, because the first version is either guessable from the brief or undiscoverable inside the time limit. That figure belongs in the budget conversation, and it recurs: cases leak, and a case everyone has seen measures preparation.

Sequence by where AI actually changed the work rather than by org chart. Two or three cases is a program; twelve is a project that ships next year. Start where a bad hire is expensive and the AI use is already heavy, run the case beside your existing round without letting it decide anything, and compare the two on real candidates first. Piloting before it becomes a gate is the cheapest way to find out the case is wrong.

When does one assessment legitimately cover two roles?

When a job analysis shows the two roles perform substantially the same major work behaviors on substantially the same kind of material 2. That happens more often than the org chart suggests and far less often than a buyer hopes. Compare the observed work behaviors and the observed work products in each job; where neither matches, treat the jobs as different, because that is the presumption you would be arguing against 1.

The grouping that survives is by work, not by department. A pricing analyst in finance and a sizing analyst in strategy both read a packet whose claims settle only when someone opens them, and one case can serve both. A designer and a data scientist share a reporting line and nothing else. The test is whether the deliverable and the material would be recognizable to someone in the other role, not whether both roles use the same assistant.

Two limits are worth saying out loud before the rollout. A shared rubric across departments does not produce a company-wide number and should not be sold as one: six findings on one person, read by the manager who owns that role, is the whole output, and what it makes comparable is kinds of evidence rather than people on different teams. And the maintenance never consolidates: the rows stay written, the cases keep aging, and every new occupation is another answer key someone has to author and keep current.

Buying rather than building becomes the honest answer somewhere around the third role, and the trade is the same either way: a vendor's case is grounded in an occupation rather than in your company, so it carries the role's work behaviors and not your private constraint. Olive is one instrument in that category and says so; multiple-choice AI literacy tests, code-collaboration graders and unwatched take-homes are all real approaches with different trade-offs. Build or buy turns mostly on how many roles you genuinely need and whether anyone in the building can write an answer key.

See what gets scored

Common questions

Can one rubric really score a marketer and an engineer on the same terms?

Yes, because the rows describe behaviors rather than outputs. "Tested a claim against something outside the conversation" is observable in both jobs; only the observable differs: a recomputed figure in one, a run test suite in the other. What cannot be shared is the anchor sentence under each row, which has to be written with someone who does that job. Get the anchors wrong and you have six rubrics wearing one name, which is worse than admitting you built six.

Does one company-wide AI assessment create legal exposure?

The exposure is job-relatedness, not the number of instruments. A step that decides who advances has to be job-related and consistent with business necessity for the job it is used on, and a generic quiz has no job analysis behind it to show that. A shared rubric is fine; a shared task across unlike jobs is the part that is hard to defend. Keep the job analysis, the case, the answer key and the reviewer notes for each role. That file is the defense, and it is per role.

How many role-specific cases do you need to start?

Two or three, chosen where AI already changed the work and where a bad hire is expensive. A program with three live cases and a working review protocol beats a twelve-role plan that ships next year, and the first three teach you what the rubric anchors should have said. Add roles once two reviewers agree on outcomes without discussion, because reviewer disagreement multiplies with every case you add.

What about roles where AI barely touches the work?

Do not assess them. An AI-skills assessment used on a job where the assistant plays no part is a step that cannot be tied to the work, which is both indefensible and a waste of the candidate's hour. Decide role by role whether the assistant is in the daily path, and write down the reason either way. The list of roles that need this is usually shorter than the mandate assumes.

Should the same reviewer read every department's sessions?

One reviewer per role reads with more context; one reviewer across roles keeps the rows consistent. Do both: assign by occupation, then double-mark a sample across departments and check the outcomes match. Where they diverge, the anchor is ambiguous and gets rewritten. Calibration is the maintenance cost nobody budgets, and it is the first thing to break as the number of roles grows.

References

  1. 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job; the setting, level and complexity should closely approximate the work situation; jobs grouped together need common work behaviors at a comparable level of complexity; where observed work behaviors and products differ, the agencies presume the jobs' behaviors are different.
  2. 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.7 (Use of other validity studies) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Validity evidence may be borrowed for another job only where incumbents perform substantially the same major work behaviors, shown by appropriate job analyses on both jobs.
  3. 3. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Work samples and simulations are selection procedures, and a selection procedure must be job-related and consistent with business necessity.
  4. 4. The O*NET Content Model O*NET Resource Center, U.S. Department of Labor Employment and Training Administration, 2026. onetcenter.org The content model pairs descriptors usable across many jobs and industries with information specific to particular occupations, including occupation-specific tasks.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.