Assessment design
How Do You Set a Defensible AI Bar for a Specific Role?
The number is not what makes a role's AI-proficiency bar defensible. It holds up when it comes from the occupation's own task list, sits at the minimum the work requires rather than the level you find impressive, and exists on paper before anyone is measured against it. Federal guidelines ask that a cutoff match normal expectations of acceptable proficiency in the workforce doing that job. Judge named behaviors pass or fail, rank by score only with a study showing higher scores predict better work, and re-check when the job changes.
The takeNotice what the documentation requirement really does. Writing the bar down forces you to say, in the words of the work, what you are screening for. Write that sentence honestly and the bar has probably been set at fluency all along: whoever sounds most at home with the tool clears a line nobody wrote. And fluency, I'd guess, is the part of any bar with the shortest shelf life, written as it is against an interface somebody is already redesigning. Interfaces turn over. The judgment about when to believe an answer does not.
Where Olive fits
Open a role and see what the work shows
Setting the level is the easy half; the half that has to survive a challenge is the evidence behind each judgment. Olive's item banks are authored per occupation and carry its SOC code, and a session returns six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a timestamped moment rather than to a number, with the candidate granted the same report.
Rank your shortlistWhat makes a bar defensible rather than just written down?
Three things, and none of them is the number itself. The bar has to come from a job analysis of the occupation you are hiring for, it has to sit at the level the work actually requires rather than the level you find impressive, and the method you used to set it has to exist on paper before a candidate is measured against it. That last one is where most internal bars fail.
The legal frame is not exotic. A threshold that decides who advances is a selection procedure under the Uniform Guidelines on Employee Selection Procedures, and where cutoff scores are used they "should normally be set so as to be reasonable and consistent with normal expectations of acceptable proficiency within the work force" 1. Read that phrase slowly: the workforce doing that job. Not the vendor's benchmark cohort, not the two people on your team who are unusually good at this, and not a number chosen because it makes the shortlist come out at ten.
The EEOC's standard for any test used in hiring is that it be job-related for the position in question and consistent with business necessity, which it explains as necessary to the safe and efficient performance of the job 2. A bar set above what the job needs is the classic failure. It screens people out over a gap that has no consequence in the work, and the gap is usually the part that was easiest to measure.
The Guidelines also ask for the paperwork, and mark it essential. Where a selection procedure is used with a cutoff score, the user should describe "the way in which normal expectations of proficiency within the work force were determined and the way in which the cutoff score was determined" 1. If you cannot produce that description, you do not have a bar. You have a preference with a number attached.
Start from the occupation's tasks, not from a tool list
Write the bar against work behaviors. A job analysis for content validity "should focus on the work behavior(s) and the tasks associated with them" 1, and a tool inventory describes neither. "Has used ChatGPT" names a product. "Reconciles the figure against the source before it goes in the memo" names a behavior, and the behavior is the thing whose absence you can point at afterwards.
The same section carries a trap worth knowing about. Content validity is not an appropriate strategy "when the selection procedure involves knowledges, skills, or abilities which an employee will be expected to learn on the job" 1. Most tool familiarity fails that outright, because a competent analyst learns your assistant's interface in a week. What does not arrive in a week is the judgment about when to believe its output, which is why a durable bar is written about framing and verification rather than about tool hours.
You do not have to start from a blank page. O*NET publishes tasks, detailed work activities and software skills per occupation, keyed to SOC codes, across 923 data-collection-level occupations; 891 of them were updated year to date through May 2026, and the next database release is scheduled for August 2026 6. Pull the task list for the SOC code on the requisition, mark the tasks an assistant now touches, and write the bar against those.
Then check the list against the job you actually run, because a published occupation record is a floor rather than a description of your team. Finding out what AI does in the role before the posting goes out is the same exercise from the other end, and the requirement you publish should match the bar you screen against, since a vague AI-skills line in a job description is hard to defend for the same reason a vague bar is.
Set the level at minimum acceptable, not at impressive
Minimum acceptable is the standard the regulation names, and every step above it is exposure you took on voluntarily 1. Make the bar pass/fail unless you can show that higher scores go with better work: using a procedure for ranking calls for evidence "that persons who receive higher scores on the procedure are likely to perform better on the job" 3. Few employers have that study for AI collaboration.
So write the bar as a short list of behaviors, each judged pass or fail against something a reviewer can point at:
- Framed the problem before generating. The opening move states what is being decided and what would make an answer wrong.
- Demanded a source for the claim the answer turns on, and opened it, rather than asking for sources in general.
- Kept at least one judgment out of the assistant's hands, deliberately, with the reason visible.
- Tested one claim against something outside the tool and changed a number, a recommendation or a stated limit when it did not hold.
Each line is observable in a work sample and each fails cleanly, which is what makes two reviewers agree. Writing an AI-use rubric two reviewers apply the same way is the same problem one level down, and the follow-up question is what separates a real answer from a fluent one.
Then measure what the bar does to your pipeline. A selection rate for any race, sex or ethnic group below four-fifths of the rate for the highest group is generally regarded by the federal enforcement agencies as evidence of adverse impact 1. And where two procedures serve the same legitimate interest and are substantially equally valid, the Guidelines tell you to use the one with less adverse impact 1. A bar you can move down without losing the signal is a bar you should move down.
Why does the bar move, and how often should you re-set it?
Because the work moves, and the regulation expects you to notice. On currency, the Guidelines say plainly that "there are no absolutes," and that changes in the relevant labor market and in the job itself decide when a validity study is outdated 1. A bar written against a 2024 task list is resting on assumptions about the job that have already shifted underneath it.
The shift is measurable. Across eight months of one large usage sample mapped to O*NET occupational categories, conversations where the user delegated a complete task rose from 27% to 39%; educational tasks rose from 9.3% to 12.4% of the sample while business and financial operations tasks fell from 6% to 3%; and inside coding, creating new code rose from 4.1% to 8.6% while debugging fell from 16.1% to 13.3% 4. That is one vendor's traffic rather than an occupational census, and it is still the shape of the problem: what people hand to an assistant changes faster than a hiring rubric gets revised.
Base rates differ by occupation as well. In late 2024, 23% of employed US respondents had used generative AI for work at least once in the previous week and 9% used it every workday, with between 1 and 5% of all work hours assisted; writing communications topped the list of tasks respondents named it most useful for, at 39.5% 5. A bar for a marketing role and a bar for supply chain rest on different tasks, and neither is anchored by what an executive read last quarter.
So give the bar a review date the way you gave it a level. Re-check when the requisition reopens, at minimum once a year, and log what changed and what you did about it. A record of a bar that was reviewed and deliberately left alone is worth as much as a record of one that moved. Which roles genuinely need an AI bar at all is worth re-asking on the same schedule.
Write the bar so someone else can check it
One page per role, finished before the first candidate is assessed. It names the occupation, the tasks the bar rests on, the behaviors it looks for, what counts as evidence for each, the level, who set it, when it gets reviewed, and what you considered and rejected. A candidate's lawyer, an agency investigator and your own successor all end up reading that page.
- Occupation and SOC code, plus where the task list came from and the date you pulled it 6.
- The behaviors, in the words of the work, with the evidence that counts for each.
- The level and its reasoning: why this is the minimum the job requires, in the terms the Guidelines ask for 1.
- Pass/fail, or the study behind ranking if you are ordering candidates at all 3.
- Who set it, when, and the review date.
- Selection rates by race, sex and ethnic group for the step the bar sits in 1.
- The alternatives weighed, and why the one you chose was not the one with more adverse impact 1.
Two of those matter more than the rest once someone challenges the decision. Running the adverse-impact arithmetic on a screening step is not hard, and not having run it is itself the finding. And keeping the alternatives you weighed answers the question a plaintiff's expert asks first: whether a detector, a structured interview or a work sample is the thing actually carrying the decision.
Date every one of these on the day you write it. A record assembled after a challenge arrives reads as a reconstruction, because that is what it is.
Common questions
What does 'good enough at AI' mean for a specific role?
It means carrying the tasks the occupation actually performs with an assistant in the loop, at the level the job requires: framing the problem before generating, demanding a source for the claim the answer turns on, keeping the judgment that should not be delegated, and testing a claim against something outside the tool. The content differs by occupation (a paralegal's version and a data analyst's version rest on different task lists), which is why the definition is written per role, from that role's own tasks.
Can a percentage score from an AI-skills test be the bar?
Only if you can say what the percentage means in terms of the job. Where a cutoff score is used, the Guidelines ask you to describe how normal expectations of proficiency in the workforce were determined and how the cutoff was set 1, and a vendor's percentile against its own test-taker pool answers neither. If the score is all you have, use it as pass/fail at the level the work requires and keep the evidence behind each judgment. Do not order candidates by it without a study showing higher scores go with better performance 3.
Should the bar differ by seniority in the same role?
Usually, because the tasks differ. Write each level against its own work behaviors rather than scaling one bar by a multiplier: a junior analyst is expected to check the figure, a senior one is expected to decide which figure the recommendation turns on. Same instrument, different pass conditions, both documented. What does not vary by level is the evidence requirement: every bar names what a reviewer has to be able to point at.
Is 'must use AI daily' a defensible requirement?
It is hard to defend, because frequency is not performance. Daily use describes a habit, and the job-related question is whether the work comes out right. It also sets a threshold on something most people are expected to pick up after hiring, and content validity is not an appropriate strategy for a skill an employee will learn on the job 1. Write the requirement against a task the role performs, then test the judgment inside that task.
How often should you re-set the bar?
When the requisition reopens, and at minimum once a year. The Guidelines say there are no absolutes on currency, and that changes in the labor market and the job decide when a study is outdated 1. There is real movement to track: over eight months in one large usage sample, conversations where the user delegated a complete task rose from 27% to 39% 4. Log the review date even in the years nothing changes, because an unlogged review is indistinguishable from no review.
What records show the bar was set properly?
Four. The job analysis the behaviors came from; the description of how the level was chosen, which the Guidelines mark essential wherever a cutoff score is used 1; the selection rates by race, sex and ethnic group for that step 1; and the alternatives considered, since a procedure with less adverse impact and substantially equal validity is the one you are expected to use 1. Date each of them. Records created after a challenge arrives carry very little weight.
References
- 1. Uniform Guidelines on Employee Selection Procedures (1978), 29 CFR Part 1607 govinfo.gov Sec. 1607.5(H) sets the cutoff-score rule; 1607.15 makes describing how the cutoff and the workforce expectations were determined essential; 1607.5(K) covers currency; 1607.14(C)(1) rules out content validity for skills learned on the job; 1607.14(C)(2) requires job analysis focused on work behaviors; 1607.3(B) prefers the substantially equally valid procedure with less adverse impact; 1607.4(D) states the four-fifths rule.
- 2. Employment Tests and Selection Procedures eeoc.gov States the job-related and consistent with business necessity standard and explains it as necessary to the safe and efficient performance of the job.
- 3. Questions and Answers to Clarify and Provide a Common Interpretation of the Uniform Guidelines on Employee Selection Procedures eeoc.gov Question 62 sets what ranking requires: evidence that persons who receive higher scores are likely to perform better on the job, and the conditions under which pass/fail use is the appropriate choice.
- 4. Anthropic Economic Index report: Uneven geographic and enterprise AI adoption anthropic.com Across eight months of Claude usage mapped to O*NET occupational categories: directive conversations 27% to 39%, educational tasks 9.3% to 12.4%, business and financial operations 6% to 3%, new code creation 4.1% to 8.6%, debugging 16.1% to 13.3%.
- 5. The Rapid Adoption of Generative AI (NBER Working Paper 32966) nber.org 23% of employed respondents used generative AI for work at least once in the previous week, 9% every workday, 1 to 5% of all work hours assisted; Figure 9 puts writing communications top of the tasks it was named most useful for, at 39.5%.
- 6. O*NET Occupation Update Summary onetcenter.org 923 data-collection-level occupations; 891 updated year to date through May 2026; next database update scheduled for August 2026; updates cover tasks, detailed work activities and software skills.
6 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.