Assessment design

The Resume Screen Is Regulated AI. A Human-Scored Work Sample Usually Is Not.

A human-scored work sample is usually legally safer than an AI resume screen, and not because assessments are gentler: the rules turn on automated versus human-decided and job-related versus proxy. An automated screen that orders or cuts candidates falls inside California's automated-decision system definition and often inside New York City's tool test. A job-analysed exercise administered to everyone and scored by a named person against a written rubric falls outside that trigger and produces the evidence a disparate-impact defense needs.

The takeThe received wisdom, earned over decades of cognitive-test litigation, is that testing is where employers get sued and resume screening is the quiet part. That was true when the test was a proctored aptitude battery and the screen was a person reading paper. It is now backwards. The stage with the audit duty, the notice duty and the named regulatory definition is the one nobody thinks of as an assessment, and the stage everyone fears has the better paper trail.

Where Olive fits

Open a role and see what the work shows

Building this in-house, the expensive parts are the answer key and the evidence trail. Olive ships item banks grounded in twelve occupations, each carrying its SOC code, and returns six findings whose wording a human reviewer writes against a timestamped excerpt from the session.

Rank your shortlist

Which stage is actually regulated?

The automated one, on both coasts. New York City's Local Law 144, in force since January 1, 2023, attaches where a tool substantially assists or replaces discretionary decision-making 1. California's regulations, effective October 1, 2025, define an automated-decision system as a computational process that makes or facilitates an employment decision, and name screening resumes for particular terms or patterns as an example 28. A resume screen that orders a queue or cuts the bottom meets both descriptions comfortably.

A human-scored exercise struggles to meet either. No computational process concludes anything that assists the decision, and the discretion sits with a reviewer forming a judgment from the work itself. The trigger fails at the first element rather than the last, which is a much stronger place to fail it.

Be precise about what carries that conclusion. It is not the word assessment, and it is not the presence of a human somewhere in the room. It is that no software concluded anything about the person. A take-home graded by an auto-scorer, an exercise ranked by a model, or a rubric where a tool fills in four of six lines puts you back inside the definitions, whatever the stage is called.

And a tool outside a city rule is not outside federal law. The Uniform Guidelines treat any measure used as a basis for an employment decision as a selection procedure, reaching everything from paper-and-pencil tests through informal or casual interviews and unscored application forms 3. Both stages are selection procedures. Only one of them is also a regulated automated tool.

What makes a human-scored work sample hold up?

A documented job analysis, and content that samples the job rather than standing in for it. The federal standard for a content-validity study asks you to identify the important work behaviours, show that the procedure is a representative sample of those behaviours or of the work products, and operationally define any knowledge, skill or ability as necessary to critical duties 3. That is a design brief as much as a legal test.

The same regulation draws two limits worth knowing before you build. A procedure resting on inferences about mental processes cannot be supported primarily on content validity, and content validity is not an appropriate strategy for knowledge, skills or abilities an employee will be expected to learn on the job 3. So an exercise that asks a candidate to do a real task passes the test on its face. One that asks a candidate to demonstrate judgment in the abstract does not.

Four build decisions carry most of the defensibility:

1. Write the job analysis first, and keep it. The document that says which behaviours matter is the one that makes everything downstream job-related instead of assumed. 2. Give it to everyone at the same stage. Selective administration reintroduces the discretion the design was meant to remove. 3. Name the scorer and the rubric. Who applied which rubric to which submission is what a challenge asks about, and a general impression cannot answer it. 4. Score the work, not the presentation. Prose polish, speed and confidence are proxies, and proxies are what the whole analysis exists to catch.

If the case for moving budget out of the screen and into the exercise is what you are actually weighing, whether to drop the resume screen entirely works that decision on funnel terms, not legal ones.

Where a work sample can still get you sued

Everywhere adverse impact lives, which is everywhere. Nothing about human scoring exempts an exercise from Title VII: a practice causing a disparate impact is unlawful unless the employer shows it is job related for the position in question and consistent with business necessity, and a plaintiff can also win by naming a less-discriminatory alternative the employer refuses to adopt 4. A work sample can produce impact as readily as a test can.

Three exposures specific to exercises, none of which the AEDT analysis touches:

  • Accommodations. It is unlawful to fail to select and administer tests in the most effective manner to ensure results reflect the skill being measured rather than an applicant's impaired sensory, manual or speaking skills 5. A timer, a screen-share requirement or a single input method can each turn into a screen-out. Offer accommodations in the invitation rather than waiting to be asked.
  • Unpaid burden. A multi-hour exercise selects on who can afford the hours. That is not a named legal test, and it maps onto protected characteristics in practice.
  • Automation creep. The moment any part of scoring becomes automatic, the classification changes and the audit and notice questions come back. This is how a compliant design drifts out of compliance without a decision.

One honesty note about the evidence, because the sales version of this argument overstates it. The often-quoted claim that work samples predict job performance at .54 is superseded; the current estimate is .33, and most of the underlying studies tested people already doing the job 6. A work sample is a good instrument with a documented job link. It is not a guarantee of a better hire, and claiming otherwise in a hearing is worse than claiming nothing.

On the classification itself, whether a validated assessment differs from a bias-audited one is the distinction that decides which evidence you are actually holding. Confirm any of this with counsel before relying on it.

Compare the two stages on five points before moving budget

Run both stages down the same five questions and the decision usually makes itself. This is a comparison a founder can do in an hour with the current process open in one window, and it is more useful than a vendor's compliance page because it is about your configuration rather than the category.

1. What does software conclude about the person? Screen: an order, a category or a cut. Exercise: nothing, if it is scored by a person. 2. Can you point at a job analysis? Screen: rarely, because the criteria accumulated. Exercise: yes, if you wrote one, and it is the single highest-value artifact in either stage. 3. What record exists a year later? Screen: a vendor's logs, which you may not be able to obtain. Exercise: the submission, the rubric and the named reviewer's notes. 4. What notice do you owe? Screen: potentially a bias audit, a posted summary and advance notice under New York City's rule 1, plus a notice duty in Illinois since January 1, 2026 7. Exercise: generally none of those, though accommodations still have to be offered. 5. What happens when it drifts? Screen: a vendor toggle changes the answer. Exercise: an auto-scorer added for throughput changes the answer.

What the comparison does not do is settle whether the exercise is any good. A defensible stage that measures the wrong thing is still the wrong stage, and the design question is separate from the legal one. What holds up if a rejected candidate challenges an AI-skills assessment covers the version of that argument you would actually have to make.

Neither stage is a place to run a tool you cannot describe. If you cannot say in a sentence what comes out of it and who acts on it, that is the finding, and it applies to the exercise as much as to the screen.

See what gets scored

Common questions

Does adding a human reviewer to an AI screen take it out of scope?

Sometimes, and less often than teams assume. New York City's test asks whether the output substantially assists or replaces discretion, and a reviewer who opens a ranked list and works down it is being assisted substantially. What defeats the trigger is independent review of every candidate, where the person forms a view from the underlying material. California's definition is wider still, reaching a computational process that facilitates a decision, so a human in the loop does less work there. Classify the behaviour, not the org chart.

Is a take-home assignment the same as a work sample?

Not for evidence purposes. The validity research on work samples largely tested people already doing the job on job-shaped tasks, and says nothing about unpaid multi-hour assignments completed at home. A take-home can be built as a work sample by keeping it short, tying it to a documented job analysis, giving it to everyone at the same stage and scoring it against a written rubric. Without those, it is an exercise with a legal profile nobody has established.

Do we need a bias audit for a human-scored exercise?

Not under New York City's rule, which attaches to automated employment decision tools rather than to assessments generally. That is not the same as being free of measurement duties. Adverse impact analysis applies to any selection procedure, and running a stage for a year without ever looking at outcomes by group is its own risk. Audit the outcomes because you want to know, not because a city told you to.

What if our assessment vendor scores the submissions?

Then ask precisely what does the scoring. A vendor employing human graders against your rubric leaves the analysis roughly where it was. A vendor whose platform produces a number, a band or an ordering has put a computational process between the candidate and the decision, which brings the definitions back into play. Get the answer in writing, including what happens on the fast path when volume spikes, because that is usually where the automatic scorer lives.

Does it help to run both stages?

Yes, though the first stage keeps every obligation it already had. Adding an exercise after an automated screen leaves the screen exactly as regulated as it was, and the candidates who never reached the exercise were still cut by the tool. If the goal is to reduce exposure rather than to add a stage, the change that does it is narrowing what the automated step is allowed to decide, then giving the exercise to everyone who clears a stated bar.

How long should a defensible work sample take a candidate?

Short enough that the burden is not itself a filter, and long enough to sample real work. Somewhere between forty minutes and two hours covers most roles, and anything past that needs a reason you would be willing to state in writing. Pay for longer exercises. The legal analysis does not turn on duration, but duration drives who can participate, and who can participate is what adverse impact measures.

References

  1. 1. Automated Employment Decision Tools: Frequently Asked Questions NYC Department of Consumer and Worker Protection (DCWP), 2023. nyc.gov Supports the substantially-assists-or-replaces standard that an automated resume screen tends to meet and a human-scored exercise tends not to, and the audit, posted summary and advance notice duties that follow.
  2. 2. Final Unmodified Text of Proposed Employment Regulations Regarding Automated-Decision Systems (Attachment B), 2 CCR sections 11008, 11008.1 California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov Supports the automated-decision system definition and the fact that screening resumes for particular terms or patterns is named in the regulation as an example.
  3. 3. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q), 1607.3(A) and 1607.14(C) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the selection-procedure definition covering both stages, and the content-validity standards: job analysis, representative sampling of work behaviours, and the limits on mental processes and on skills learned on the job.
  4. 4. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases Office of the Law Revision Counsel, United States Code (prelim), 1991. uscode.house.gov Supports the job-related and consistent-with-business-necessity standard and the less-discriminatory-alternative route, both of which apply to a work sample.
  5. 5. 29 CFR 1630.11 - Administration of tests (Regulations to Implement the Equal Employment Provisions of the Americans with Disabilities Act) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 2023. govinfo.gov Supports the duty to administer tests so results reflect the skill measured rather than an applicant's impaired sensory, manual or speaking skills.
  6. 6. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the correction of the work-sample validity estimate from .54 to .33 and the point that most underlying studies tested people already in the job.
  7. 7. HB3773 Enrolled (Public Act 103-0804), amending the Illinois Human Rights Act Illinois General Assembly, 2024. ilga.gov Supports the January 1, 2026 Illinois notice duty when AI is used in recruitment or hiring, cited as the state-level notice obligation an automated screen can carry and a human-scored exercise does not.
  8. 8. Rulemaking Actions - Civil Rights Council California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov The Council's own record of the automated-decision-system employment regulations: approved by OAL and filed with the Secretary of State, effective October 1, 2025.

8 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.