Policy
Vary the Data, Never the Task or the Bar
Every candidate for the same role at the same stage gets the same task, the same materials, the same stated time and the same bar. You may vary the inputs: a different dataset from the same generator, a different customer inside the same scenario. Varying the task or the standard gives you a second selection procedure applied to some people and not others, which is the shape a challenge attacks first. Accommodations are a separate matter and never a variation you chose.
The takeTwo pieces of good advice collide here and almost nobody says which one wins. Consistency guidance tells you to run the same assessment for everyone. Anti-leak advice tells you to rotate, because briefs circulate. Rotation wins only when it is scheduled ahead of time, applied to everyone inside a window, and calibrated so the versions are equivalent. A brief swapped for one candidate who complained is not rotation. It is an exception, and exceptions are what a challenge looks for first.
Where Olive fits
Open a role and see what the work shows
Uniform treatment is easier to show when the record is evidence rather than a cutoff. Olive returns six separately evidenced findings written by a human reviewer, stated as demonstrated, partly demonstrated or not demonstrated, and every released report exports with its rubric and bank versions attached.
Rank your shortlistDoes everyone have to get the same exercise?
Yes, for candidates at the same stage of the same role. A work sample is a selection procedure in the federal sense: the Uniform Guidelines define one as any measure or procedure used as a basis for an employment decision, and say the term covers the full range of assessment techniques, from paper tests through informal or casual interviews and unscored application forms 1. On that definition, two versions of it are two procedures.
Two qualifications keep that from being a slogan. The Guidelines impose a validation burden only where adverse impact appears, and they are a 1978 regulation applied by analogy to tools nobody had then. The statutory test sits elsewhere: under 42 U.S.C. 2000e-2(k), once a complaining party shows a particular practice causes disparate impact, the employer has to demonstrate the practice is job related for the position in question and consistent with business necessity 2.
That is Title VII, so it reaches race, color, religion, sex and national origin; age runs under the ADEA and disability under the ADA, on different standards. None of it is legal advice and all of it belongs in front of counsel before it becomes policy. What it establishes for a hiring team is the shape of the question you will be asked, which is normally about a particular practice applied to particular people.
An exercise varied per candidate answers that question badly. The practice is now several practices, the comparison between two finalists rests on nobody's calibration, and the reason one person got a different brief is whatever the recruiter remembers. Consistency is the cheaper posture, and it is also what makes your own data readable when you ask whether the assessment would hold up under challenge.
Which parts can you vary safely?
Inputs, freely. Task, standard and stated time, not at all. Run any proposed change through one question: is a reviewer scoring two submissions applying the same rubric to the same demand? A different dataset from the same generator, a different customer inside the same scenario, or a different quarter of the same synthetic ledger all pass that test. A different question does not, and neither does a different deliverable.
Safe to vary:
- the dataset instance, provided it comes from one generator and carries the same difficulty
- names, identifiers, and the specific customer or ticket at the centre of the case
- the order of the supporting documents
- the calendar window a candidate has to complete it in
Not safe to vary:
- the deliverable, or the number of deliverables
- the stated time
- the rubric, its weights, or the bar for a pass
- whether an assistant is permitted
- how many people read the submission
The generator is what makes the first list workable. Build the exercise so instances can be minted from one specification and varying the data costs nothing and changes nothing about what is measured. Build it around a single hand-made spreadsheet and every variation becomes a redesign, which is how teams end up holding three versions of unknown relative difficulty and no record of which finalist sat which.
Run rotation on a schedule, not on request
Rotation is legitimate when the schedule exists before the candidate does. Decide the window, generate the versions together, calibrate them against the same rubric with the same reviewers, then keep every comparison inside a single version. Announced in advance and applied to everyone inside the window, that is one procedure with instances. Swap it for one person mid-process and you have a second procedure plus an exception with a name attached to it.
Calibration is the step teams skip, and the interview literature records the same habit. A content analysis of 104 interviews reported in 103 articles published between 1997 and 2010 found that studies describing a structured interview used an average of 5.74 of 15 structure components, and two of the scoring-side components were among the least common: detailed notes in 19% of the interviews, statistical prediction in 11% 3. Those counts describe what researchers built into published designs, so read them as a ceiling on ordinary practice.
What calibrating two versions actually involves:
1. Score a handful of past submissions against both versions, with the same two reviewers. 2. Compare the pass rates, and then the disagreements. Versions producing different disagreement patterns are not equivalent, whatever the pass rates say. 3. Fix the version, not the rubric. If version B needs a softer bar to produce the same pass rate, version B is harder and should be rewritten. 4. Record which version each candidate received, in the same place you record the findings.
The alternative is the treadmill, which is where teams land when leaks set the calendar. Rotating faster than you can calibrate produces a stack of versions of unknown difficulty, and the same trap catches question banks, so it is worth reading how the rotating question treadmill plays out before committing to a cadence.
How do accommodations fit without breaking consistency?
Accommodations are not a variation you chose, and they do not break consistency. Under 29 CFR 1630.11 an employer must select and administer tests so the results reflect what the test claims to measure rather than an applicant's impaired sensory, manual or speaking skills, except where those skills are the factor the test purports to measure 4. The working rule that follows: adjust the format or the clock, and keep the task and the bar identical.
The EEOC's 2022 technical assistance on algorithmic assessment names extended time and an alternative version of the test, including one compatible with a screen reader, among accommodations that may be effective 5. That document was removed from the agency's site in January 2025 and is read here from an archive, so it is not current agency position. The regulation underneath it has not changed, and what is owed on any particular request is a question for counsel on your own process.
Three practical rules keep the record clean:
- A request needs no particular words. An applicant who says a medical condition may make the test hard to take has already asked 5. Where the need is not obvious you may ask for reasonable supporting documentation, and a broad pre-offer demand for medical detail is a separate problem of its own.
- Write down what was adjusted and why, in the same file as the submission, so a later reader can see the task and the bar were untouched.
- Never treat an accommodated submission as a different tier. Same exercise, same criteria, finding written the same way.
Extended time is the request worth settling before it arrives, because the answer turns on what your exercise claims to measure. If speed is genuinely part of the requirement, say so in the brief and be ready to defend it. If it is not, and for most take-homes it is not, the clock is a convenience you imposed rather than a construct you measured. Working that out in advance is most of what ADA accommodations mean for an AI-based assessment.
Before the next debrief, write down which version every live finalist received. If two of them are being compared across two versions nobody calibrated, you have two findings sitting in one spreadsheet and no comparison.
Common questions
Can senior and junior candidates get different exercises?
Yes, because they are different roles at different levels, and consistency is owed within a role and stage rather than across a company. Write the split down before candidates arrive: which level gets which exercise, and what the bar is for each. The problem is not two exercises, it is two exercises with no rule saying who gets which, because then the assignment is being made case by case and the reason for it lives in somebody's memory.
Is rotating the exercise the same as changing it?
Not if the rotation was scheduled and calibrated. A version generated in advance, checked against the same rubric by the same reviewers, and given to everyone inside a defined window is one procedure with instances. A brief swapped mid-process for a single candidate is a different thing entirely, however reasonable the trigger was. The test is whether you could describe the rule to a stranger without naming an individual candidate in it.
What if a candidate has already seen the exercise?
Move them to the next scheduled version if one exists, and record that you did. If no other version exists, the honest options are to accept the exposure and note it, or to run the stage live instead. Do not invent a harder brief on the spot: a one-off substitute is the least defensible artifact in the process, since nobody can say afterwards whether it was equivalent. This is the argument for building instances from a generator before you need one.
Do candidates in different countries need the same exercise?
The same task and the same bar, with translation and local material handled as inputs. Where local law adds requirements about notice, data handling or automated decisions, those attach to the process rather than to the content of the exercise, and local counsel is who names them. Keep one rubric. A second rubric per region is how two candidates end up measured on different standards for a role that has only one.
Does an accommodation have to be recorded?
Record what was adjusted and that the task and bar were unchanged. Do not record medical detail, and do not put a diagnosis in a hiring file. The purpose of the note is narrow: a later reader should be able to see that the exercise was the same exercise. That record also protects the candidate, because it is what stops an adjusted submission from being quietly discounted by somebody who noticed the timestamp and drew their own conclusion.
References
- 1. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) govinfo.gov Supports the claim that a work sample is a selection procedure under the Uniform Guidelines, which reach the full range of assessment techniques including informal interviews.
- 2. 42 U.S.C. 2000e-2(k) - Burden of proof in disparate impact cases uscode.house.gov Supports the statutory disparate-impact test quoted here: job related for the position in question and consistent with business necessity, once a particular practice is identified.
- 3. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature doi.org Supports the ceiling reading of how rarely the scoring side of a procedure gets built at all, from the paper's Table 2 content analysis of studies conducted 1997 to 2010.
- 4. 29 CFR 1630.11 - Administration of tests (Regulations to Implement the Equal Employment Provisions of the Americans with Disabilities Act) govinfo.gov Supports the rule that test administration must reflect the factor the test purports to measure rather than an applicant's impaired sensory, manual or speaking skills.
- 5. The Americans with Disabilities Act and the Use of Software, Algorithms, and Artificial Intelligence to Assess Job Applicants and Employees web.archive.org Supports the named examples of accommodation on an algorithmic assessment, extended time and a screen-reader-compatible alternative version, cited as archived guidance.
5 sources, numbered by first appearance. How Olive sources claims
General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.