Assessment design
Detector, Interview, or Work Sample: Which Defends the Decision?
A hiring decision is defended by a work sample or a structured interview, never by an AI detector. Both can be built to produce validity evidence the Uniform Guidelines accept; a detector produces none, misfires hardest on non-native English writers, and answers a question no job description asks. Choose between the two by whether the work leaves a product a reviewer can hold. Where the deliverable is a recommendation, score the observable acts, because a rubric written in terms of judgment names a construct content validity cannot carry.
The takeThe detector's appeal was never accuracy. It was that somebody else does the deciding, on a number, for pennies, and nobody has to write an answer key. Writing that key is the actual expense in both of the methods that survive: the hours a person spends settling in advance what a good answer looks like, then defending it to a reviewer who scored it differently. Skip that step and you probably end up with a work sample graded on taste, which is an unwritten rubric with better manners. One task scored well beats three instruments nobody could score twice the same way.
Where Olive fits
Open a role and see what the work shows
If a work sample is the instrument that survives this comparison, the expensive parts are the answer key and the evidence trail. Olive ships twelve authored cases per occupation, each bank carrying its SOC code, and returns six findings a person wrote by hand, each anchored to the moment in the session it rests on.
Rank your shortlistWhich of the three actually defends a hiring decision?
A work sample or a structured interview, depending on the job. Both can be built to produce the kinds of evidence the Uniform Guidelines recognize: criterion-related, content, or construct 1. A detector produces none of the three. It scores a document for how predictable its wording is, which is not a work behavior, not a job requirement, and not something you can put in front of counsel.
| Method | Validity evidence you can produce | Adverse-impact profile | Candidate experience | Cost |
|---|---|---|---|---|
| AI detector | None tied to the job | Concentrates on non-native English writers 4 | A silent rejection with no appeal | Cheap to run, expensive to defend |
| Structured interview | Content, plus criterion-related with enough hires 1 | Travels with the rubric and the panel | Familiar, and rehearsable | Interviewer hours per candidate |
| Work sample | Content, strongest when the task is the work 1 | Travels with the task and the accommodations | Longest time cost, clearest relevance | Reviewer hours per candidate |
The detector's problem is not that it is inaccurate, though it is. It is that accuracy about authorship is not evidence about the job. What a detector measures is a property of the text, and the measured accuracy of that property is well below what most buyers assume.
The other two are not interchangeable either. A structured interview captures a candidate describing work. A work sample captures work. Both are defensible, and they defend different claims: if the decision rests on "this person can produce X," a description of producing X is a weaker record than X.
What counts as validity evidence for each one?
Three strategies, and only three: criterion-related, content, and construct 1. A structured interview is usually defended on content (the questions sample the job's work behaviors) or on criterion evidence, if you have enough hires to correlate scores against performance. A work sample is defended on content, and the Guidelines say the closer the procedure sits to the actual work, the stronger that case gets 1.
That is the whole menu, and a vendor's accuracy figure is not a fourth item on it. One of the most comprehensive public tests of detection tools put 14 of them against human, AI-written, machine-translated, human-edited and paraphrased documents: the best reached 76% overall, and across the field they averaged 30% on human-edited text and 15% on paraphrased text 3. Even at 100% those numbers would say nothing about performance on a job.
Adverse impact is where the detector stops being merely unhelpful. Seven detectors run over 91 TOEFL essays by non-native English speakers returned an average false positive rate of 61.22%, while classifying US eighth-grade essays almost perfectly 4. A neutral selection procedure that disproportionately excludes on a protected basis has to be job-related and consistent with business necessity 2, and a vendor threshold is not that showing. Run the four-fifths arithmetic on any instrument before it gates anyone.
Interviews and work samples carry adverse-impact risk too. It travels with the content rather than with the writing style: a coding task heavy on unfamiliar tooling, an interview rubric that rewards one idiom of confidence. The difference is that you wrote those, you can inspect them, and you can change them.
Why is a work sample easy to defend in software and hard in consulting?
Because content validity rests on resemblance, and the resemblance is obvious in one case and arguable in the other. The Guidelines ask you to show that the behavior in the selection procedure is a representative sample of the job's behavior, or that it yields a representative sample of the work product 1. A patch to a repository is a work product. A market-entry recommendation is an opinion about the future.
Judgment is on the regulation's own list of constructs, alongside intelligence, aptitude, personality, commonsense and leadership, and content validity is not an appropriate strategy for procedures that purport to measure traits or constructs 1. That clause catches most consulting cases by itself. A consulting work sample whose rubric reads "demonstrates strong judgment" has named a construct, and the content argument stops carrying it.
The fix is not to drop the work sample. It is to score observable acts instead of the construct: which sources were opened, which assumption was written down before the recommendation, which number was recomputed when the first answer looked convenient. Those are behaviors a reviewer can point at in the record, and an observable act is the level at which the Guidelines ask a skill to be defined 1.
Two further clauses are worth designing against. The manner, setting, level and complexity of the procedure should closely approximate the work situation, and content validity is not available for something an employee is expected to learn on the job 1. So a task sized to a real deliverable defends better than a puzzle, and a task testing tooling your new hire will pick up in week two defends worse than one testing what they must already have. If the format itself is still open, work sample against structured interview is the narrower comparison.
What does each one cost you and the candidate?
The detector costs nothing to run and the most to defend. A structured interview costs interviewer hours, the scarcest resource on most teams, and returns a description of work rather than work. A work sample costs candidate hours, which is the cost that shows up later as attrition, plus reviewer hours to grade. Price the grading, not the task.
Candidate experience separates the two survivors more sharply than cost does. A work sample takes the most candidate time and is the easiest to justify while taking it, because the relevance is visible from the first screen. A structured interview takes the least candidate time and rewards rehearsal: the practiced answer and the earned one look identical in a transcript.
The detector sits in a different category again. A candidate cut by one is never told, cannot appeal, and did nothing wrong. Given where the measured errors concentrate 4, the people that happens to are not randomly distributed, which is a fairness problem before it is a legal one.
Two costs get underpriced almost everywhere. Grading is the first: a work sample with no answer key written in advance turns into an opinion poll among reviewers, and the disagreement stays invisible until someone audits it. Accommodation is the second: any timed instrument needs an accommodation path planned before the first invite goes out, not improvised on the third.
Pick the instrument the job's own evidence supports
Write down the job's critical work behaviors first, then pick the instrument that can sample them 1. If the work leaves a product a reviewer can hold (code, a model, a memo, a denial appeal), run a work sample. If it does not, or if the observable behavior is genuinely conversational, run a structured interview and defend it on the same job analysis.
Then keep four records, because together they are what a defense consists of: the job analysis naming those behaviors, the rubric with its scoring anchors, the log of who scored what, and the reason each behavior is on the list. The Guidelines expect the analysis to focus on work behaviors and the products they result in 1, and the EEOC expects the procedure to be job-related and consistent with business necessity if it excludes disproportionately 2.
Use the detector for nothing. If the worry underneath the question is that take-homes now arrive finished and polished from everyone, that is a real problem with a different answer: change what the task asks for, so that a polished artifact stops being the evidence and the record of producing it becomes the evidence instead. At the level of the role, the same move is hiring for what a person verifies rather than what they can produce.
Olive publishes what it has validated and what it has not. See how Olive measures this
Common questions
Can you reject a candidate on an AI detector result?
Not defensibly. A detector becomes a selection procedure the moment it changes an outcome, and it produces no validity evidence about the job, only a claim about the text. The measured error also concentrates: seven tools tested in Patterns returned an average 61.22% false positive rate on essays by non-native English speakers while classifying native-speaker essays almost perfectly. That is national-origin exposure you would have to justify as job-related and consistent with business necessity, on evidence a vendor threshold cannot give you.
Which defends better in an EEOC inquiry, a structured interview or a work sample?
Whichever one you can show sampled the job. Both are defensible, and the defense is the job analysis, the rubric and the scoring record rather than the format. A work sample usually has the shorter argument, because the Uniform Guidelines say the closer the procedure's content and setting are to the actual work, the stronger the content-validity case. A structured interview matches it when the questions map to critical work behaviors and every candidate is scored against the same written anchors.
What makes a work sample content-valid?
Resemblance, defined narrowly. The behavior in the task has to be a representative sample of the job's behavior, or the task's output a representative sample of the work product, with the manner, setting, level and complexity close to the real work situation. Two things break it: scoring a construct such as judgment or aptitude rather than an observable act, and testing something the person would be taught in their first month. Score what a reviewer can point at in the submission.
Do you need both a work sample and a structured interview?
Usually one, plus a short structured conversation about it. The work sample produces the evidence; a 20-minute structured follow-up on the submission tests whether the candidate can account for their own choices, which is the part a take-home cannot show. Running both as full independent rounds doubles candidate time for a small gain in signal, and candidate time is where completion rates go. Use the same rubric anchors in both, so the two records stay comparable.
How do you defend a work sample for a role with no tangible work product?
Move the evidence from the product to the acts. When the deliverable is a recommendation, the defensible record is what was done to reach it: which sources were opened, which assumption was stated before the answer, which figure was recomputed. Those are observable behaviors you can tie back to the job analysis, and they avoid the construct problem that sinks any rubric written in terms of judgment or business acumen. Score them separately rather than blending them into one verdict.
References
- 1. 29 CFR 1607.14 — Technical standards for validity studies (Uniform Guidelines on Employee Selection Procedures) ✓ ecfr.gov The three acceptable validity strategies; content validity requires a representative sample of job behavior or work product and is not appropriate for constructs such as judgment or for skills learned on the job.
- 2. Employment Tests and Selection Procedures ✓ eeoc.gov A neutral selection procedure with a disproportionate exclusionary effect must be job-related and consistent with business necessity.
- 3. Testing of detection tools for AI-generated text ✓ edintegrity.biomedcentral.com 14 detection tools tested; best 76% accuracy, none above 80%; field averaged 30% on human-edited and 15% on paraphrased text.
- 4. GPT detectors are biased against non-native English writers ✓ doi.org Seven detectors over 91 TOEFL essays returned a 61.22% average false positive rate while classifying US eighth-grade essays near-perfectly.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.