Assessment design

Work Sample, Interview, or Assessment: Which Catches AI Dependence?

A work sample, a structured interview, an AI-skills test: none catches AI dependence if all you see is what comes back. Ranked by how easily somebody who cannot do the work passes anyway, the unwatched take-home is easiest, the AI-skills test next, the structured interview hardest, and even the interview rewards candidates who prepared with AI. What holds up is a work sample with the session recorded, and that is one occasion, not a forecast. Dependence is an absence of acts, and an absence is invisible in a deliverable.

The takeThe three-way comparison is the wrong argument, and it survives because it is cheaper than the alternative. Choosing a format is a purchase. Rebuilding a task so the working record is the thing scored is a project somebody has to own. So teams relitigate take-home against interview, quote a validity number at each other, and ship the round they already had. Nobody has measured which rounds lean on that validity number hardest, and the ones that have changed least are where I would look first. Format was never the variable. Whether anyone can see the work happen is.

Where Olive fits

Open a role and see what the work shows

The recorded variant of this comparison is what Olive runs: a 40-to-60-minute assignment written for one occupation, taken with an AI assistant, with the working record kept. It comes back as six findings a person wrote by hand (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session, and the candidate is granted the same document.

Rank your shortlist

Which of the three actually catches AI dependence?

None of them, if the only thing you look at is what comes back. Rank the three by how easily a candidate who cannot do the work passes anyway, and the unwatched take-home is easiest, a multiple-choice AI-skills test is next, and a structured interview is hardest without being hard. Dependence is a property of how somebody works, so it shows up only in a record of the work happening.

MethodHow a candidate who can't do the work passesWhat the artifact still hides
Unwatched take-home or work samplePaste the brief, return the outputEverything that happened between the brief and the answer
AI-skills test, multiple choice or scenarioAnswer it with the assistant it asks aboutWhether any of it transfers to the occupation's work
Structured interviewRehearse the answers with a model 3Whether the check they described was ever run
Work sample with the session recordedThe record shows what was done and what wasn'tWhether the same person works this way in month three

Treat this as a capability question rather than an integrity one. A candidate who hands the assistant the part they cannot do is not cheating you; they are showing you the shape of what they cannot do yet, and that is the thing worth knowing before an offer. The register matters because it changes the instrument: you are trying to see a working pattern, not to catch someone out.

The fourth row is the honest answer to the headline, and it is a variant of the first rather than a fourth method. A work sample whose acts are visible (the questions asked before generating, the source opened, the direction refused) answers the criterion. The same task with only a file at the end does not. That distinction, not the choice of format, is where the signal lives, and it is the whole argument behind rebuilding what a work sample tests.

Why does high validity not mean resistance to an assistant?

Because validity ranks how well a method predicts performance, not how well it holds up when the candidate has help. In the 2023 re-analysis of the meta-analytic record, structured interviews top the list at a mean validity of .42, with an 80% credibility interval running from .18 to .66 1. The spread inside one method is wider than the distance between methods.

That interval is the practical finding, and it is usually read past. "Structured interview" is not one instrument. A shared behavioral bank that every mock-interview tool has trained against sits near the bottom of that range; a bank built from decisions your own team made last quarter sits near the top. If your round is running on the universal questions, the number you are relying on describes a different round than the one you have, which is the whole of what AI coaching does to a question bank.

The second thing the estimates do not carry: none of that research was measuring what happens when a candidate has an assistant open. It was measuring prediction of job performance from a selection procedure. So the ranking is a real answer to a real question, and it is not this question. A method can predict performance well and still be trivially passable by someone with a model running, because those are separate properties.

Use the validity literature the way it was meant: to decide whether a method can support a decision at all, and to justify the round to legal and to the hiring manager. Then run the assistant question separately, method by method, because nothing in the meta-analysis answers it for you.

What does an assistant change about each of the three?

It changes what the output proves, and it raises the candidate's confidence at the same time. In a controlled study, participants with an AI coding assistant wrote significantly less secure code than those without one, and were more likely to believe their code was secure 2. That pairing is the thing worth seeing: not the assistant's involvement, but a person who cannot evaluate what it handed them.

  • The take-home or unwatched work sample loses the part of the brief that was fully specified. Whatever the instructions pinned down, the assistant can finish, so every submission converges on the same competent artifact and the grading distribution flattens. What remains gradeable is the part the brief deliberately left open: an unstated assumption, a defect in the data, two constraints that cannot both hold.
  • The AI-skills test names the wrong construct. A multiple-choice instrument measures declarative knowledge about tools, which the tools answer, and which decays as the tools change. It also tells you nothing about the occupation: knowing what a context window is has no bearing on whether an underwriter opens the survey before binding. Set the bar on the work instead, which is the argument in what an AI assessment should measure.
  • The structured interview degrades in a specific direction: toward a better-sounding candidate. In the research an MIT Sloan Management Review piece reports, candidates who prepared with generative AI received higher overall interview ratings than unassisted candidates 3. The answers get more polished and more contextualized, and a rating scale that rewards a well-organized STAR answer rewards exactly that.

The common failure is the same in all three: each one collects a description or a product, and dependence is neither. It is the absence of an act: nothing asked before generating, no source opened, nothing refused, nothing checked against anything outside the conversation. An absence is invisible in a deliverable and visible in a record.

Match the method to what the role's work leaves behind

If the work leaves an inspectable output (code, a model, a query, a denial appeal), run the work sample, because content validity is a short argument when the task is the task 4. If the work is judgment under constraint, with a recommendation as the artifact, the interview carries more load, and the Uniform Guidelines say why: a procedure aimed at a mental process cannot be supported primarily on content validity, and judgment is named among the constructs excluded 4.

That is the clause most three-way comparisons skip, and it is the reason the right combination changes with the requisition rather than with the company. The Guidelines also ask that the manner, setting, level and complexity of the procedure closely approximate the work situation 4. A software task meets that bar easily. A market-entry recommendation is an opinion about the future, and a rubric that scores it as "strong judgment" has named the construct the regulation excludes.

The fix for the second case is not to drop the work sample. It is to move the scoring off the recommendation and onto the acts that produced it: which sources were opened, which assumption was written down before the answer, which figure was recomputed when the first one looked convenient. Those are observable, they are gradeable, and they are what screening for AI judgment in non-technical roles comes down to.

One consequence for the process: a single instrument bought once and applied to every opening will fit some requisitions and misfit others, which is the trade in running one assessment or one per role. Pick per family of work (inspectable output, or judgment under constraint) rather than per department.

Score the criterion, then build the round around it

Write the pass condition down first: a candidate who cannot do this work should not be able to produce a passing submission with an assistant open. Then build the round that makes that true: ask for the record alongside the deliverable, add a short structured follow-up on the candidate's own submission, and score acts a reviewer can point at rather than a verdict on their judgment.

1. Author one real ambiguity into the task. Something the assistant will resolve confidently and wrongly: a defect in the dataset, a claim that only settles when a source is opened, two requirements that cannot both hold. 2. Ask for the working record, not only the output. The intermediate artifact, the exchanges, the thing that was thrown away. State plainly that the assistant is allowed, because a policy nobody believes produces a hidden assistant and no record at all. 3. Follow the submission with 20 structured minutes. Ask what was rejected and why, and what would have changed the answer. Those are answerable only by whoever did the work, and the shape of that round is take-home against live working session. 4. Write the answer key before the first invite. Two reviewers scoring the same submission differently is the failure that stays invisible for months, which is why a rubric two reviewers score the same is the expensive part of building this yourself. 5. Run it beside the current round before it gates anyone. Compare the two on candidates you already decided on, and keep the old round until the new one earns it. The method is piloting an assessment before making it a gate.

None of this delivers certainty, and a round sold internally as certainty will lose its credibility on the first bad hire. What it delivers is evidence about one candidate on one occasion: what they framed, what they checked, what they refused. Olive publishes the six behaviors its reviewers write findings against, if a starting list is useful. See how Olive measures this

See what gets scored

Common questions

Can a work sample tell you whether a candidate depends on AI?

Only if the session leaves a record. The finished artifact cannot, because whatever the brief specified is now free: an assistant produces a competent version of it either way. What separates candidates is the set of acts around the artifact: the questions asked before anything was generated, the source actually opened, the direction refused, the figure recomputed. Ask for those alongside the deliverable, or run the task somewhere they are visible. A work sample graded on output quality alone measures the assistant.

Do structured interviews still work when candidates prepare with AI?

They still work, and they still degrade. Structured interviews top the revised validity estimates at .42 mean validity, but the 80% credibility interval runs from .18 to .66, so implementation decides where you land. Candidates who prepare with generative AI have been found to receive higher overall interview ratings than unassisted candidates, which means polish rises faster than signal. Rebuild the question set from decisions your own team made recently, ask what happened rather than what would, and follow every claim with a specific question about the same event.

Is a multiple-choice AI-skills test worth running?

Rarely, and never as a gate. It measures declarative knowledge about tools, which the tools themselves answer and which goes stale as the tools change. It also has no occupational grounding: naming a technique correctly does not predict whether someone opens the source before repeating what a model told them. If you need a fast top-of-funnel filter, filter on something the job actually requires. If you need evidence about AI judgment, that takes a task, not a quiz.

Should you ban AI during the assessment instead?

A ban converts an unmeasurable behavior into an unenforceable rule. You cannot verify it on an unwatched take-home, and the candidates who comply are the ones you disadvantage. The stronger move is to allow the assistant explicitly and change what gets graded, so that using it well and using it badly stop looking alike. Where the job genuinely forbids AI on some material (protected data, a licensing requirement), say so and scope the task to the part where the rule is real.

Do you need all three methods in one loop?

No. Two is usually the ceiling: one task that produces work, plus a short structured conversation about that same work. Running an AI-skills test, a take-home and a full behavioral loop triples candidate time for signal that overlaps, and completion drops as the total climbs. Pick the task format from what the role's work leaves behind (an inspectable output or a recommendation), and let the follow-up round do the rest, using the same rubric anchors so the two records stay comparable.

How do you defend the method choice to legal?

With the job analysis, not with the format. Write down the critical work behaviors first, then show that the procedure samples them: for a work sample, that the manner, setting, level and complexity approximate the real work situation; for an interview, that the questions map to those same behaviors and every candidate is scored against the same written anchors. Score observable acts rather than constructs such as judgment or aptitude, because a content-validity argument does not carry a procedure aimed at a mental process.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology (Cambridge University Press), Sackett, Zhang, Berry & Lievens, 2023. doi.org Structured interviews top the revised list of widely used predictors at a mean validity of .42, with an 80% credibility interval running from .18 to .66.
  2. 2. Do Users Write More Insecure Code with AI Assistants? Perry, Srivastava, Kumar and Boneh (Stanford University), ACM CCS 2023, 2023. arxiv.org Participants with access to an AI coding assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure.
  3. 3. When Candidates Use Generative AI for the Interview MIT Sloan Management Review, 2025. sloanreview.mit.edu Reports research in which candidates who prepared with generative AI received higher overall interview performance ratings than unassisted candidates.
  4. 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (eCFR), 2026. ecfr.gov Content validity holds to the extent a procedure is a representative sample of the job's content, with manner, setting, level and complexity closely approximating the work situation; it cannot primarily support a procedure aimed at a mental process, and judgment is among the constructs named.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.