Policy

Disclosure, Detection, or Observation: Which AI-Use Policy Holds Up?

An employer setting policy on candidates' AI use has three options, and only one holds up: watching the person work. Detection fails because detectors misclassify human writing at rates no employment decision can carry, and errors land hardest on non-native English writers. Honor-system disclosure works only where AI isn't part of the job; roughly half of AI users won't admit to it on the tasks that matter most, so the answer can't feed a decision. Observation costs the most to build and is the only one that returns evidence.

The takeNone of this is new law. The Uniform Guidelines were published in 1978 and they already decide it: a detector threshold is a selection procedure in exactly the way a typing test was, and no vendor has a validation study because there is nothing to validate the output against. What changed is that the cheap check became the indefensible one, and it is still the one a team with no budget for the defensible one reaches for. When an employer finally has to defend a detector rejection, expect the loss to come on validity long before anyone argues about accuracy.

Where Olive fits

Open a role and see what the work shows

A policy that says "watch the work" still has to produce something a reviewer can point at afterwards. Olive returns six separately evidenced findings, each written by a person and anchored to a timestamped excerpt from the session, and every released report exports with its rubric, scorer and bank versions attached.

Rank your shortlist

Which of the three actually holds up?

Observation. Ranked on enforceability, legal exposure and information value, watching a candidate work with AI wins on all three, and the other two only trade places. Detection is unenforceable and carries the worst exposure. Honor-system disclosure is enforceable in the trivial sense that nothing gets checked, carries almost no exposure, and tells you nothing you can act on.

Enforceable?ExposureWhat it tells you
Honor-system disclosureNot at allLow, until the answer decides somethingThat the candidate read the question
AI detectionNo, and the errors are patternedHighest of the threeA property of the prose
ObservationYes, within the session you runOrdinary selection-procedure exposureWhat the person did with the assistant

The third column is where most policy drafts go wrong. Disclosure and detection are both attempts to answer a question about a document: was this written by a model. Observation answers a different question entirely, which is the one the hire turns on: can this person work with a model and still be the one making the decisions.

One rule governs all three, and it is worth reading before the policy is drafted. Under the Uniform Guidelines, a selection procedure is "any measure, combination of measures, or procedure used as a basis for any employment decision," and the definition explicitly reaches "informal or casual interviews and unscored application forms" 4. A disclosure checkbox that routes someone out of the pile is a selection procedure. So is a detector threshold. So is an observed work sample. Once any of them produces adverse impact, it is discriminatory unless it has been validated 4.

So the honest framing is not "which policy is safest." All three sit inside the same rule. The question is which one you can defend once it is in there.

Why detection loses before the policy is written

Because it cannot be made accurate enough to carry a decision, and its errors are not random. Fourteen detection tools tested against human and machine text were found to be neither accurate nor reliable, biased toward classifying output as human-written, and made materially worse by ordinary obfuscation such as paraphrasing 1. Seven detectors run over 91 human-written TOEFL essays produced an average false-positive rate of 61.3 percent 2.

That second number is the one that ends the conversation. In the same study, 97.8 percent of those TOEFL essays were flagged by at least one detector, while essays by US eighth-graders were classified with near-perfect accuracy 2. When the TOEFL essays were rewritten with enriched word choices, the false-positive rate fell to 11.6 percent 2. The tool is reading register, and register tracks first language, education and profession.

National origin is a protected class, so an error pattern that lands on non-native English writers is the adverse-impact case assembling itself. And a procedure with adverse impact is discriminatory unless validated 4. No detector vendor publishes a validation study relating its output to job performance, because there is nothing to relate: the output describes a paragraph, not a person.

The usual rescue is a human in the loop. It does not work, for a mechanical reason. A reviewer who forwards every flag to a rejection has not decided anything, and in New York City the test for an automated employment decision tool is whether it "substantially assists or replaces discretionary decision-making" 6. Where that test is met, a bias audit must be done before use and NYC candidates notified 10 business days ahead 6. A rubber stamp assists substantially.

More on whether detectors work at all in a hiring context, including the arithmetic of a 1 percent false-positive rate against a real funnel.

When is honor-system disclosure enough?

When AI genuinely isn't part of the job, and the policy exists to set an expectation rather than to gather evidence. A disclosure line costs nothing, reads as respect rather than suspicion, and gives a candidate a clean place to say what they used. It stops being enough the moment the answer starts deciding something.

The failure is not that people lie. It is that the incentive structure guarantees the quiet answer. Microsoft and LinkedIn's 2024 Work Trend Index, fielded to 31,000 knowledge workers across 31 markets, found 75 percent using AI at work, 78 percent of those users bringing their own tools in, and 52 percent reluctant to admit using it for their most important tasks 3. That is people describing their current employer, where the job is already theirs. A candidate answering the same question with an offer on the line has strictly more reason to say less.

So an honor-system question collects the least signal from exactly the people whose AI use you most wanted to understand. It is not a screen. Treat it as a stated expectation and nothing more, and keep the answer out of the advance decision, because a disclosure answer used as a basis for a decision is a selection procedure with no validity evidence behind it, on the same footing as the detector 4.

Disclosure does two things well. It tells candidates what the rules are before they spend their evening on your take-home, which is a fairness point rather than a compliance one. And it gives you a written record that the expectation was communicated, which matters later if a dispute arises about what was allowed.

What you can and cannot put in the question is its own problem: see what a disclosure question can legally ask.

What does observation buy that the other two cannot?

The thing the policy is actually about: what the person did. Work samples carry high content validity and high criterion-related validity, applicants perceive them as fair, and subgroup performance differences are generally little or none, though that depends on the competencies assessed 5. Watching someone work with an assistant turns an unanswerable authorship question into an observable one.

What becomes visible is behavior, not authorship. Whether the first move went after understanding the problem or straight for the output. Whether a source was demanded for the claim the recommendation rested on, and whether anyone opened it. What was kept by hand and what was handed over. Whether anything the assistant produced was refused, and on what grounds. Whether one claim got tested against something outside the conversation. None of that survives contact with a finished document, which is why detection and disclosure both come up empty.

The honest costs are real and worth stating in the policy. Work samples may be costly to develop in both time and money, may require periodic updating, and can be time-consuming and expensive to administer, since individuals have to observe and sometimes rate performance 5. That is the trade: the only option that produces evidence is also the only one you have to build.

Two constraints to write in before anyone builds it. Keep the observation scoped to the assignment itself rather than to the candidate's machine. A policy that reaches further than the work is a monitoring policy wearing an assessment's clothes, and it will be read that way. And if a tool rather than a person produces the judgment, you are back inside the automated-decision rules: New York City's definition turns on machine learning, statistical modeling, data analytics or artificial intelligence generating a prediction or classification that substantially assists the decision 6.

If you are deciding what the exercise itself should permit, whether to allow AI on the take-home is the same question one level down.

Write the policy from one question: is AI use part of this job?

Answer that first, per role, before anything is drafted. If the work is done with AI, disclosure produces data you cannot use and detection penalizes the behavior you are hiring for, so observation is the only coherent policy. If the work genuinely isn't done with AI, a one-line disclosure expectation is enough and nothing more is worth its cost. The answer is per requisition, not per company.

  • Write the answer down for each role, with a reason. "Analysts here draft with an assistant and check the numbers by hand" is an answer. "We're an AI-forward company" is not, and it will not survive a question from counsel. Start from what AI actually does in the role.
  • Name the decision each policy element feeds. A disclosure question that feeds nothing is a courtesy, and should be written as one. A disclosure question that feeds an advance decision is a selection procedure and inherits every obligation that comes with it 4.
  • If a detector stays in the process, demote it to routing. A flag sends a file to a human read and never rejects on its own, and the reviewer records the job-related reason they gave. If nobody can name a case where the reviewer overruled the flag, the tool is deciding 6.
  • Scope observation to the work. Say what is captured, for how long, and what is never looked at. Declining should be a supported outcome with a stated consequence, not a silent disqualification.
  • Write the rejection reason before you write the question. A reason that existed before you met the candidate is the one that holds up when they ask for it, and they will. See how to explain an AI-related rejection.
  • Keep impact records per role. Adverse impact attaches to a procedure as used, and a single policy applied across ten requisitions is ten exposures, each with its own applicant pool 4.

The rest of the document (scope, retention, who reviews, what happens on a dispute) is ordinary policy drafting. What belongs in an AI hiring policy covers the sections most first drafts leave out.

Read the evidence

Common questions

Is a disclosure question a selection procedure?

If the answer feeds an employment decision, yes. The Uniform Guidelines define a selection procedure as any measure, combination of measures, or procedure used as a basis for any employment decision, and the definition explicitly reaches informal or casual interviews and unscored application forms 4. A disclosure box that only sets expectations and never routes anyone is not one. The moment somebody is screened out on the answer, it is, and the ordinary validity and impact obligations attach.

Can you keep a detector if a human reviews every flag?

Only if the human genuinely decides. New York City's test for an automated employment decision tool is whether it substantially assists or replaces discretionary decision-making 6, and a reviewer who forwards every flag to a rejection is not exercising discretion. The accuracy problem also survives the review: the false positives fall hardest on non-native English writers, at a 61.3 percent average rate in one study of seven detectors 2, so the human is now adjudicating a biased input. Keep a record of what each reviewer read and the job-related reason they gave.

What if a candidate declines to be observed working?

Write the consequence into the policy before anyone declines, and make it a stated outcome rather than a silent disqualification. Declining is information about the process, not about the person. Offer an equivalent path where you can, tell the candidate what the reviewer will see, and record the decision. A policy that has no answer for this question will produce an inconsistent one under time pressure, which is the shape that later reads as disparate treatment.

Does an observed work sample trigger a bias audit in New York City?

That turns on what produces the judgment. Local Law 144 covers a computer-based tool that uses machine learning, statistical modeling, data analytics or artificial intelligence, helps make an employment decision, and substantially assists or replaces discretionary decision-making 6. A human reviewer reading a recorded session against a written rubric is not that. A model that outputs a prediction or classification about a candidate's fit is. Where the law applies, a bias audit is required before use and NYC candidates get 10 business days' notice 6.

Does one AI-use policy work across every role in the company?

No, and a single policy is usually the tell that nobody asked the underlying question. AI use is central to some jobs and absent from others, so the same clause is a fair expectation in one requisition and an irrelevance in the next. Write the per-role answer down with a reason, then attach the policy elements that answer requires. Impact is also measured per procedure as used 4, so one policy across ten requisitions is ten separate exposures rather than one.

Is observation just surveillance with a better name?

Not if it stays scoped to the work. The distinction is what gets captured and why: a work sample records the assignment a candidate agreed to do, for a stated period, reviewed against a rubric written in advance. Monitoring reaches past the task into the person's machine, their other applications, or their body. Say in the policy exactly what is captured and what is never looked at. If the scope cannot be written in two sentences a candidate would accept, it is too wide.

References

  1. 1. Testing of Detection Tools for AI-Generated Text Weber-Wulff, Anohina-Naumeca, Bjelobaba, Foltynek, Guerrero-Dib, Popoola, Sigut and Waddington (arXiv:2306.15666), International Journal for Educational Integrity, 2023. arxiv.org Fourteen detection tools evaluated; the authors conclude the available tools are neither accurate nor reliable, show a main bias toward classifying output as human-written, and perform significantly worse under content obfuscation.
  2. 2. GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu and Zou, Patterns (Cell Press), 2023. pmc.ncbi.nlm.nih.gov Seven GPT detectors over 91 human-written TOEFL essays: 61.3 percent average false-positive rate, 97.8 percent flagged by at least one detector, near-perfect accuracy on US eighth-grade essays, and 11.6 percent after the essays were rewritten with enriched word choices.
  3. 3. AI at Work Is Here. Now Comes the Hard Part (2024 Work Trend Index Annual Report) Microsoft and LinkedIn, 2024. microsoft.com Survey of 31,000 knowledge workers across 31 markets fielded February 15 to March 28, 2024: 75 percent use AI at work, 78 percent of AI users bring their own tools, and 52 percent are reluctant to admit using it for their most important tasks.
  4. 4. Uniform Guidelines on Employee Selection Procedures (1978), 29 CFR Part 1607 Equal Employment Opportunity Commission, via govinfo, 2023. govinfo.gov Sec. 1607.16(Q) defines a selection procedure as any measure, combination of measures, or procedure used as a basis for any employment decision, expressly including informal or casual interviews and unscored application forms; 1607.3(A) makes a procedure with adverse impact discriminatory unless validated; 1607.4 requires impact records.
  5. 5. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov States that work samples have a high degree of content validity and criterion-related validity, that applicants perceive them as very fair, that subgroup performance differences are generally little or none depending on competencies assessed, and that they may be costly to develop and time-consuming to administer because individuals must observe and sometimes rate performance.
  6. 6. Automated Employment Decision Tools: Frequently Asked Questions NYC Department of Consumer and Worker Protection, 2023. nyc.gov Defines an AEDT under Local Law 144 as a computer-based tool using machine learning, statistical modeling, data analytics or artificial intelligence that helps make employment decisions and substantially assists or replaces discretionary decision-making; sets the bias audit before use and the 10 business day notice to NYC candidates.

6 sources, numbered by first appearance. How Olive sources claims

General guidance, not legal advice. Hiring rules differ by state and country and change often; check anything here against your own counsel before you act on it.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.