Teams

How Do You Stop Hiring Managers Rejecting Candidates for 'Sounding Like AI'?

A memo won't stop hiring managers rejecting candidates for sounding like AI. Change what a rejection must contain instead: the claim in the work that's wrong, thin or unverifiable, the check the reviewer ran, and what it returned. Style can't fill those fields, and anything fixable by rewriting sentences isn't evidence about judgment. Read flag rates by requisition and by reviewer monthly; the instinct fires on conventions that vary by discipline and first language. The fields fix the reason, not the signal. For the signal, run a work sample.

The takeReviewers who flag prose are reaching for an instrument that stopped working. A rejection form makes that loss legible, which is more useful and much less comfortable. Nobody has measured what a mandatory reason field does to a flag rate. The honest guess is that it falls close to zero, not because the instinct went quiet but because it never had anything to write down. That silence is the finding.

Where Olive fits

Open a role and see what the work shows

"It sounded like AI" names a style; a reason that survives review names an act. Olive returns six findings a person wrote, each anchored to the moment in the session it rests on, and every released report exports with its rubric, scorer and bank versions attached.

Rank your shortlist

Why does a memo asking people to stop never work?

Because "sounds like AI" isn't a rule anyone chose to follow. It's a perception, and you can't revoke a perception by email. What you can change is what a rejection is allowed to contain. Make the reviewer write the specific claim in the work that's wrong, thin or unverifiable, and where they checked it. Style alone doesn't fill that field, so the instinct has nowhere to land.

The instinct is also the weakest instrument in the funnel. Across three studies of unstructured interviews, people who unknowingly ran interviews in which the answers were generated at random came away as confident in their impressions as people who ran real ones; having interviewed anyone made their predictions worse rather than better; and participants still reported they would rather have a random interview than no interview at all 2. A confident read of a stranger's prose is that same machinery with less to work from.

A memo produces one of two outcomes, both bad. Reviewers stop saying it out loud and keep doing it, or the rejection note becomes "not a fit" and the reason leaves the record entirely. Neither is a screen anyone can audit six months later, when someone asks why one req family was rejected at four times the rate of another.

Settle the underlying policy in writing first. Whether an AI-written application is disqualifying at all is a decision the company makes once, and a rejection form with no stated policy behind it only moves the argument onto the form.

What does the rejection form have to make them write?

Three fields, and the rejection can't be submitted without them. One: the specific claim, number or decision in the submission that's wrong, thin or unverifiable. Two: what the reviewer did to check it. Three: what the check returned. A reviewer who fills all three has found something real. A reviewer who can only say the prose felt generated has found nothing anyone can act on.

Worked examples travel further than the rule, so put two of them in the form itself:

  • Fills the fields. "The letter says they rebuilt the team's forecasting model. In screen I asked what the old model got wrong and what they changed; the answer named neither, and the two versions they described don't reconcile."
  • Doesn't fill the fields. "Reads like ChatGPT. Every paragraph is the same length and it uses 'moreover' twice."

The second is a statement about sentences. The first is a statement about a claim, and it survives being read back to the candidate, which is the practical test, because an AI-related rejection you can't explain to the candidate is one you will end up explaining to somebody else.

One rule closes the loophole: anything a candidate could fix by rewriting sentences is not evidence about their judgment. Em dashes, tidy parallel structure, an unusually clean summary paragraph: all style, all coachable in an afternoon, and none of it tells you whether the person checked anything. Brief the panel on that distinction before the form ships rather than after the first dispute, since getting a panel to read AI use the same way mostly comes down to everyone holding one sheet of paper.

Why do engineering and commercial reqs flag different people?

Because the flag fires on writing conventions, and conventions vary by discipline and by first language. Even, low-variance prose is what the instinct reads as machine-written, and it is also what you get from a competent engineer writing in a second language, from anyone trained on a house style, and from a candidate who ran the letter past a template. One unexamined rule, two req families, two different rejection patterns.

The measurement of this exists and it is not close. Seven widely used GPT detectors were run over TOEFL essays written by non-native English speakers and over essays by US eighth-graders: the average false-positive rate was 61.22% on the TOEFL essays against 5.19% on the eighth-graders', and 97.80% of the TOEFL essays were flagged as AI-generated by at least one of the seven 1. Prompting a language model to rewrite the same essays with richer vocabulary made most of the flags go away, which is the finding that matters for a hiring panel: the signal is linguistic range, not authorship.

Your reviewers are running a worse version of that classifier from memory, with no threshold and no record. Whether detectors work well enough to use in hiring is a separate question with the same answer: buying one doesn't convert a hunch into evidence, it converts it into a hunch with an invoice attached.

Which is why the by-req cut below is the first thing worth building. A single company-wide flag rate hides the exact pattern you need to see.

Measure the flag rate by req and by reviewer

Two cuts, monthly. Flags per hundred submissions by requisition family, and flags per reviewer. The first shows whether one discipline's writing conventions are being read as a defect; the second shows whether one person is generating most of them. Both are trivial to produce once the reason field exists, and neither is available at all while the reason lives in a reviewer's head.

There is a legal reason to want the record, separate from the accuracy one. The Uniform Guidelines define a selection procedure as any measure used as a basis for an employment decision, and define it expansively: "the full range of assessment techniques from traditional paper and pencil tests, performance tests, training programs, or probationary periods and physical, educational, and work experience requirements through informal or casual interviews and unscored application forms" 3. A resume reviewer's impression sits inside that definition. It is not outside the rules because nobody wrote it down.

The Guidelines also expect an employer to keep records disclosing the impact its selection procedures have on identifiable race, sex and ethnic groups, and treat a selection rate under four-fifths of the highest group's rate as evidence of adverse impact 4. Where a procedure does screen out a protected group, the EEOC's position on employment tests and selection procedures is that the employer has to show it is job-related and consistent with business necessity 5. "It sounded like AI" is not a showing anyone has managed to make. Where a tool does the screening, the duties get more specific still: New York City requires notice to a candidate living in the city at least 10 business days before an automated employment decision tool is used 6. A reviewer's hunch triggers no such notice, which is exactly why the form has to hold the record instead.

Set the threshold before you look at the numbers, so the review isn't a negotiation about whether a gap is large. One reviewer at twice the panel median, or one req family at twice the others, triggers a read of the actual notes. A read, not a retraining. Most months the notes explain themselves, and when they don't, the fix is usually one more worked example on the form.

What a written standard still can't tell you

Whether the person can do the work. The form fixes the reason, not the signal. A candidate whose letter names a claim and a source may have had every word of it produced for them, and a candidate whose prose reads flat may be the sharpest reader in the pile. Writing was never the measure of judgment; it was the part of the work that happened to be visible at screen.

So stop asking the application to carry weight it can't hold, and move the decision to something where the answer isn't stylistic. A short task with an assistant available, a confident claim inside it that happens to be wrong, and a record of what the candidate did about it will separate those two candidates in twenty minutes, and it returns a rejection reason nobody has to argue about. Whether a detector, an interview or a work sample is the right instrument is worth settling before you spend another quarter on the form.

Two things stay in place either way. Say what's allowed with AI in the posting, in one sentence: an unstated rule gets guessed at, and the guessing measures how well someone read your company rather than how well they work. And keep the reason field mandatory after the work sample lands, because the screen still exists and it is still the least examined thing in the funnel.

See how Olive measures this

See the benchmarks

Common questions

Can I just tell managers to ignore how something is written?

Telling them won't hold past the second candidate whose letter reads oddly polished. An instruction to un-notice something produces no output anyone can check. Changing the rejection form produces one: either the three fields are filled or the rejection doesn't go through. Keep the instruction (say plainly that prose style is not a rejection reason here), but let the form be the thing that enforces it, and read the notes monthly.

What if the application really was entirely AI-generated?

It still tells you nothing about the candidate on its own, because you can't establish it and the honest ones look identical to the rest. What works is asking, in a screen, about one specific claim in the document: where the number came from, what the earlier version got wrong, what they would cut. Someone who wrote it, or who directed and checked whatever wrote it, answers immediately and in detail. Someone who pasted a prompt does not.

Should the rejection reason go to the candidate?

The reason has to be one you would be willing to send, whether or not you send it. That is the test the form is built around. A note naming a claim, a check and a result reads as a real reason; "sounded AI-generated" reads as a verdict on someone's writing and produces the reply you don't want. Separately, New York City requires notice to candidates living in the city at least 10 business days before an automated employment decision tool is used 6. A human reviewer's note is a different obligation. Check your states with counsel.

Which requisitions does this hit hardest?

Any req drawing candidates who write in a second language, and any function whose deliverable is prose: marketing, communications, customer success, recruiting itself. Technical reqs aren't exempt: an engineer's cover letter is often the only writing in the file and often the piece they cared least about. The pattern isn't really about the role. It's about how much of the screen rests on writing style, which is why the by-req flag rate is the diagnostic rather than a guess.

Does buying a detector solve this instead?

It makes the record worse, not better. The tool's output becomes the stated reason, so a false positive is now documented as evidence rather than left as an impression, and the measured error rates fall unevenly across writers. A reviewer's note about a specific claim is a better artifact than a percentage from a classifier, and it is the one you can defend when somebody asks.

Is there an assessment that produces a reason a panel can defend?

One shape does: a work sample where the candidate uses an AI assistant on a real occupational task, reviewed by a person who writes down what happened. That is what Olive returns: six findings, each anchored to a timestamped excerpt from the session, with no number standing for a person and no hiring recommendation attached. The candidate is granted the identical report. It's an input to your decision rather than a verdict, which is also what makes the reason quotable in a rejection conversation.

References

  1. 1. GPT detectors are biased against non-native English writers Patterns (Cell Press); Stanford University. Preprint at arXiv, 2023. arxiv.org Seven GPT detectors averaged a 61.22% false-positive rate on TOEFL essays by non-native English writers against 5.19% on US eighth-graders' essays; 97.80% of the TOEFL essays were flagged by at least one detector, and prompting for richer vocabulary removed most flags.
  2. 2. Belief in the unstructured interview: The persistence of an illusion Judgment and Decision Making, Vol. 8, No. 5, pp. 512-520, 2013. journal.sjdm.org Across three studies, interviewers given randomly generated answers were as confident in their impressions as those given real ones; interviews reduced predictive accuracy, and participants preferred a random interview to no interview.
  3. 3. 29 CFR 1607.16 - Definitions Uniform Guidelines on Employee Selection Procedures, via Cornell Legal Information Institute, 1978. law.cornell.edu A selection procedure is any measure used as a basis for an employment decision, expressly including informal or casual interviews and unscored application forms.
  4. 4. 29 CFR 1607.4 - Information on impact Uniform Guidelines on Employee Selection Procedures, via Cornell Legal Information Institute, 1978. law.cornell.edu Employers should keep records disclosing the impact of selection procedures by race, sex and ethnic group; a selection rate below four-fifths of the highest group's rate is generally regarded as evidence of adverse impact.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov A selection procedure that screens out a protected group must be shown job-related and consistent with business necessity.
  6. 6. Notice of Adoption of Final Rule: Automated Employment Decision Tools (§ 5-304, Notice to Candidates and Employees) NYC Department of Consumer and Worker Protection, rules implementing Local Law 144 of 2021, 2023. rules.cityofnewyork.us § 5-304 sets out how an employer notifies a candidate for employment who resides in New York City under Admin. Code § 20-871(b), including notice in a job posting or by mail or email at least 10 business days before use of an AEDT.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.