Interviewing

Recognition Error Is Not Evenly Spread. Do Not Score the Transcript

The accuracy figure an interview transcription vendor quotes is an average over whoever they tested. Measured error is stratified by speaker: in laboratory reading tests across four commercial speech services, mean word error rate was 52.6% for d/Deaf and hard-of-hearing speakers against 5.0% for hearing speakers, roughly ten times higher, and it moved enormously with intelligibility. So treat the transcript as an instrument with a known failure mode, keep the audio, and judge what the candidate said rather than how cleanly it came back.

The takeThe dangerous artifact is the summary built on top of the transcript, because a summary launders a recognition failure into a judgment about a person and drops the evidence on the way. Nobody reading "communicated unclearly" in a hiring file can tell whether the candidate was unclear or the microphone was. A quoted sentence can be checked against a recording. An adjective cannot be checked against anything, which is exactly why it survives to the debrief.

Where Olive fits

Open a role and see what the work shows

Olive's think-aloud is spoken or typed, chosen at the start of the session and equally valid either way, and a human reviewer writes each of six findings against the timestamped moment behind it, so a disagreement can be about that moment rather than about a rating. The candidate reads the identical report.

Rank your shortlist

What does a vendor accuracy number actually measure?

It measures the population the vendor tested, on the audio conditions they tested it in. Word error rate is a property of a speaker and a recording as much as of a model, so one figure is an average hiding its own distribution. Ask two questions before accepting it: which speakers were in the sample, and was the speech spontaneous or read from a script.

The second question does more work than people expect. Tested on read sentences from a corpus of non-native speakers of English, recent systems land close to human accuracy: Whisper and AssemblyAI reached mean match error rates of 0.054 and 0.056 on 2,400 single-sentence recordings from 24 speakers 3. Read speech in a research corpus is near the easiest condition available, and a job interview is near the hardest, being unscripted and recorded on whatever microphone and room the candidate has. The six first languages in that corpus are also well resourced in training data, and match error rate is not word error rate, so those two numbers do not sit on the same scale as the ones below. That paper is also a preprint.

That rules out the blanket claim that speech recognition fails for everyone with an accent. It does not. The finding the vendor slide never breaks out is where the failures actually concentrate.

Whose speech does it get wrong?

Deaf and hard-of-hearing speakers, speakers with disordered speech, and speakers of under-represented dialects carry the measured error. Across four commercial speech-to-text services, mean word error rate was 52.6% for d/Deaf and hard-of-hearing speakers against 5.0% for hearing speakers, about ten times higher 1. That is laboratory read speech from 24 deaf or hard-of-hearing speakers and nine hearing ones, so it sizes a known gap under easy conditions.

The average also hides the result that matters most. The same study reports 85.9% word error rate for speech rated low in intelligibility, 46.6% for medium and 9.5% for high, the last comparable to the hearing group 1. So the tool is close to usable for one candidate and useless for the next, and a single number for either is wrong. Variation is the story.

Disordered speech shows the same shape with a trap attached. In a study of 432 people with self-reported disordered speech who each recorded at least 300 short phrases, the two speaker-independent models had median word error rates of 31.5% and 29.4%, while models personalized to each individual speaker reached a median of 4.6% 2. Only the first pair applies to hiring. Personalization needs hundreds of that person's own recordings, which no tool meeting a candidate for the first time has, so 4.6% should never be quoted as evidence that the problem is handled. One of those two speaker-independent models is a commercial API and the other was built by the authors to hold architecture constant, and both are 2021-era, so current systems do better than these figures.

Dialect carries a measured gap of its own. Five commercial systems transcribing matched sociolinguistic interviews returned an average word error rate of 0.35 for Black speakers against 0.19 for white speakers, and the authors traced it to the acoustic models and to how little training audio came from Black American speakers 6. A word error rate of 0.35 means roughly one word in three came out wrong, and these were the commercial systems of 2019, tested on research recordings of conversation and not on job interviews, so read it as a property of the tools of that period.

What this looks like from the candidate's side is on the record in one complaint. An Indigenous Deaf woman who had worked seasonally for Intuit since 2019 alleges she was required to take a HireVue automated video interview for a promotion, that parts of the audible content lacked subtitles, that her request for human-generated captioning went unmet, and that the feedback after her rejection focused on her communication style and recommended she practice active listening 4. Those are allegations in an administrative complaint, and no agency and no court has ruled on them. Read it for the shape of the failure: a recognition gap and an evaluation of communication, compounding inside one process.

Stop scoring adjectives about delivery

Adjectives are the layer most contaminated by recognition error and the hardest to argue with. A summariser reading a degraded transcript cannot separate a candidate who rambled from a tool that could not hear them, and nobody downstream can either. Quoted sentences survive that. Ratings of delivery do not. Keep nervous, unclear, rehearsed and low-energy out of the hiring file entirely.

Delivery analysis also carries legal exposure that predates anybody's AI policy. In Baker v. CVS Health Corporation, decided in February 2024, a federal court in Massachusetts declined to dismiss a putative class action alleging that a HireVue video interview whose recordings went to a third-party platform for analysis of facial expressions, eye contact, voice intonation and inflection amounted to a lie detector test administered without the notice a state statute requires 5. Denying a motion to dismiss decides only that the allegations, taken as true, state a claim. It decides nothing about the technology. The statute at issue is a Massachusetts lie-detector notice law, so its bearing on a discrimination claim is indirect, and what has happened in the case since is not reported here.

The scope question does not depend on that case. The federal Uniform Guidelines on Employee Selection Procedures, issued in 1978, define a selection procedure as any measure used as a basis for any employment decision, and say so down to informal or casual interviews 7. They carry a validation burden only where adverse impact appears, so a scoring layer nobody classified as a selection procedure can still be one. Whether any of that reaches your own process is a question for counsel.

Scoring demeanour is also the channel a great deal of other bias travels through, which is the wider argument in naming what each interview round measures. Whether a summary carrying that kind of adjective can go into the hiring file at all is worked through in what an AI notetaker's summary can be used for.

Read one transcript against its recording

Pick a recent interview with a candidate whose speech the tool struggled with, and read the transcript against the audio line by line. The gap you find is the size of the error already sitting inside your decisions. Do it once and the vendor conversation changes, because the question becomes who the product is accurate for.

Four changes follow, and none of them needs a new vendor:

  • Ask for accuracy broken out by speaker group, and for how it was measured. A single global figure, or a refusal, is itself the answer.
  • Keep the audio for as long as you keep the transcript, so a disputed quote can be checked against the thing it came from.
  • Let candidates correct their own transcript before anyone deciding reads it. It costs one email and it repairs the error at the only point where somebody knows what was actually said.
  • Require any quote a decision rests on to be confirmed by a person who was in the room. If nobody was in the room, that is itself the finding.

The fourth one survives contact with a busy panel, because it applies only to the handful of sentences doing real work. The other three are configuration.

Transcription itself is not the problem. A record everyone can point at beats four people's memories of the same hour, and it is what makes an interview reviewable at all. Nor will a transcript tell you whether an answer was drafted somewhere else, which is a separate question with a separate answer in whether AI detectors work on interview transcripts. The argument here is against treating the record as the candidate. Where a decision rests on words, go back to the words. Where it rests on how the words sounded, there was never anything to go back to.

See how it works

Common questions

Is one accuracy percentage from a vendor enough?

No, because the number is an average over a sample you did not choose. Word error rate varies by speaker, by microphone, by room and by whether the speech was scripted, and published research on laboratory reading tests finds gaps of roughly ten times between hearing and deaf speakers on the same systems. Ask which speakers were in the evaluation, whether the audio was read or spontaneous, and whether accuracy is reported by group. A vendor that will only give one figure has told you how carefully the question was asked internally.

Should candidates be allowed to correct their own transcript?

The candidate is the only person who knows what they meant to say, which makes a short correction window the cheapest repair available. It catches recognition errors at the source, before they reach a summary nobody can audit. Send the transcript, invite corrections with a deadline, and record both versions. It also changes the tone of the process: a candidate who can fix a misheard sentence is being treated as a participant rather than as material.

Does an interview transcript count as a selection procedure?

Treat it as part of one. The federal Uniform Guidelines on Employee Selection Procedures, issued in 1978, define a selection procedure as any measure used as a basis for any employment decision, and say so down to informal or casual interviews, so a transcript feeding a scored summary sits inside that scope. The practical consequence is that anything derived from the transcript needs to be job related and defensible, which is a high bar for a rating of how somebody sounded and a low one for a quoted sentence about the work. Check the specifics with counsel.

What about accented speech from non-native English speakers?

The evidence does not support a blanket claim in either direction. On read sentences from a well-known non-native speech corpus, recent systems approached human accuracy, so accent alone is not automatically a wall. That corpus covers six well-resourced first languages and easy recording conditions, and it says nothing about under-represented dialects, where measured gaps are large. Treat accent as a case for keeping the audio and confirming quotes, not as a reason to assume the transcript is broken.

Can a summary describing how a candidate came across go in the hiring file?

Keep it out. A description of delivery is the part of the output most affected by recognition error and the part nobody can check later, which is a bad combination in a document that may be read by a lawyer. Record what the candidate said and what the interviewer concluded about the work. If a panel member believes delivery is job relevant, make them write the specific behaviour and the moment it happened, so somebody else can go and look at it.

References

  1. 1. Quantification of Automatic Speech Recognition System Performance on d/Deaf and Hard of Hearing Speech The Laryngoscope (American Laryngological, Rhinological and Otological Society), author copy hosted at Cornell University, 2025. koenecke.infosci.cornell.edu Supports the ten-times gap in mean word error rate between deaf and hearing speakers, and the spread by speech intelligibility.
  2. 2. Automatic Speech Recognition of Disordered Speech: Personalized Models Outperforming Human Listeners on Short Phrases Proceedings of Interspeech 2021, ISCA Archive, 2021. isca-archive.org Supports the speaker-independent median error rates for disordered speech, and the point that personalized models are unavailable to a hiring tool.
  3. 3. Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling arXiv (arXiv:2503.06924), 2025. arxiv.org Supports the claim that non-native accent is not automatically a wall, on read speech, with the corpus limits stated.
  4. 4. Complaint of Discrimination (D.K., against Intuit, Inc. and HireVue, Inc.), redacted public copy American Civil Liberties Union (assets.aclu.org), read via the Internet Archive Wayback Machine, 2025. web.archive.org Cited strictly as allegations, for the account of a captioning request going unmet and feedback focused on communication style.
  5. 5. Baker v. CVS Health Corporation, Civil Action No. 23-11483 (D. Mass. Feb. 16, 2024) United States District Court for the District of Massachusetts, read via FindLaw Caselaw, 2024. caselaw.findlaw.com Cited as a pleading-stage ruling only, for the claim that affect analysis of face and voice in an interview creates exposure under statutes not written for it.
  6. 6. Racial disparities in automated speech recognition Proceedings of the National Academy of Sciences (PNAS), read via PubMed Central, 2020. pmc.ncbi.nlm.nih.gov Supports the measured dialect gap in word error rate between Black and white speakers on the commercial systems of 2019, with the research-recording condition stated.
  7. 7. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the scope claim that a selection procedure reaches informal or casual interviews, and that validation is required only where adverse impact appears.

7 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.