Interviewing
How Do You Judge Communication After an AI-Translated Interview?
A candidate who ran the interview through live AI translation can still be judged on what they decided. A translation layer moves words between languages. It does not pick which trade-off to defend, demand the number nobody supplied, or repair an answer a follow-up broke. Read the record for those and most of your round survives. A garbled sentence belongs to the pipe: mark it unreadable and ask again. If the job genuinely needs unmediated English, name it in the posting and test it openly, the same segment for everyone.
The takeThe tool didn't create this problem. It exposed one. Interviews have graded English fluency as a stand-in for judgment for about as long as there have been interviews, and nobody had to look at that, because the stand-in was invisible. A live translation layer makes it visible, and most of the discomfort in this question is the discomfort of being asked what you were scoring all along. Watch which rounds object hardest to the tool. It tends to be the ones running without a rubric, because an impression was the whole instrument, and the layer took it away.
Where Olive fits
Open a role and see what the work shows
An interview in any language captures a candidate describing how they would check a confident claim rather than checking one. Olive puts that in front of them as work: a role-grounded assignment done with an AI assistant, written up by a human reviewer as six evidenced findings, with the candidate granted the same report.
Rank your shortlistWhat does a translation layer actually change?
Less than it felt like at the time. The layer changes the surface: word choice, idiom, rhythm, and the small talk that used to make the first five minutes easy. It does not choose which trade-off to defend, name the constraint nobody mentioned, or produce a second answer that holds when a follow-up moves the target. Most of what the round existed to measure is still there.
Six things come through intact, and they are the ones worth writing down:
- The decision, and the reason given for it before you asked for one.
- What they demanded first: which figure, which document, which person.
- What they refused to guess at, and what they said they would have to check.
- The correction, when a follow-up made the first answer wrong.
- The question asked back, and whether it was about the constraint or about the format.
- The order of operations, and what each step would settle.
What does not come through is fluency, and fluency is what you are most likely to grade by accident. In a University of Chicago experiment, native English listeners rated identical trivia statements as less true when the speaker had an accent: 6.95 for a mild accent and 6.84 for a heavy one, against 7.59 for native speakers, on a 14 cm scale where a higher number means more truthful. Every speaker was reciting statements the experimenter wrote, so prejudice about the source of the claim could not explain it, and listeners told that processing difficulty was the cause corrected for a mild accent but not a heavy one 1.
That is the reason to run the ladder rather than trust the impression. The same follow-up questions that expose whether someone understands the answer they just gave work through an interpreter, because each one asks for a different fact rather than a better sentence.
Does the job actually require unmediated English?
Answer that in writing before you judge anything. In the United States, the EEOC's enforcement guidance on national origin discrimination, issued 18 November 2016, holds that an English fluency or proficiency requirement "is permissible only if required for the effective performance of the position for which it is imposed," and that the degree of fluency lawfully required varies from one position to the next, assessed case by case 2.
The same guidance sets a two-part standard before an accent may factor into a decision: evidence that effective spoken communication in English is required to perform the duties, and evidence that this person's accent materially interferes with their ability to communicate in spoken English 2. A discernible accent is not evidence of either. The regulation behind the guidance is blunter still. National origin discrimination includes denying equal employment opportunity because an individual has "the physical, cultural or linguistic characteristics of a national origin group" (29 CFR 1606.1), and when the Commission investigates a selection procedure for adverse impact on the basis of national origin it applies the Uniform Guidelines on Employee Selection Procedures 3. Those guidelines define a selection procedure as any measure or procedure used as a basis for an employment decision, running from paper and pencil tests "through informal or casual interviews and unscored application forms" 4. Your round is one of those. This is public legal fact rather than advice, and the specific call belongs with counsel.
So write the requirement down and be honest about which side of the line the role sits on.
- Genuinely required. An on-call incident commander who runs a live English bridge at 3am. A revenue cycle specialist who phones a US payer and argues a denial in real time. A consultant who takes questions from a client's board with no deck to hide behind.
- Usually not. A role whose output is written and reviewed, whose standups are already captioned, and whose team runs in three languages on a normal Tuesday. Here the interview's language was a courtesy, not a duty.
If it is required, name it in the posting, test it as its own short segment on a real task, unaided, with the same prompt for everyone, and say in the invite that the segment exists. Testing it openly is also the only version you can defend later. The move is the same one that separates a condition from the construct in ADA accommodations on an AI-based assessment: what changes is how the person reaches the task, and the task still measures what the job needs.
Where does the translation pipe put errors, and whose are they?
In two places, and neither of them is the candidate. Speech recognition transcribes first, a model translates second, and both stages fail unevenly across speakers and language pairs. The damage arrives as a dropped clause, a technical term swapped for a general one, or a hedge flattened into a flat claim. Grade any of that as imprecision and you are grading the pipe.
How uneven the first stage is has been measured. Across five commercial systems from Amazon, Apple, Google, IBM and Microsoft, transcribing 19.8 hours of structured interviews with 42 white and 73 black speakers, the average word error rate was 0.35 for black speakers against 0.19 for white speakers, and the gap held on a subset of identical phrases spoken by both groups 5. That study measured race in American English rather than non-native speech, so the transferable finding is the narrow one: recognition error is a property of the speaker as much as of the system, and a garbled answer says nothing about a person until you know which stage garbled it.
The second stage errs too, and it errs on prepared text. In a 2019 study of 100 sets of emergency department discharge instructions containing 647 sentences, machine translation rendered 594 sentences (92%) accurately into Spanish and 522 (81%) into Chinese; 2% of all sentences in Spanish and 8% in Chinese carried errors with potential for clinically significant harm 6. Written, edited, reviewed text, and the accuracy still moved with the language pair. A live interview is the harder case, not the easier one.
Four habits keep the pipe out of your notes:
- Mark it unreadable, not wrong. When an answer arrives incoherent, say so and ask for it again. An unreadable answer written down as a weak one is a measurement error you introduced.
- Put figures and names in writing. Share a document. Ask for the number typed rather than spoken. Digits survive translation; a spoken decimal often does not.
- Send a glossary with the invite. Ten terms the round will use, in English, to every candidate.
- Confirm the term, not the sentence. "When you said the base case, did you mean the 2024 actuals or the plan?" costs fifteen seconds and settles most of it.
Rewrite the round so the evidence isn't a sentence
Ask for things that survive a bad translation: a number, a choice between two named options, a source, an order of operations. "It depends on the discount rate" carries the same information through an interpreter that it carries in English. Then push once on the reason, and once more on what would have to change for the answer to flip.
Five roles, and what to hand over in each:
- Financial analysis. Hand over a two-page packet in which one growth figure is not supported by the filing behind it. Ask which claim they checked, what they found, and what it moved in the recommendation.
- Software engineering. Show a failing test and a stack trace. Ask which file they would open first and what they expect to find in it. A stack trace is the same document in every language.
- Management consulting. Give two entry approaches and one budget. Ask which they would drop first, and what number would have to move for them to swap.
- Healthcare revenue cycle. Give three denials with their codes. Ask which one they would concede, and why that one rather than the other two.
- Marketing. Give them the most on-message statistic in the brief. Ask what population it was measured on, and what they would do if the answer were forty customers.
Write the answers into the same shared document for every candidate, in the same columns, before anybody debriefs. That is what makes a translated round comparable to an untranslated one: the same questions, the same rubric, the same record. The mechanics are the ones behind a structured interview about AI use, where the rubric gets written first, and they carry more weight here, because the part you are most tempted to write from memory is the exact part the layer distorted.
What if the candidate never said they were using it?
Silence here is a gap in your invitation, not a lie. An ordinary interview invite says nothing about whether captions, an interpreter or a translation tool are allowed, so a candidate who brought one broke no rule anybody wrote. Write the line for next time, offer the choice to everyone, and judge this candidate on the answers you already collected.
The line, short enough to paste into the invite:
- This round can run in English, or with live translation or captions. Say which you would like when you book.
- The questions and the rubric are identical either way.
- Where the role carries a language requirement, it is named here, along with the segment that tests it.
Asking up front does more than close the gap. A candidate hides the tool because needing it looks like a cost to them, which is also why disclosure demanded after the call is worth so little: it arrives as an accusation with a form attached. Offering it in writing turns the medium into a choice you made available rather than a secret somebody kept.
It also stops the failure this question usually becomes, where the medium stands in for a doubt about the person. Whether the doubt is about identity, about AI use in the answers, or about candidates recording the interview and running it through a model, each is a separate question with its own answer, and none of them is settled by how the words reached you.
If the job's real question is whether this person catches a confident wrong number before it ships, no interview was ever going to answer it. A translated round, like an untranslated one, catches a candidate describing what they would check rather than checking it. A work sample settles that, in either language.
Common questions
Can you require candidates to interview in English without a translator?
Yes, where the job requires it and you can show why. The EEOC's 2016 enforcement guidance permits an English fluency requirement only if it is required for the effective performance of that position, assessed case by case, and sets a separate two-part standard before an accent may factor into a decision. In practice: name the requirement in the posting, test it as its own short segment on a real task, apply it to every candidate, and keep the record. A requirement that first appears after one interview is the hardest version to defend.
Should candidates have to disclose that they are using AI translation?
Ask in the invite instead of after the call. Offer the round in English, with live translation, or with captions, and say the questions and the rubric are identical either way. Disclosure asked up front is a choice; disclosure demanded afterwards is an accusation with a form attached, and it produces the least honest answer available. Where the role carries a genuine language requirement, name it in the same message along with the segment that tests it, so nobody has to guess what the medium costs them.
How do you judge written communication if the interview was translated?
Test it directly and separately. A short written task in the language the job uses, the same prompt for everyone, in the time the job would allow: a customer reply, a one-paragraph status update, a summary of the packet you already sent. That measures what you need in the register the role writes in, instead of asking a live conversation to stand in for it. Grade it against what the team sends out on a normal day, not against native-speaker prose.
The translation kept dropping technical terms. Does that count against them?
No. Term loss is the classic machine translation failure and it belongs to the pipe, not the person. Send a ten-term glossary with the invite, keep a shared document open for figures and product names, and ask for numbers typed rather than spoken. When a term does go missing mid-answer, confirm the term rather than the sentence: "when you said the base case, did you mean the 2024 actuals or the plan?" Then write down the answer and not the noise.
Is it unfair to the candidates who interviewed without translation?
Not if everyone was offered the same choice and answered the same questions against the same rubric. What matters is comparability of evidence, not sameness of medium. The unfair version is a round where one candidate got a glossary and a shared document and another did not, or where the notes on the translated interview are about how it sounded while the notes on everyone else are about what got decided.
What if they only got through the round because of the tool?
Then the question is whether the job needs the language, which is answerable, rather than whether the interview was real, which is not. Write the language requirement down as a duty or drop it. If it is a duty, test it as one: a short unaided segment on a real task, same prompt for every candidate, judged against what the role does daily. If it is not a duty, the tool did what a well-designed round would have offered anyway, and the answers you collected still stand.
References
- 1. Why don't we believe non-native speakers? The influence of accent on credibility ✓ accent-american.com Native English listeners rated identical trivia statements less truthful when the speaker had a mild accent (M = 6.95) or a heavy accent (M = 6.84) than when the speaker was native (M = 7.59), on a 14 cm scale; speakers were reciting statements written by the experimenter, and awareness of the processing difficulty corrected the mild-accent effect but not the heavy-accent one.
- 2. EEOC Enforcement Guidance on National Origin Discrimination ✓ eeoc.gov Issued 11-18-2016. Sections V.A and V.B: an English fluency or proficiency requirement "is permissible only if required for the effective performance of the position for which it is imposed" and the degree of fluency lawfully required varies from one position to the next (V.B); an accent-based decision requires evidence both that effective spoken English is required for the duties and that the accent materially interferes with spoken communication (V.A).
- 3. 29 CFR Part 1606, Guidelines on Discrimination Because of National Origin (sections 1606.1 and 1606.6) ✓ govinfo.gov Section 1606.1 defines national origin discrimination as including denial of equal employment opportunity because an individual has "the physical, cultural or linguistic characteristics of a national origin group"; section 1606.6 applies the Uniform Guidelines on Employee Selection Procedures when investigating a selection procedure for adverse impact on the basis of national origin.
- 4. 29 CFR Part 1607, Uniform Guidelines on Employee Selection Procedures (section 1607.16, Definitions) ✓ govinfo.gov Section 1607.16(Q) defines a selection procedure as "Any measure, combination of measures, or procedure used as a basis for any employment decision", covering assessment techniques "through informal or casual interviews and unscored application forms".
- 5. Racial disparities in automated speech recognition ✓ pmc.ncbi.nlm.nih.gov Five commercial ASR systems from Amazon, Apple, Google, IBM and Microsoft, transcribing 19.8 hours of structured interviews with 42 white and 73 black speakers: average word error rate 0.35 for black speakers against 0.19 for white speakers, with the gap holding on identical phrases spoken by both groups.
- 6. Assessing the Use of Google Translate for Spanish and Chinese Translations of Emergency Department Discharge Instructions ✓ pmc.ncbi.nlm.nih.gov 100 sets of emergency department discharge instructions containing 647 sentences: 594 (92%) accurately translated into Spanish and 522 (81%) into Chinese, with 2% of Spanish and 8% of Chinese sentence translations carrying potential for clinically significant harm.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.