Interviewing

Are Candidates Using AI During Live Interviews?

Some candidates are answering live interview questions with AI, and you cannot count them: interview copilots are shipping products that run in a window on the candidate's own machine your call never sees. They are strongest on questions with one public answer: definitions, complexity bounds, a named framework. They hold nothing about a decision the candidate personally made, so they invent, and the invention dies on your second question. People read AI text at chance, and a hunch you act on is a selection procedure you have to defend.

The takeThe copilot didn't break the live interview. It sent the bill for something that was already true: a round built on questions any well-read stranger could answer was never measuring the person in front of you, and prep sites were selling those answers a decade ago. The honest guess is that the surviving loops get smaller and stranger, and that the ones answering this with gaze tracking lose their best candidates first. A rehearsed answer and a retrieved one look identical from the outside, so the suspicion lands on whoever prepared hardest and whoever is speaking their second language. That is a hiring filter nobody chose.

Where Olive fits

Open a role and see what the work shows

A live round can capture a candidate saying how they would check a confident claim; it cannot capture the check itself. Olive puts that in front of them as work: a 40-to-60-minute occupational assignment, an assistant willing to do all of it, and a human reviewer who writes six findings with the moment each one rests on.

Rank your shortlist

Which tools are candidates actually using?

The category is called an interview copilot, and the products are real businesses. Cluely, built by the two Columbia students behind Interview Coder, raised a $15 million Series A led by Andreessen Horowitz in June 2025, at a valuation investors put near $120 million 1. What it sells is a hidden in-browser window that the interviewer or test giver cannot view 7.

That mechanic decides what you can do about it. A window the call cannot see is also a window screen sharing does not show you, because screen sharing shows the surface the candidate chose to share. The rest follows from where the software runs rather than from anything a vendor has published: what sits on the candidate's own display never crosses to your side of the call, so there is nothing in your video vendor's logs to find.

Underneath the named products sits a much larger grey market that needs no product at all: a second laptop, a phone propped below the camera, a chat window open behind the call. Above them sits identity fraud proper. The FBI's advisory on deepfaked applicants for remote positions describes interviews where lip movement does not coordinate with the audio, and where coughing or sneezing does not line up with what is on screen 2. That is a different problem with different remedies, and conflating the two is how interviewers end up policing nerves.

So the honest answer to whether it is happening is yes, at a rate you cannot measure, in a round that gives you nothing to measure it with. Which makes the useful question not how many, but which of your questions are answerable this way.

Which questions can an assistant answer for them?

Anything with a canonical answer. A definition, a complexity bound, the difference between two named patterns, the standard framework for a market-entry case, the textbook tradeoff between two architectures. Each has a public right answer, and a model returns it in seconds, phrased better than most candidates manage under pressure. A definitions-heavy screen for an engineer or an analyst is now close to uninformative.

Sort your current question set into two piles. The first pile is everything a well-read stranger could answer without knowing the candidate:

  • What is the difference between X and Y?
  • How would you approach a problem of this shape?
  • Which metric would you track for a product like this?
  • Walk me through how that algorithm, framework or process works.
  • What are the tradeoffs of this design?

That pile was already the weak half before any of this. It rewarded preparation over practice, and interview-prep sites published its answers a decade ago. What changed is that the preparation no longer has to happen in advance, which removes the last thing the question was accidentally measuring, which was whether the candidate cared enough to study. If your loop leans on this pile, the polish is not a candidate signal, and uniformly excellent answers are the tell.

The same holds for a live coding round built on a problem with a published solution. The realistic response is not to hunt for a harder puzzle, which lasts one quarter, but to open the assistant deliberately and watch what the candidate does with it. Running a coding interview with AI open is its own piece of work.

Which questions are unhelpable?

Questions about something the candidate personally did. A decision they made, what they knew when they made it, what they traded away, what happened afterward, and what they would do differently. No assistant holds that record. OPM defines a structured interview as systematically asking about a candidate's behavior in past experiences, their proposed behavior in hypothetical situations, or both 3, and the past-experience half is the half that still holds.

Four that travel across roles, because each needs a fact only the candidate has:

  • The undo. Tell me about something you shipped and then had to pull back. What made you pull it?
  • The number you stopped trusting. When did you last decide a figure in your own analysis was wrong, and how did you settle it?
  • The thing you should have kept. What did you hand to a model in the last month that you should have done yourself, and how did you find out?
  • The disagreement. Name a call your manager made that you thought was wrong. What did you do about it?

An assistant will generate a plausible story for any of these. What it cannot do is survive the second question, because the second question depends on what the candidate just said. Ask which number moved, by how much, and who else saw it. A retrieved answer has depth of exactly one, and the pause before the probe is the part that shows. Follow-up questions expose understanding faster than any transcript analysis will.

One more family is worth adding: questions about work the candidate has already submitted. Hand back their take-home or a portfolio piece, change one constraint, and ask what breaks. That cannot be prepared and cannot be retrieved, and anyone who did the work answers it in ninety seconds.

Should you watch for tells instead?

No, not as the main instrument. The cues people trust run backwards. Across six experiments with 4,600 participants, people identified AI-generated professional self-presentations at 50-52% accuracy, barely above chance, and the texts carrying grammatical problems (which participants were more likely to rate as machine-written) were in fact 15% less likely to be AI-generated 4.

The live-interview tells are worse, because each has an innocent cause more common than the guilty one. Eyes moving off camera: notes, a second monitor, a nervous habit, an attention difference. A pause before answering: processing in a second language, a bad connection, a stutter, thinking. A register that turns suddenly formal: a rehearsed answer, which is the thing you asked them to prepare. Eyes off camera are not evidence on their own.

Running the transcript through a detector afterwards does not rescue it. An independent review of fourteen detection tools found them neither accurate nor reliable, biased toward calling text human-written, and badly degraded by ordinary rewording 6. On a speech transcript the input is already a machine's guess at what was said. The longer version is in whether AI detectors work on interview transcripts.

Exposure is the other cost. A suspicion you act on is a selection procedure, and EEOC guidance holds that a procedure used to make an employment decision must be job-related and consistent with business necessity where it screens out a protected group 5. "The answer sounded too smooth" is job-related to nothing, and the people it screens out are the ones who prepared hardest and the ones whose speech does not match your default.

Change the round, not the surveillance

Two changes carry most of the weight. State the rule in the invitation, in one sentence, so nobody has to guess. An unstated rule gets guessed at, and the guessing tracks how much coaching a candidate has had. Then rewrite the question set so most of it asks about work the candidate has already done, with the follow-up written before the round rather than improvised in it.

The rest is bookkeeping:

  • Pick a side and publish it. Banned, allowed, or allowed and scored are all defensible; unstated is not. Letting candidates use AI in the interview is a decision to make once, not per interviewer.
  • Score against a written key. Decide what a good answer contains before the first candidate, apply it to everyone, and record why each answer landed where it did. That record survives a challenge. A hunch does not.
  • Give every question one prepared probe. The probe is the instrument; the question is only the setup.
  • Move canonical-answer questions out of the live round. If a fact matters, check it in writing in two minutes, or accept that it gets checked on the job.

What none of this buys you is the thing the interview was standing in for: whether this person works well with an assistant. A live round can capture a candidate describing how they would check a confident claim. It cannot capture them checking one, because there is nothing in the room to check against. That gap is what a work sample with the assistant open is for. See how Olive measures this.

See how it works

Common questions

Can you tell from the video that a candidate is using AI?

Not reliably. There is no window into the candidate's own machine, and the behavioral cues (eyes off camera, a pause, a formal register) have innocent causes far more common than the guilty one. People asked to judge AI-generated professional text managed 50-52% accuracy, and the cues they trusted pointed the wrong way. A live deepfake is a different matter and does leave visible artifacts, but that is identity fraud rather than answer assistance.

Is a candidate using AI during an interview cheating?

Two things decide it: what you told them, and what the job is. If the role involves working with a model all day, watching someone do it is information rather than a violation, provided you said the round allows it. If a question exists to test unaided recall, say so before the round starts and pick a format where the answer is checkable: their own prior work, a decision they made, a probe no script survives.

What should the interview invitation say about AI use?

One sentence, plainly, before they prepare. Name what is allowed in each round and say how it will be read: assistant open and part of what gets assessed, or closed and the answer has to be theirs. Candidates told nothing fall back on what their last three interviews implied, which is a fact about their coaching and not about their work. Say it and the round becomes readable, because how a candidate used an assistant only means something once they knew it was on the table.

Do live coding rounds still tell you anything?

Only if the problem has no published solution, or if the assistant is open on purpose. A closed round built on a standard algorithm question now measures whether the candidate has a second screen. The version that still works hands over a small unfamiliar codebase, lets them use whatever they use at work, and watches which parts they refuse to delegate and what they check before calling it done.

Does Olive tell you whether a candidate used AI during an interview?

No, that is not what it looks at. Olive is an employer-purchased assessment: the candidate works a task from their own occupation with an AI assistant available, and a human reviewer writes six findings on what happened: how the problem was framed, what evidence was demanded, what was kept, what was built in between, what was refused, and what was tested outside the conversation. There is no composite number, and the candidate is granted the same report the employer reads.

References

  1. 1. Cluely, a startup that helps 'cheat on everything,' raises $15M from a16z TechCrunch, 2025. techcrunch.com Cluely was founded by Roy Lee and Neel Shanmugam, suspended from Columbia over Interview Coder; $15M Series A led by Andreessen Horowitz in June 2025, valuation put near $120M by two investors, after a $5.3M seed.
  2. 2. Deepfakes and Stolen PII Utilized to Apply for Remote Work Positions FBI Internet Crime Complaint Center (IC3), Public Service Announcement I-062822-PSA, 2022. ic3.gov Law-enforcement record of deepfaked candidates in live online interviews: lip movement not coordinating with audio, and coughing or sneezing not aligned with what is presented visually.
  3. 3. Structured Interviews U.S. Office of Personnel Management, 2026. opm.gov Defines a structured interview as systematically inquiring about behavior in past experiences and/or proposed behavior in hypothetical situations, with the same predetermined questions and rating scale for every candidate.
  4. 4. Human heuristics for AI-generated language are flawed Proceedings of the National Academy of Sciences, 2023. pmc.ncbi.nlm.nih.gov 4,600 participants across six experiments identified AI-generated self-presentations at 50-52% accuracy (51.2% in the professional context); self-presentations with grammatical issues were 15% less likely to be AI-generated, the reverse of the heuristic participants used.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2024. eeoc.gov A selection procedure that screens out members of a protected group must be shown job-related and consistent with business necessity.
  6. 6. Testing of detection tools for AI-generated text International Journal for Educational Integrity, 2023. edintegrity.biomedcentral.com Fourteen tools tested: neither accurate nor reliable, biased toward classifying output as human-written, and significantly degraded by content obfuscation.
  7. 7. Columbia student suspended over interview cheating tool raises $5.3M to 'cheat on everything' TechCrunch, 2025. techcrunch.com Describes the product mechanic: Cluely offers users the chance to cheat on exams, sales calls and job interviews through a hidden in-browser window that cannot be viewed by the interviewer or test giver.

7 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.