Interviewing

Brilliant in the Screen, Flat Onsite: Which Candidate Is Real?

When a candidate is brilliant in the remote screen and flat in the onsite, neither version is the real one. The two rounds are different tests, and the medium moves interviewer ratings before the candidate does anything. Weight the round whose conditions resemble the job: the screen if the work happens at a keyboard with tools open, the onsite if the work is a room full of people. Help that was never banned in writing before the call broke no rule; ask what it produced and what she checked.

The takePanels almost always decide the weaker round was the real one, and that instinct deserves less trust than it gets. Difficulty reads as authenticity: the round that was harder to sit through feels truer, so the conference room wins the debrief whatever the job looks like. Nobody has measured that preference directly, though it would explain how readily a remote round gets discounted where the desk, not the room, is the job. A loop that ends "she was better on the phone, so watch out" is grading its own conference room. If the room is not the job, the room is not the evidence.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing what she did with an assistant, in either format, but it cannot capture her doing it. Olive puts that in front of her as work: a 40-to-60-minute assignment grounded in the occupation, with an AI assistant that will overreach, returned as six findings a human reviewer writes with the moment behind each one attached, and the candidate is granted the same report.

Rank your shortlist

Which round should you believe?

The one whose conditions look like the job. If the work happens at a screen with a browser, an editor and an assistant open, the remote round is the closer sample and the conference room is the artificial one. If the job is a room, a whiteboard and four people disagreeing, the onsite is the sample. Pick before you reread your notes, so the choice is not a rationalization of the impression you already formed.

Which round is closer changes by occupation, and it is worth being literal about it:

  • Software engineering. The daily instrument is an editor, a terminal, documentation and an assistant that will write the whole change if nobody stops it. A marker at a whiteboard tests recall of syntax nobody types from memory any more.
  • Financial analysis. The memo gets built in a spreadsheet against a filing somebody has to open. A case discussion with no file open measures how well a person talks about analysis, which is a different skill and a real one.
  • Underwriting, claims and revenue cycle. Queue work, alone, at a screen, all day. A panel round measures almost nothing the job asks for.
  • Management consulting. Rooms are half the job, so the onsite is a genuine sample of that half. The other half is done on a laptop at 11pm.
  • Sales, teaching, clinical work, anything where the room is the deliverable. The onsite is the sample and the remote screen was the artificial one. The gap runs the other way and means the opposite thing.

If the role genuinely contains both, then you have one observation of each and no contradiction between them. What you know is that this candidate is stronger with tools and time than in front of strangers with neither. That is information about how to deploy someone, not a verdict on which version was authentic. It is also why candidates who interview brilliantly can struggle in their first quarter: the round that impressed you was not measuring the conditions the work imposes.

How much of the gap is your interview format?

More of it than a debrief usually credits, and the direction is not fixed. A meta-analysis of twelve studies (13 samples, 1,557 interviewees) found interviewer ratings lower in technology-mediated interviews than in face-to-face ones, a moderate negative effect (d = -.41), with applicant reactions lower too (d = -.36) 1. The medium moved the ratings before any candidate did anything.

Nerves in the room is the next explanation a panel reaches for, and one controlled comparison does not support it. Eighty-eight interviewees were randomly assigned to a face-to-face, telephone or videoconference interview, with strain measured physiologically as well as by self-report. There were no differences between the three media on strain or interview anxiety, and ratings of interviewee performance were still lower in the technology-mediated conditions 2. Strain did not separate the formats. The ratings did.

Which direction the format pushes belongs to your design, not to the medium. Across 1,026 applicants to one pharmacy program, candidates in the virtual multiple mini interview scored higher than the in-person cohort on three stations, with medium effects on two teamwork stations (D = 0.44 and D = 0.47) and a small one on integrity (D = 0.28) 3. Same intent, opposite direction from the meta-analysis. Nobody can tell you in advance which way your own loop leans, and the only thing that settles it is your own ratings.

And the medium was not the only thing that changed between the two rounds. Count the rest:

  • Who asked. One recruiter who runs screens daily, then four people who interview twice a year.
  • What was asked. A screening script, then role-specific depth nobody calibrated in advance.
  • What was allowed. Notes, a browser, a second monitor, then a marker.
  • The audience. One listener, then a panel watching each other watch her.
  • The clock. Forty-five minutes, then six hours and a lunch that is also being observed.

Everything moved at once and one impression came back different. There is no honest way to attribute that to the candidate.

What if the remote round had AI in it?

Assume it did, and notice how little that settles. Unless the invitation said not to, in writing, before the call, there was no rule to break. A second window with an assistant in it is ordinary equipment. What you want to know is what the candidate did with what it produced, and the onsite round is the one that removed the evidence rather than revealing it.

How ordinary depends on the job, which is worth checking instead of assuming. In nationally representative US surveys reported in 2024, 23% of employed respondents had used generative AI for work at least once in the previous week and 9% used it every work day 4. That is a real share and not a universal one. If it describes the role you are filling, the assisted round is the closer sample of the work. If the role's work genuinely happens without a model in the loop, the onsite is.

The delta itself tells you nothing about assistance. A candidate who prepared hard, who writes better than she speaks, who had a bad morning, or who faced four people at once produces exactly the same shape. Inferring AI use from a performance drop is the same move as inferring it from prose style, and it lands hardest on people whose fluency does not match what an interviewer expected. If the question is live and specific, answer it directly: whether candidates are using AI to answer your questions during live interviews has a real answer, and a candidate's eyes moving off camera is not it.

What is readable is the substance of the remote answers, reread with the polish set aside. A fluent answer with nothing behind it fails on its own terms: no claim traced to a source, no alternative weighed, false precision where she could not have had the data. Name that if it is there, and name it about the answer rather than about the tool. Every candidate suddenly giving the same polished STAR answer is the version of this problem waiting in your next loop anyway.

Ask about the delta, don't infer it

Book twenty minutes and put the two rounds side by side out loud, with no accusation in the sentence. Say what you observed, then ask about the work rather than about the tool: what a specific claim from the screen answer rests on, what she would check first if that number were wrong, and what was different for her between the two rooms. Specifics arrive fast from someone who did the thinking.

Three questions carry it:

  • "In the screen you said the migration would take a quarter. What is that resting on, and how would you check it?"
  • "What did you use to prepare, and where did it get something wrong?"
  • "Which of the two conversations was closer to the work you actually do?"

Ask all three of every finalist, not only of the one whose rounds disagreed. A question asked of one person is a different test given to one person, and it is the version that becomes expensive to explain later. The answers are also usable either way: someone who did the thinking names the source and the correction inside a minute, and someone who did not moves to generalities and stays there.

Better than asking is watching a little of it happen. A short working session in the format the job uses, with the tools the job supplies, run identically for every finalist, retires the question entirely, because both rounds become the same test. Redesigning the interview so AI assistance becomes signal instead of cheating is the general form of that fix. See how Olive measures this.

Write down what would change your mind before the session, not after. "She defends the screen answer with a source she opens in front of me" is a check that can fail. "She seemed more confident this time" is the impression you already had, arriving twice.

Fix the loop before the next candidate splits

Hold the medium constant inside a stage. The meta-analysis that measured the mode effect closes on exactly this point: organizations should be wary of varying interview mode across applicants, because inconsistency in administration raises fairness questions 1. If some finalists get a video call and others get a room, part of what you are comparing is your own calendar.

If a loop runs two formats on purpose, ask one identical question in both. Without a shared item, "stronger in the screen" has no unit: two different sets of raters heard two different conversations, and subtracting them produces a number about the loop. One repeated question turns a mood into an observation with something to hold against it.

Then keep one rubric across both rounds, written before either happens, with the evidence each rating needs named in it. Consistency is doing more work than it used to, because the answers themselves have converged: a well-prepared candidate now arrives with a coached version of every standard question, which is why a structured interview no longer separates candidates the way it did without something behind the questions.

Do all of that and neither round is predictive on its own. A screen is forty-five minutes on a laptop, an onsite is one afternoon in front of strangers, and both are samples of behavior under conditions you built. The useful thing to do with a candidate who split across them is to write down which conditions the job actually imposes, run one round under those conditions for everyone, and let the split stop being mysterious.

See how it works

Common questions

Does a big gap between the remote screen and the onsite mean the candidate cheated?

No. That gap is what a change of medium, interviewers, questions, tools and audience produces on its own, and interviewer ratings are known to move with the medium before the candidate does anything. Cheating is a specific claim: a rule stated in writing before the round, and a candidate who broke it or denied it when asked directly. Absent that, you have two observations taken under different conditions and one question worth asking out loud. Ask what she used to prepare and where it got something wrong, and let the answer decide it.

Which round predicts on-the-job performance better?

The one whose conditions resemble the job. For work done alone at a screen with a browser and an assistant open, the remote round is the closer sample and the panel round is the artificial one. For work done in rooms full of people who disagree, it runs the other way. Neither is predictive by itself: a screen is forty-five minutes and an onsite is one afternoon, both under conditions you constructed. Pick the closer sample before you reread your notes, so the choice is not assembled out of the impression you already formed.

Should you re-run the onsite in the same format as the remote screen?

Only if you re-run it for everyone at that stage. A second chance offered to one candidate is a different test given to one person, and it is the version that is hard to explain afterwards. If the format is the problem, fix the stage: same medium, same questions, same tools, for every finalist in the round. If you want one more observation of this candidate specifically, make it a short working session in the format the job uses, and give the same session to the others.

What do you ask a candidate whose two rounds disagree?

Point the questions at the work, not at the tool. Three questions do it: what a specific claim from the screen answer rests on and how they would check it, what they used to prepare and where it got something wrong, and which of the two conversations was closer to the work they actually do. Someone who did the thinking answers in specifics quickly. Someone who did not moves to generalities and stays there. Ask the same three of every finalist, so the round stays comparable and the record shows that it was.

Is it unfair to give one candidate a remote round and another an onsite?

It is at least inconsistent, and a 2016 meta-analysis of technology-mediated interview studies warns organizations to be wary of varying interview mode across applicants because inconsistent administration raises fairness questions. Ratings shift with the medium, so two candidates in two formats were not measured on the same scale. Fix it at the stage rather than per candidate: decide the medium for that round, apply it to everyone in it, and record what was asked. A candidate who needs a different format as an accommodation is a separate process, handled on its own terms.

What if the candidate was simply nervous in the room?

Possible, and harder to assume than it looks. In a 2021 comparison of face-to-face, telephone and videoconference interviews with 88 interviewees, strain and interview anxiety did not differ across the three media, yet performance ratings were still lower in the technology-mediated ones. Nerves are real in individual cases and they are not the default explanation for a format gap. If you think the room cost her, test the substance: the same claims and the same reasoning, delivered less fluently. Substance that thinned out is a different finding from delivery that did.

References

  1. 1. Technology in the Employment Interview: A Meta-Analysis and Future Research Agenda Personnel Assessment and Decisions, Blacksmith, Willford and Behrend, 2016. scholarworks.bgsu.edu Meta-analysis of twelve studies (K=13 unique samples, N=1,557): mean effect of interview medium on interviewer ratings d=-.41 and on applicant reactions d=-.36, both lower in technology-mediated interviews; the authors write that organizations should be especially wary of varying interview mode across applicants because inconsistency in administration could lead to fairness issues.
  2. 2. A Comparison of Conventional and Technology-Mediated Selection Interviews With Regard to Interviewees' Performance, Perceptions, Strain, and Anxiety Frontiers in Psychology, Melchers, Petrig, Basch and Sauer, 2021. pmc.ncbi.nlm.nih.gov 88 interviewees randomly assigned to a face-to-face (n=30), telephone (n=26) or videoconference (n=32) interview: no differences between the three media on psychological or physiological indicators of strain or interview anxiety, while ratings of interviewee performance were lower in the technology-mediated interviews.
  3. 3. Validity evidence for a virtual multiple mini interview at a pharmacy program BMC Medical Education, Hammond, McLaughlin and Cox, 2023. pmc.ncbi.nlm.nih.gov 1,026 applicants to one pharmacy program (438 in-person, 588 virtual): virtual candidates scored higher on the teamwork-giving (D=0.44), teamwork-receiving (D=0.47) and integrity (D=0.28) stations, the opposite direction to the technology-mediated meta-analysis.
  4. 4. The Rapid Adoption of Generative AI National Bureau of Economic Research Working Paper 32966, Bick, Blandin and Deming, 2024. nber.org Nationally representative US surveys: 23 percent of employed respondents had used generative AI for work at least once in the previous week, and 9 percent used it every work day.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.