Pipeline

Why Do Candidates Who Interviewed Brilliantly Struggle in Their First Quarter?

A candidate who interviewed brilliantly, then struggled in the first quarter usually faked nothing: the round sampled the one state the job almost never occupies, supervised, time-boxed, on a problem you framed, with nothing riding on being right. The quarter is unsupervised work on a confident draft nobody else reads. Some of the gap is the instrument: structured interviews top out near .42 validity, so one hard quarter is ordinary. The rest tracks how much of the role starts as a generated draft, weeks in some jobs, never in others.

The takeInterviews got harder to read at the same moment everything else did. A round without a rubric rewards whoever sounds most settled, and the settled sound has probably never been cheaper to produce. The loop is not too short. It is pointed at the wrong minute: it watches the part of the work that was always going to look good and skips the part that was always going to be boring. Verification is not a performance, which is exactly why it is worth watching and exactly why a round almost never does.

Where Olive fits

Open a role and see what the work shows

Olive is priced per attempt rather than per seat, so a role whose first quarter keeps going wrong can be piloted beside your existing loop (ten attempts a month cost nothing) and the two compared against what the hires actually did. An attempt returns six evidenced findings about one candidate, written by a human reviewer, as an input to your decision rather than a verdict on the person.

Rank your shortlist

Why does a brilliant interview predict so little about the first quarter?

Because the interview samples the one state the job almost never occupies: supervised, time-boxed, on a problem someone else framed, with an audience and nothing riding on whether the answer is true. The first quarter is unsupervised, open-ended, and full of confident drafts that have to be checked by the person who ordered them. A round can watch someone reason. It cannot watch them verify.

Start with the ceiling, because it reframes the surprise. In the revised validity estimates that replaced decades of overcorrected meta-analysis, structured interviews are the best-ranked widely used predictor at a mean operational validity of .42, with an 80% credibility interval running from .18 to .66, and unstructured interviews at .29 1. That is the best case: a proper structured round, scored against a rubric. A correlation at that level says something real about a group of hires and very little that is certain about any one of them. A brilliant interview followed by a hard quarter is not an anomaly in that data. It is the ordinary width of the interval.

The second mechanism is that people cannot reliably report their own process. In a randomized trial, sixteen experienced open-source developers took 19% longer to complete issues when AI tools were allowed, and still believed afterwards that the tools had sped them up by about 20% 2. They were working on their own repositories, on real issues, with nobody interviewing them. If self-assessment fails under those conditions, a candidate narrating how they work with AI in a 45-minute round is not a witness you can cross-examine into accuracy. They are describing a process they cannot see.

The third runs the wrong way from what a round rewards. Across 319 knowledge workers describing 936 first-hand examples of AI-assisted work, higher confidence in the AI predicted less critical thinking, while higher confidence in one's own ability at the task predicted more 5. Interviews reward the first kind, because it sounds like fluency. The quarter charges for the absence of the second. The polished answer everyone now gives is exactly the surface on which those two are indistinguishable.

Which part of the job is unsupervised generation?

The part where a model produces the first version and nobody else reads it before the work leaves your hands. Count it as a fraction: of the deliverables a person ships in a normal week, how many now begin as a generated draft, and how many of those get a second reader? That fraction is the share of the job your interview did not sample, and it is not the same number in two roles.

Adoption is broad but uneven. As of late 2024, 23% of employed US respondents had used generative AI for work in the previous week and 9% used it every workday 3. Every workday is the figure that matters here: it marks the roles where the first version of most artifacts is now machine-produced, so a weak framing habit compounds daily rather than monthly.

The shape of the use differs too, not just the amount. A classification of 200,000 anonymized Copilot conversations against O*NET work activities found AI applicability cutting across sectors, because most occupations contain information work. It also separated occupations likely to delegate a task to AI outright from those using it to assist an existing workflow 6. Delegation and assistance fail differently. A delegated task fails when nobody framed it; an assisted task fails when nobody checked it.

Two roles make that concrete. A marketing brief starts as generated copy, gets read by three people who are not its author, and its weakest point is usually the most on-message statistic, the one least likely to have been opened. That fails in weeks, in public, as a positioning line nobody can source. A diligence memo also starts as a draft, but the check that decides it, recomputing a figure against the filing, requires a practitioner's judgment about which figure matters at all. That fails later and more quietly, in a recommendation that reconciles to nothing. Same tools, different interval, different tell. Work out what AI actually does in the role before deciding which round in your loop is underweight.

What actually breaks between week two and week ten?

Four things, in roughly that order: the brief never gets framed, the confident claim never gets checked, the judgment that should have been kept gets handed over, and nothing the model produced is ever refused. None of the four is visible in an hour with an interviewer. All four are expensive at forty hours a week with a deadline attached.

  • The brief was never framed. In the round you supplied the problem. On the job the request arrives vague, and the model resolves that vagueness confidently and in the wrong direction. The tell: a finished, well-made deliverable that answers a question nobody asked, arriving early.
  • The claim was never checked. Almost-right is the dominant failure mode, not obvious error. In the 2025 Stack Overflow developer survey, the most common frustration with AI tools, at 66%, was output that is almost right but not quite, and 46% of developers said they distrust the accuracy of AI output against 33% who trust it 4. The tell: a number, a citation or a claimed behavior in the work that nobody can say where they got.
  • The judgment was delegated. Total delegation reads as speed in week one and as a person who cannot answer a question about their own work in week six. The tell: they go quiet when the assistant is wrong or unavailable, and their rework arrives with no account of what changed.
  • Nothing was ever refused. Accepting everything is the fastest route to a finished artifact, which is exactly what a timed round selects for. The tell: fluent editing and zero rejection (prose gets tidied, framings never get thrown out).

What those share is that each is an act, with a time and a place, that only exists once the work is long enough to be boring and real enough to cost something. An interview can capture a candidate describing all four with total conviction. The follow-up question is where the description thins out, and it is the cheapest correction available inside a loop you are not allowed to lengthen.

Make the round predict work, not performance

Change what the round asks for. An interview that asks a candidate to perform reasoning measures performance; an interview that hands them finished, flawed work and fifteen unstructured minutes measures whether they check. Producing more output cannot score anything once the artifact already looks done, so checking becomes the only move available. That one change converts an audition into a sample of the behavior the quarter bills for.

Build the exercise out of your own material. Take a real deliverable from the role, have a model produce it, then plant one error only a practitioner would catch: a figure on the wrong basis, a source that does not say what the paragraph claims, a dependency that does not do what its name implies. No preparation reaches your error, which is the point, because coaching reliably produces fluency about AI work, not the instinct to open the filing.

Write the rubric before the round, in the words of someone who has done the job: what strong, adequate and weak look like on each of the four behaviors above. Without that sheet, panels fall back on polish and confidence, which is what the interview already measured. Score acts rather than impressions (what was framed, what was opened, what was kept, what was refused) and record the reason beside each rating, because the reason is the only part another interviewer can audit.

Two smaller changes cost nothing at all. Say in the invitation that AI is expected rather than tolerated, so you stop measuring how well a candidate guesses your policy. And ask about one specific task with a date attached instead of a general account of how someone works; generalities are rehearsed, a particular Tuesday is not. If the loop cannot absorb another stage, add the judgment signal without lengthening the loop by replacing a round rather than appending one.

What can an interview never tell you?

Whether the person does any of it when nobody is watching. Every answer in a round is a self-report from someone with an obvious stake in it, and people misread their own process even while doing real work on their own repositories 2. Closing that gap takes a longer sample with a genuine cost of being wrong, or a first-quarter measurement you actually collect and feed back into the loop.

So instrument the other end. Take four deliverables from a new hire's first eight weeks and record how many rounds of correction each needed and of what kind: wrong question, unsourced claim, no account of what the person did themselves. Three or four of those files per role, read side by side, tell you which stage of your loop is not doing its job. The answer is usually specific, along the lines of nothing here tests verification, rather than a general verdict that the interview was too easy.

Be careful what you conclude from that record. First-quarter output is confounded by manager, onboarding, project difficulty and luck, and a handful of hires will never separate them; what an assessment score predicts about performance is a claim that needs its own evidence rather than one you can borrow. Use the first-quarter record to find the behavior your loop never sampled, not to grade individual hires.

The last honest thing to say is that part of this gap does not close inside a hiring process. An interview buys 45 minutes; the quarter buys 500 hours, and no sample of the first predicts the second cleanly. The realistic ambition is smaller and worth having anyway: make one of those 45 minutes contain a real act (one check, actually performed, on material that arrived finished and wrong), because that is the smallest thing a round can hold that the quarter will also ask for.

See a sample report

Common questions

Did the candidate oversell in the interview?

Usually not. The round measured something real: structured reasoning, communication, and how a person handles a problem someone else framed. That ability exists. It just does not include the behavior the quarter bills for, which is deciding what a confident, finished draft got wrong while nobody is watching. Treat it as an incomplete measurement rather than a deception. The correction is not more suspicion in the round; it is one exercise where the artifact arrives already finished, so producing more of it cannot earn anything.

How long does the first quarter take to expose this?

Usually the third or fourth week, when onboarding tasks with known answers give way to real requests with vague briefs and nobody between the work and its recipient. The earlier signal is quieter: a new hire whose first deliverables arrive fast and then need three rounds of correction, or who cannot say where a number in their own work came from. Watch the correction rate on the first four deliverables rather than the speed of the first one.

Should structured interviews be dropped, then?

No. Structured interviews remain the best-validated widely used predictor available, and they beat the unstructured version by a wide margin, so running a looser round is strictly worse. The point is what a validity coefficient means: real signal about a group of hires, wide uncertainty about any single one. Keep the structured round, narrow what you conclude from it, and add one exercise that reaches a behavior an interview cannot: checking a claim, refusing a direction, or keeping work back from the model.

Which roles show the gap fastest?

The ones where a model writes the first version of most deliverables and the work leaves the building without a second reader. Marketing, content, support and analyst-adjacent work tend to surface it within weeks, because the artifact is public and its weakest claim is often the most quotable one. Roles with a slower review cycle (audit, underwriting, diligence) hide it longer and cost more when it lands, since the error is found downstream of a decision that already used it.

Can an assessment close the gap?

It narrows it. A longer, role-grounded task done with an assistant puts a candidate in the state the interview skipped: unsupervised, holding a confident draft, with a real reason to check it. Olive works that way: a 40-to-60-minute occupational assignment, then a human reviewer who writes six findings, each carrying the moment in the session it rests on. The candidate is granted the same report the employer reads, and the report makes no hiring recommendation. It is an input to your decision rather than the decision.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. doi.org Structured interviews top the revised list of widely used predictors at a mean operational validity of .42, with an 80% credibility interval running .18 to .66; unstructured interviews are estimated at .29.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial with 16 experienced open-source developers: 19% longer to complete issues with AI tools allowed, while the developers believed afterwards that AI had sped them up by about 20%.
  3. 3. The Rapid Adoption of Generative AI (NBER Working Paper 32966) Bick, Blandin and Deming, National Bureau of Economic Research, 2024. nber.org As of late 2024, 23% of employed respondents had used generative AI for work at least once in the previous week and 9% used it every work day.
  4. 4. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  5. 5. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ACM CHI Conference on Human Factors in Computing Systems (CHI '25), 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in generative AI is associated with less critical thinking, while higher confidence in one's own task ability is associated with more.
  6. 6. Working with AI: Measuring the Applicability of Generative AI to Occupations Tomlinson, Jaffe, Wang, Counts and Suri (Microsoft Research, arXiv), 2025. arxiv.org 200,000 anonymized Copilot conversations classified against O*NET work activities; applicability cuts across sectors because most occupations contain information work, and the method distinguishes occupations likely to delegate tasks to AI from those using it to assist existing workflows.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.