Teams

Why Do Candidates Who Demo Brilliantly With AI Fall Apart in Month One?

Candidates who demo brilliantly with AI and then fall apart in month one usually faked nothing. The demo measured production under supervision, and the job bills for verification nobody watches. In an hour, with an audience and a brief you wrote, four behaviors stay optional: framing the problem, checking a claim outside the conversation, keeping back the work that shouldn't be handed over, refusing bad output. Month one makes all four load-bearing. The exception is a role that really is mostly production, where the demo predicts the job fine.

The takeThe uncomfortable part is that the demo round is doing exactly what it was built to do. It was designed when producing the artifact was the hard part, and producing the artifact is no longer the hard part. So panels keep grading the one thing that got cheap, then blame the person when month one arrives. I'd expect most demo rounds to survive this anyway, because they are pleasant: everyone leaves impressed, and nobody has to sit through fifteen minutes of someone quietly reading. That comfort is the cost. A round that never feels tedious is measuring the wrong hour.

Where Olive fits

Open a role and see what the work shows

The same six dimensions name what breaks in month one: framing before generating, demanding a source for the claim that matters, keeping the judgment that shouldn't be handed over, and testing a claim against something outside the conversation. Olive reads those from a recorded occupational session that a human reviewer writes up, rather than from a demo the candidate controlled.

Rank your shortlist

Why does a brilliant AI demo predict so little?

Because a demo is a supervised hour with a known answer and no downstream cost, and month one is none of those things. The candidate performs to an audience, on a problem you framed, with nothing riding on whether the output is true. Every behavior that fails in week three (checking a claim, refusing a direction, keeping the part that shouldn't be handed over) is optional in a demo and load-bearing in the job.

The self-report you collected afterwards doesn't close the gap either. In a randomized trial, sixteen experienced open-source developers took 19% longer to finish issues with AI tools allowed, and still believed afterwards that the tools had sped them up by roughly 20% 1. If people cannot correctly report the direction of their own effect while doing real work on their own repositories, a candidate narrating a demo they just won is not a reliable witness to their own process.

What month one adds is volume and quiet. In the 2025 Stack Overflow developer survey, the most common frustration with AI tools, at 66%, was output that is almost right but not quite, and 46% of developers said they distrust the accuracy of AI output against 33% who trust it 2. Almost-right is invisible at demo scale. It becomes expensive at forty hours a week, when nobody is watching the screen and the artifact goes to a client, a repository, or a board deck.

The third mechanism is confidence, and it runs the wrong way. Across 319 knowledge workers describing 936 first-hand examples of AI-assisted work, higher confidence in the AI predicted less critical thinking, while higher confidence in one's own ability to do the task predicted more 5. A candidate who has only ever produced with an assistant has the first kind of confidence and not the second. The demo rewards it. The job punishes it, about four weeks in.

Which four failures show up in month one?

Four, and they arrive in roughly this order. The problem is never framed before generation starts. No claim is checked against anything outside the conversation. Everything of consequence gets handed over. Nothing the assistant produced is ever refused. Each one is a behavior with a time and a place attached, which means each one has an interview question that would have surfaced it.

1. The problem was never framed. In the demo you supplied the brief. On the job the request arrives vague, and the assistant resolves the vagueness for them, confidently and in the wrong direction. The tell in month one is a finished deliverable answering a question nobody asked. Ask: "Describe a request last month that arrived unclear. What was your actual first message to the assistant, not the tidy version?" A strong answer opens with what was being decided, or what would make an answer wrong. A weak one opens with the deliverable. 2. Nothing was checked outside the conversation. Demo claims never have to be true; month-one claims reach a client or a pull request. The tell is a number, a citation or an API that nobody opened. Ask: "Which specific claim in that output did you check, and where did you check it?" A real check has a location and a result: "I opened the filing and the figure was annual, not quarterly." "I reviewed it all" is not a check. 3. Everything of consequence was delegated. The demo hour makes total delegation look like speed. Month one makes it look like a person who cannot answer a question about their own work. The tell is the new hire who goes quiet the moment the assistant is unavailable or wrong. Ask: "What did you deliberately do yourself on that task, and why?" Note what a good answer is not: volume of AI use is not the measure, and someone who judged the model was the wrong instrument for a step has demonstrated exactly the thing you want. 4. Nothing was ever refused. In a demo, accepting everything is the fastest path to a finished artifact. In a role, it means the first bad framing ships. The tell is a person who edits fluently and rejects nothing. Ask: "What did the assistant give you that you threw out, and on what grounds?" Tidying prose is not refusing; removing a section because it was doing no work is.

Ask all four of every finalist, in the same order, and write the reason for each rating rather than the rating alone. Fluency in describing AI work is the one thing interview coaching reliably produces, so the follow-up question is where understanding shows.

Why is the missing signal different for every role?

Because the check that separates careful from careless is field-specific, and so is what it costs when nobody runs it. The structure is constant: an assistant produces something fluent, and one verification only a practitioner would think to run decides whether it holds. The consequence is not constant. An unverified claim costs a consultant a client meeting and costs an engineer a production incident, so the signal you were missing in the round is a different signal per role.

  • Consulting and legal work. The failure is a confident citation that does not exist or does not say what the draft claims. This is measured, not hypothetical: three legal research tools sold on eliminating hallucination were found producing hallucinated output between 17% and 33% of the time 4. Month one delivers it to a partner or a client. The check to interview for: a source opened and read before it entered a document, and what it turned out not to support.
  • Software engineering. The failure is code that runs clean and is still wrong. In a controlled study, participants with an AI code assistant wrote significantly less secure code than those without, and were more likely to believe their code was secure; the participants who trusted the assistant least and reworked their prompts produced the fewest vulnerabilities 3. The check to interview for: a test the candidate wrote themselves, an input they tried, a dependency they actually read.
  • Analysis and finance. The failure is a model that reconciles to nothing while being described beautifully. The check to interview for: one figure recomputed by hand or against the filing, and the recommendation that moved because of it.
  • Product and marketing. The failure is a spec or a brief whose most on-message statistic is the one least likely to have been opened, and whose ambiguity the assistant quietly resolved. The check to interview for: the vagueness the candidate refused to let the model settle, and who they went to instead.

This is why one generic AI-fluency round under-serves a mixed loop. Start from what AI actually does in this role, name the artifact it produces fluently, then name the check that decides whether the artifact holds. That pair is the round.

Change the round so a check is visible

Stop asking for output in the round and start asking for the act performed on output. Two changes cover most of it. Hand the candidate a short piece of AI-assisted work with one error planted in it (your error, in your field, so no preparation reaches it) and watch what they do with fifteen unstructured minutes. Then ask about the last thirty days of their work rather than the last hour of yours.

The planted-error exercise is the cheapest instrument here because it inverts the demo's incentive. The artifact already looks finished, so producing more is worthless and checking is the only move that scores. Write down beforehand what a strong, adequate and weak response looks like, in the words of someone who has done the job; without that sheet, panels fall back on prose style and confidence, which is what the demo already measured. Testing whether a candidate catches an AI error is a different round from testing whether they can produce with one.

For the thirty-day questions, insist on one task with a name and a date. Generalities about "how I use AI" are rehearsed; a specific Tuesday is not. Follow every answer with one probe: where, exactly, and what changed as a result. The second answer is the informative one, because the first is the one candidates prepare.

Say in the invitation that the round covers working with AI and that using it is expected rather than tolerated. An unstated rule gets guessed at, and the guessing measures how well a candidate reads a company rather than how they work. It also removes the perverse case where the strongest verifier in your pipeline hides the assistant because they think that is what you want.

What can't an interview show you?

Whether the person would actually do any of it unwatched. Every answer above is a report on a check, given by someone with an obvious stake in it, and the METR result says people misread their own process with no interviewer present 1. A structured round is still worth running, and revised estimates put structured interviews at .42 operational validity against .33 for work samples 6, but validity at that level is a correlation, not a guarantee about one hire.

The second limit is duration. A demo hour and a forty-five-minute round both sample the supervised state. Month one is the unsupervised state, and the behaviors that fail there (sustained framing, checking something on the fourth hour of a boring task, refusing a direction with a deadline on it) only appear when the work is long enough to get tedious and real enough to have a cost.

So close the gap where it is cheapest to close. Extend a work sample until it is boring. Make the assistant available and make the confident answer wrong. Then read the record of what the candidate did rather than the artifact they handed in, because what a score predicts about performance depends entirely on whether the thing scored was the behavior or the output. If the round still ends in an argument about who was more impressive, nothing was measured.

See the benchmarks

Common questions

Was the candidate faking the demo?

Usually not. The demo measured something real (the ability to produce a finished artifact quickly with an assistant), and that ability exists. What it did not measure is what happens when the output is confident and wrong, which is the state the job spends most of its time in. Treat it as an incomplete measurement rather than a deception. The correction is not more suspicion in the round; it is a round that makes checking the only scoring move, since producing more output cannot earn anything once the artifact already looks finished.

How long does month one take to expose this?

Long enough for the first unsupervised deliverable to reach someone with standing to object. In practice that is the second or third week, when work stops being onboarding tasks with a known answer and starts being real requests with vague briefs. The earlier signal is quieter: a new hire who cannot answer a question about their own submitted work, or whose output arrives fast and then needs three rounds of correction. Watch the correction rate on their first four deliverables rather than the speed of the first one.

Should the demo round be scrapped?

No. Narrow what you conclude from it. A demo is decent evidence about production speed and communication and near-zero evidence about verification. Keep it, score it for what it can carry, and add one exercise where the artifact arrives finished and flawed. Two changes make the existing demo more informative at no extra time: give a deliberately vague brief instead of a clean one, and require the candidate to name one thing they refused. Both convert a production test into a judgment test without adding a stage.

What if the role really is mostly production?

Then the demo is closer to the job and this failure pattern is rarer, but check the assumption before relying on it. Roles described as pure production usually contain one decision per week where a wrong, confident answer is expensive, and that decision is where the month-one failure lands. Name it explicitly: what does this job get wrong when a competent person is rushing? If the honest answer is nothing costly, hire on the demo and spend your assessment budget on a role where the answer is not nothing.

How do you tell interview coaching from real competence?

Ask for coordinates. Coaching produces fluent general accounts of process; competence produces a specific task, a specific claim, a place it was checked, and a result that changed something. Follow each answer with one probe: where exactly, and what moved. Prepared answers thin out at the second question, while a real check has detail the candidate did not expect to need. It also helps to ask what they threw out, because rejection stories are much harder to invent than production stories.

Does an assessment solve the demo-to-delivery gap?

It narrows it rather than closing it. A longer, role-grounded task with an assistant available puts the candidate in the state where verification is optional and boring, which is the state the demo skipped. Olive works that way: a 40-to-60-minute occupational assignment, and a human reviewer who writes six findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each carrying the moment it rests on. There is no number standing for a person, and the candidate is granted the same report the employer reads. It is an input to your decision, not the decision.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial with 16 experienced open-source developers: 19% longer to complete issues with AI tools allowed, while developers believed afterwards that AI had sped them up by about 20%.
  2. 2. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  3. 3. Do Users Write More Insecure Code with AI Assistants? arXiv (Stanford University; published at ACM CCS 2023), 2023. arxiv.org Participants with an AI code assistant wrote significantly less secure code and were more likely to believe it was secure; those who trusted the assistant less and reworked their prompts produced fewer vulnerabilities.
  4. 4. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools arXiv (Stanford RegLab and Institute for Human-Centered AI), 2024. arxiv.org Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI each hallucinated between 17% and 33% of the time, despite vendor claims of eliminating hallucination.
  5. 5. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ACM CHI Conference on Human Factors in Computing Systems (CHI '25), 2025. advait.org Survey of 319 knowledge workers and 936 first-hand examples: higher confidence in GenAI is associated with less critical thinking, while higher confidence in one's own task ability is associated with more.
  6. 6. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. doi.org Revised operational validity estimates: structured interviews .42, job knowledge tests .40, work samples .33, general mental ability .31.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.