Pipeline

How Do You Test AI Skills Without Adding an Hour to the Loop?

Testing AI skills doesn't require a longer interview loop: swap a stage instead of adding one. Find the round whose output a model now writes in seconds (the from-scratch coding round, the SQL screen, the writing sample) and spend its minutes watching a candidate decide what a finished artifact got wrong. Two limits hold. Never cut the structured interview to make room, and where nothing is redundant, merge two overlapping rounds rather than call the hour free. Which stage goes differs by role, so audit your own loop first.

The takeLoops don't grow because anyone believes five rounds beat four. They grow because adding a stage needs no defense and deleting one needs a confession: that a round you have run for three years has rejected nobody. My read is that most teams have never counted how many people each stage uniquely rejected, because the number would end a stage somebody owns. The hour you say the loop cannot spare is already inside it, being spent on a round that has never changed a decision.

Where Olive fits

Open a role and see what the work shows

Olive is priced per attempt rather than per seat, so the replacement stage can run beside the round you are thinking of cutting before you cut it, and ten attempts a month cost nothing. An attempt returns six evidenced findings on one candidate, written by a human reviewer, as an input to your decision rather than a verdict on the person.

Rank your shortlist

Which stage should the AI exercise replace?

The stage whose whole job was watching someone produce an artifact a model now produces in seconds. In most loops that is one specific round (the from-scratch coding exercise, the SQL screen, the writing sample, the spec take-home), and since the tools arrived it has been measuring the part of the work nobody does alone any more. The minutes it holds are the minutes you were told the loop could not spare.

Run three questions at every stage you have. What decision does this stage change that no earlier stage already changed? When did it last reject someone who had passed everything before it? Could a competent person with a model pass it without understanding what they submitted? A stage that fails all three is not a stage, it is a habit with a calendar invite.

The validity evidence makes the ranking more uncomfortable than most loops assume. In the revised estimates that replaced decades of overcorrected meta-analysis, structured interviews lead the widely used predictors at a mean operational validity of .42, with an 80% credibility interval running .18 to .66, and unstructured interviews sit at .29 1. Most loops carry two or three rounds of the second kind (the culture conversation, the walk-through-the-resume call, the second manager chat), and those are the cheapest hour in the building to give up. Before cutting anything, be clear about what a brilliant round was predicting in the first place, because the honest answer narrows what you are giving up.

How do you find the redundant hour in your own loop?

Pull the last twenty candidates for one role and write four columns: the stage, minutes of interviewer time, minutes of candidate time, and how many people that stage rejected who would otherwise have passed. The fourth column comes back zero for at least one stage in almost every loop that has been running a year. That stage is your hour, and the audit takes an afternoon.

Read the table three ways. A stage with no unique rejections and real cost is a cut. A stage whose rejections duplicate an earlier stage's is a merge, not a cut. Fold its two good questions into the round before it. And the stage everyone dreads scheduling is where your calendar days go, which matters because an added round costs more than its own length: candidate time, two interviewers' time, and the days between the invitation and the slot.

Count the whole cost when you compare. A 45-minute exercise a candidate does on their own clock spends 45 candidate minutes and no interviewer minutes until someone reviews it; a 45-minute panel spends 45 candidate minutes and 90 interviewer minutes, plus scheduling. Swapping a live round for an asynchronous one can hold the candidate's total flat while giving your team back its afternoon, which is often the constraint people actually mean when they say the loop cannot get longer.

Run the new stage in shadow before it decides anything. Send it to candidates already in the loop, review the results, and compare them against what the stage you plan to cut concluded about the same people. If the two agree, you have a swap. If they disagree, you have something worth arguing about before it gates anyone. The case for piloting an assessment before it becomes a gate is strongest exactly here, where a swap is reversible and a bad gate is not.

Which stage goes, by role?

The stage that existed to make a person produce a document, and which document that is differs by function. For a software engineer it is the from-scratch algorithm round; for a data analyst the SQL screen; for a salesperson the written prospecting sequence rather than the live call; for a marketer the writing sample; for a product manager the spec take-home. Each of those was a reasonable proxy for judgment while producing the artifact was the hard part.

  • Software engineering: cut the watched from-scratch round. In a randomized trial with 48 computer science students, 61.5% failed a whiteboard task while being observed against 36.3% in private, and median correctness fell from three passing test cases to one 2. It was a noisy measurement before models could write the function. Replace it with a change that arrives already written and carries one defect the tests do not catch, because the dominant failure of AI output is not obvious error: 66% of developers named almost-right-but-not-quite their top frustration in the 2025 Stack Overflow survey, and 46% said they distrust the accuracy of AI output against 33% who trust it 3.
  • Data and analytics: cut the SQL screen. Keep a result with a fluent, confident, wrong interpretation attached and a dataset that will not correct it.
  • Sales: cut the written prospecting exercise, keep the live call. The email sequence is the artifact a model writes best. Hand them a model-written account plan built on a stale funding figure instead, and watch whether anyone opens the source.
  • Marketing: cut the writing sample. Keep a positioning brief whose most quotable statistic is the one least likely to be checked.
  • Product management: cut the spec take-home. Keep a vague request an assistant will resolve confidently in the wrong direction.
  • Legal operations, revenue cycle, underwriting: cut the drafting exercise. Keep a packet whose documents disagree, and see which one the candidate treats as settled.

The rule underneath all six: cut where the tool now does the producing, keep where the tool now does the misleading. That requires knowing what AI actually does inside the role rather than what it does in general, which is a half-hour conversation with two people who hold the job. For engineering specifically, the trade between a watched session and work done privately has its own evidence and its own failure modes, and the live-versus-take-home comparison is worth reading before you decide which of the two rounds survives.

What has to be in the replacement round for the swap to pay?

Three things: an artifact that arrives finished and wrong, an assistant the candidate is allowed to use, and a rubric written before anyone sits down. Producing more output has to earn nothing inside the exercise, or you have rebuilt the round you just cut with better tooling. Thirty to forty-five minutes is enough when the planted error is one only a practitioner catches.

Make the material yours. Take a real deliverable from the role, have a model produce it, then plant a defect a competent person would find by opening something: a figure on the wrong basis, a source that does not say what the paragraph claims, a dependency whose name promises what it does not do. No amount of interview coaching reaches a specific error in your own material, which is the property a generic AI-fluency quiz does not have.

Ask for performance, not description. Work sample tests require applicants to perform tasks that mirror the tasks employees perform on the job 4, and the distinction is the whole reason the swap works: a conversation about how someone uses AI is a self-report, and self-reports about AI work are unreliable in a way that is now measured. In a randomized controlled trial, sixteen experienced open-source developers took 19% longer to complete issues when AI tools were allowed, and still believed afterwards that the tools had made them about 20% faster 5. People describing their own process are not lying; they cannot see it.

Write the rubric in the words of someone who does the job, and score acts rather than impressions: what was framed before anything was generated, what source was actually opened, what work was kept back from the assistant, what output was refused and on what grounds. Record the reason beside each rating, because the reason is the only part a second reviewer can audit. If you are weighing this against simply adding an AI round to the existing loop, the rubric is what makes the swap defensible and the addition merely longer.

What does the swap actually cost?

Development time you were not spending before. The federal guidance on work samples lists the costs plainly: development that can be costly in time and money, periodic updating, administration that is time consuming and expensive, and someone to observe and sometimes rate performance 4. The candidate's total stays flat. Your side gets more expensive, in the one place a swap cannot hide it.

The new stage is also a selection procedure, with everything that follows. Any step used to decide who advances has to be job-related and consistent with business necessity, applied the same way to every candidate, and reviewed for adverse impact 6. That obligation does not arrive because the exercise involves AI; it arrives because the exercise decides something. Write down what it measures and why that matters for this job before the first candidate sees it, and keep the same case for everyone in the role, because a swapped-in exercise that varies by interviewer is worse evidence than the round you cut.

Moving an hour from one stage to another moves signal; it does not create more of it. A loop that was already too short to decide will still be too short, and a 45-minute exercise samples less than a three-hour one, which is a real trade rather than a rounding error. What the swap does buy is a different kind of evidence for the same money: an act you can point to instead of an account you have to believe.

The last cost is discipline. The cut stage will try to come back, usually as a 30-minute chat someone adds to be thorough, and six months later the loop is five stages again with the exercise on top. Put the deletion in writing with the reason attached, and re-run the four-column audit each time someone proposes a new round. Grading also gets harder before it gets easier, because a polished take-home is difficult to read for judgment until the rubric names the acts you are looking for.

See a sample report

Common questions

Can the AI exercise just be added as a fifth round?

You can, though the cost lands on completion and calendar days rather than on your schedule, which is why it feels free from the inside. Every added stage is another scheduling round trip, another chance for a candidate holding two offers to withdraw, and another week of time-to-fill. If you add rather than swap, add it early and asynchronously, drop a live round of similar length in the same change, and measure drop-off at the new stage for a full cycle before deciding it was worth the hour.

How long should the replacement exercise be?

As long as the stage it replaces, which is the whole constraint. Thirty to forty-five minutes is enough when the artifact arrives finished and the planted defect is specific, because the candidate is checking rather than building. Longer exercises sample more but cost completion, and a three-hour task is a different decision that needs its own justification. Fix the length first from the stage you are cutting, then design the exercise to fit it, rather than designing the exercise and discovering it needs ninety minutes.

What if nothing in the loop is safe to cut?

Then the loop is genuinely at minimum, and the honest options are to accept the added time or to merge rather than cut. Merging usually works: two 45-minute rounds with overlapping questions become one 60-minute round with the best questions from both, which frees 30 minutes without removing any signal. If neither is possible, say plainly that testing AI skills costs an hour this quarter, and pick which role is worth spending it on rather than pretending the change is free.

Does this replace the structured interview?

No. Structured interviews remain the top-ranked widely used predictor in the revised validity estimates, so cutting the structured round to make room is the one swap that reliably loses you signal. Cut the unstructured conversation instead, or the artifact-production stage a model has made trivial. The exercise reaches a behavior an interview cannot (whether someone checks a confident claim rather than describing how they would), and the two are complements, not substitutes.

Who should review the replacement exercise?

Someone who does the job, reading against a rubric written before the round. That is the real cost of the swap and the part that decides whether it works: a reviewer who cannot tell a well-aimed refusal from a tidy edit will score polish, which is what the stage you cut already measured. Budget 15 to 20 minutes of practitioner time per candidate, calibrate two reviewers against the same three sessions before anyone is graded for real, and keep the reasons alongside the ratings.

Can Olive be used as the replacement stage?

For a take-home-sized slot, yes. Olive's assignment runs 40 to 60 minutes of candidate time on an occupational task done with an AI assistant, so it swaps cleanly for an unwatched take-home and less cleanly for a 45-minute live round, which is worth saying before you plan around it. A human reviewer writes six findings, each carrying the moment in the session it rests on, and the candidate is granted the same report the employer reads. The report carries no hiring recommendation.

References

  1. 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. doi.org Structured interviews top the revised list of widely used predictors at a mean operational validity of .42, with an 80% credibility interval running .18 to .66; unstructured interviews are estimated at .29.
  2. 2. Does Stress Impact Technical Interview Performance? Behroozi, Shirolkar, Barik and Parnin, ESEC/FSE 2020 (NC State author copy), 2020. chrisparnin.me Randomized controlled trial with 48 computer science students: 61.5% failed the whiteboard task in the watched setting against 36.3% in private, and median correctness fell from three passing test cases to one.
  3. 3. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  4. 4. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2026. opm.gov Work sample tests require applicants to perform tasks that mirror the tasks employees perform on the job; development costs may be costly in time and money and may require periodic updating, and administration may be time consuming and expensive and requires individuals to observe and sometimes rate performance. Undated guidance, page text verified 2026-08-24.
  5. 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org Randomized controlled trial with 16 experienced open-source developers: 19% longer to complete issues with AI tools allowed, while the developers believed afterwards that AI had sped them up by about 20%.
  6. 6. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Any step used to decide who advances is a selection procedure: it must be job-related and consistent with business necessity, applied consistently, and reviewed for adverse impact. Guidance dated December 1, 2007; page text verified 2026-08-24.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.