Assessment design

The Algorithm Screen Now Measures Prep Time, Not Programming

The algorithm-style coding screen is still worth running as a cheap floor check on whether a candidate can write working code at all, and no longer worth running as a ranking. A model clears the standard format in seconds, so what the screen separates now is how much unpaid preparation time someone had, which tracks who has spare evenings more than who can engineer. Move the ranking to multi-file work in a repository you maintain, where the interesting decisions are what to leave alone.

The takeThe dead-or-not-dead argument treats the model as the thing that broke the screen, and the model is not what broke it. A puzzle with one known answer and an automated test suite is exactly the shape a model is best at, and it was already the shape that rewarded practice volume over engineering. What changed is that the practice advantage is now free to anyone who thinks to ask for it, which removes the last defense the format had.

Where Olive fits

Open a role and see what the work shows

Build this yourself and the cost lands on the answer key and the evidence trail. Olive ships a role-grounded assignment with an AI assistant that will do all of it if nobody stops it, and returns six findings written by a human reviewer, each carrying the timestamped excerpt behind it.

Rank your shortlist

Keep it as a floor, drop it as a ranking

Keep it if it is short, automated and cheap, and use it for one thing: confirming a candidate can write working code at all. Stop using the result to order a shortlist. A floor check carries one bit of output, passed or not, and one bit is all this format can still carry honestly. Everything a ranking would rest on is now available to anyone who asks a model for it.

Each version has a price, and the two are not close. A floor at one easy problem in twenty minutes costs a candidate a lunch break and costs the team nothing to grade. A ranking round costs ninety minutes of somebody's evening, produces a number that gets argued about in the debrief, and separates people on something nobody intended to measure.

Say which one it is in the invitation. Naming it as a check rather than a competition changes what candidates spend on it, and that is the point: nobody should be doing four weeks of preparation for a gate whose entire output is yes. It also removes the argument about whether assistants are allowed, since on a floor check the answer can simply be that they are.

The general form of this decision, ranking candidates against holding them to a bar, applies at every stage of a loop and is cheaper to settle once than to relitigate per stage.

What does the screen separate now?

Preparation time, mostly. The format rewards recognising a problem type quickly, and recognition comes from volume of practice. Practice volume tracks how many uncommitted evenings a person has had, which tracks caregiving, second jobs and how recently they were a student. That was always partly true. What changed is that the recall advantage is now free to anyone who thinks to ask an assistant for it.

The evidence that the format is a toy sits inside the most-quoted study in favour of AI assistance. The famous 55.8% speedup with Copilot comes from a single trial in which 95 freelance developers recruited on Upwork were asked to write an HTTP server in JavaScript. It was measured among those who finished, with a 95% confidence interval running from 21% to 89% 1. One synthetic task, a known answer, an automated test suite, no legacy code and no review. That is the shape of an algorithm screen, and it is the setting the case for assistance is built on.

Now put the same tooling in the setting the screen claims to predict. In a randomised trial where 16 experienced developers completed 246 real tasks on mature repositories they had worked on for about five years, early-2025 AI tools made them 19% slower, after they forecast a 24% speedup and still believed in a 20% one afterwards 2. Sixteen developers in one setting, so it does not transfer to juniors or greenfield work.

Read the two together and the gap is the whole argument. The toy task and the real repository are different problems, assistance moves them in opposite directions, and only one of them is what the job consists of.

How do you build the replacement?

Start from code you already own. Take a real change from your own repository, seed it with generated code that is plausible and subtly wrong, and ask the candidate to get it ready to merge using any assistant they like. Score two things: what they verified before trusting it, and what they accepted without checking. Both are visible in a transcript, and neither is guessable in advance.

Choose the fault carefully, because the fault is the instrument. It should be the kind a model produces confidently: an off-by-one at a boundary, a plausible but non-existent library method, an error path that swallows the error, a query that works on sample data and not on production shapes. Avoid anything a linter catches, and avoid puzzles. The candidate should have to go outside the conversation to find out, which is the behaviour worth hiring.

This is the property the old format lacks. The exercise gets harder when the assistant is open rather than easier, because a model will defend a wrong version fluently, and a candidate who reads fluency as evidence fails while producing more output than someone who stops and checks. That failure has been measured once, on a task deliberately chosen to sit outside AI capability, where consultants using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group against 60% and 70% in the two AI conditions 3. One task, one sitting and a 2023 model, so read it as a demonstration that people could not tell which side of the line they were on, not as a rate you can expect.

Two practicalities. Cap it at forty-five minutes and say so, since a multi-file exercise expands to fill whatever it is given. And read the transcript before the diff, which is most of what running a coding interview with an assistant open asks of an interviewer. How much of the working you get to see at all depends on live coding against a take-home.

Should you check the submission for AI authorship?

No, because nothing supports it, including the tools built for code. An ICSE 2025 study evaluated five AI-content detectors plus the code-specific GPTSniffer on code generated by four models across three benchmarks, and found accuracy mostly below 0.6, which the authors describe as ineffective 4. Their own purpose-built model reached an F1 of 82.55 on their own data, which is not a product anyone can buy and nowhere near a bar that could support declining a candidate.

Those numbers come from benchmark functions, and no repository or real submission was tested. Benchmark functions are the easy setting, so a real take-home is the harder case these tools were never measured on. Nothing there supports checking a take-home for authorship, and the two errors are not symmetrical: a missed assistant costs an extra interview, while a wrong accusation costs a person a job and the team a defensible process.

The design answer is to stop needing the question. An exercise that assumes the assistant, grades the working, and puts the fault where a model reliably falls into it produces evidence regardless of what wrote the first draft. Authorship stops mattering once the artifact stops being the thing that gets graded.

If the thing to replace right now is a bought online assessment, that swap has its own answer. And if the goal is narrower, testing whether a candidate notices when a model gets something wrong is the exercise to copy.

See what gets scored

Common questions

Is it fair to keep an algorithm screen at all?

As a floor with a low bar and a short clock, yes, and it is defensible for the same reason a typing test is: it checks one narrow, job-relevant thing and nothing else. It stops being fair when the result is used to order candidates, because the ordering reflects preparation hours rather than engineering judgment, and preparation hours are distributed by circumstance. Keep the bar low enough that anyone who writes code for a living clears it, and the fairness question mostly goes away.

What about proctoring the screen instead of replacing it?

Proctoring defends the format rather than fixing it. Even a perfectly enforced no-assistant rule leaves a screen that sorts on memorised patterns, which is the underlying complaint, and it adds real costs: candidate objection, the accommodation questions a timed and monitored format raises for a United States employer, which is ADA ground and a question for counsel, and a class of false accusations that are expensive to be wrong about. If the format is worth keeping only when nobody can use a tool the job uses daily, that is information about the format.

Is a technical screen still needed before the onsite?

Something cheap in front of a panel day is usually worth it, but it does not have to be an algorithm puzzle. A twenty-minute conversation about a change the candidate shipped, with two follow-ups about what they checked before merging, filters as well and costs a fraction of the candidate's time. The purpose of the stage is to avoid spending five people's afternoon, and almost any structured, scored, job-relevant filter serves that purpose.

How do you grade an exercise where everyone reaches the right answer?

Grade the route rather than the destination. When the assistant is open, most candidates arrive at working code, so the separation lives in what they checked, what they rejected, what they decided not to touch, and what they could explain afterwards without the transcript in front of them. Write the rubric against those behaviours before running the exercise, and grade from the recording. A rubric written afterwards fits whoever was graded first.

Does this apply to non-engineering roles?

The mechanism does. Any assessment built around a task with one known answer, delivered unsupervised, now measures something different from what it used to, whether the artifact is code, a spreadsheet model, a memo or a case response. The replacement pattern is the same in every case: supply something plausible and wrong, allow the assistant, and read what the candidate verified. Only the content of the fault changes with the occupation.

References

  1. 1. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (arXiv:2302.06590) arXiv (Microsoft / GitHub researchers), 2023. arxiv.org Supports the claim that the headline speedup figure was measured on a single synthetic task with a known answer and an automated test suite.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the contrast between toy-task assistance and measured performance on mature repositories, including the gap between belief and measurement.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that people accept a confident wrong answer on a task that resembles the ones AI handles well, which is what the replacement exercise observes.
  4. 4. An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We? arXiv (published at ICSE 2025), 2025. arxiv.org Supports the claim that measured detectors, including the code-specific one, cannot establish AI authorship of a code submission.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.