Assessment design
What Survives Is Knowing Whether the Output Is Right
When AI can produce the output, the skill that still holds value is telling when a plausible-looking result is wrong, and knowing which parts of it need checking in the first place. That is not the vague "creativity and critical thinking" list career advice has repeated for a decade unchanged. It is a specific, testable behavior, and it appreciates precisely because generation got cheap. Producing a plausible draft is now fast for almost anyone. Deciding whether it is right is still done one claim at a time, by a person.
Where Olive fits
Open a role and see what the work shows
If the assessment an employer sends you is Olive, the assignment is done openly with an AI assistant, and a person writes six findings describing how you framed the problem, what you delegated and what you verified, in words rather than a number. You are granted the identical report, free.
Rank your shortlistAsk What Actually Got Cheap
Generating plausible output got dramatically cheaper. Checking whether that output is actually right did not, and that gap is the shift a decade of durable-skills advice keeps circling without naming. The two most-quoted studies about AI at work are worth reading closely for what they measured, because the half they left out is the half a reader can still build a career on.
In a pre-registered online experiment, 444 college-educated professionals did occupation-specific writing tasks in a single sitting; the half given ChatGPT finished 37% faster and scored 0.45 standard deviations higher from their graders, with the largest gains going to the weakest writers 1. The most-quoted developer figure comes from a narrower setting: 95 developers recruited on Upwork were asked to write an HTTP server in JavaScript, and among those who finished, the GitHub Copilot group was 55.8% faster than the control group, a result whose confidence interval runs from 21% to 89% 2. Both describe the same thing from different angles. Producing a passable first draft, of writing or of code, stopped being where the scarcity is.
Speed is the easier half to measure, which is why both studies measure it well and neither settles the question this article is about. The Copilot task came with a known answer and an automated test suite, so correctness was established by the test suite rather than by the developer, and real work rarely arrives with one attached. The writing tasks were graded by other professionals on the finished piece, done once, online, for pay, with no revision cycle and no consequences. Neither setting is the one where a confident wrong claim survives into something that matters, and that setting is what the surviving skill is for.
Don't Trust the Standard Durable-Skills List
"Creativity, emotional intelligence, critical thinking" has survived a decade of automation panics unchanged, and that durability is itself a warning sign rather than a reassurance. A list nobody can fail is not a list anyone can practice against either. Ask what a person would actually do differently on a Tuesday to build "critical thinking," and the list has nothing concrete to offer.
The replacement is narrower on purpose: the ability to check a plausible output against something outside it, and to know in advance which parts of a task most need that check. It sounds smaller than "creativity." It is also something you can point to a specific instance of, describe out loud in an interview, and get better at through deliberate repetition.
The test for whether a skill belongs on a usable list is simple: can a stranger watch you do it and agree afterward that you did it well, or badly. "Critical thinking" fails that test. "Caught that the cited statistic didn't match the source it pointed to" passes it. That is the difference between an adjective and a behavior.
What Checking Actually Looks Like
The failure mode this skill guards against is specific: a confident wrong answer on a task that looks exactly like the ones AI handles well. In a pre-registered field experiment, 758 consultants at one firm used GPT-4 on 18 tasks chosen to sit inside the model's capability, and finished 25.1% faster with more than 40% higher quality than a control group 3.
On one task chosen deliberately to sit outside that capability, the same 2023 model left people 19 percentage points less likely to reach the correct answer, averaged across the two AI conditions 4. Nothing in the output's tone marked the difference. The sentence structure, the certainty and the polish stayed identical whether the underlying claim was solid or not, and the people using it could not tell which side of the line the task was on.
Noticing which side of that line a task sits on, before trusting the output, is the actual skill. It is not about distrusting AI generally or refusing to use it. It is about knowing, task by task, where the tool's confidence stops being a reliable signal of correctness, and building the habit of checking specifically in those places rather than everywhere or nowhere.
Where Speed and Formatting Lost Their Value
Speed at producing a first draft lost most of its distinguishing power, because that is precisely the task AI now does fastest for almost anyone, regardless of prior skill level. Naming what appreciated only means something if it is honest about what depreciated too.
Self-reported speed gains are unreliable on top of that. In a randomized trial, 16 experienced open-source developers working on mature repositories they had known for years finished 19% slower with AI tools, after forecasting a 24% speedup beforehand and still estimating a 20% speedup afterwards 5. Sixteen developers in one setting is not a general finding about AI and speed, and it does not transfer to unfamiliar code or to work you have not done before. What travels is the gap between belief and measurement: if experienced practitioners misjudged their own pace before the work and again after it, a speed claim about yourself is not evidence of much.
Formatting fluency, knowing the shape a polished memo or a clean function is supposed to take, lost value for the same reason: it is now the default output of the tool itself. Neither loss is a tragedy. It just means a resume built entirely around speed and polish is competing on ground that stopped being scarce, while checking is competing on ground that just got scarcer.
Test Yourself the Way Employers Now Test Candidates
The nearest thing to a practice instrument is the test employers are already being handed. Some employers have started designing assessments specifically to test whether a candidate notices when AI gets something wrong, planting a real error in the source material rather than in the model's output, then grading whether it was named, when, what it was checked against, and what changed in the deliverable as a result.
Knowing the shape of that test is not a way around it. The error is chosen so a competent practitioner in the field would question it and a fluent generalist would not, so the only preparation that helps is the domain knowledge and the habit of going outside the conversation to settle a claim.
A rough version works alone: take a task from your own field, generate an answer with an assistant, and find one thing wrong with it before accepting it. Do this on a task where you already know the right answer well enough to judge the check itself. Showing that you catch a mistake, specifically and out loud, carries further in an interview than claiming good judgment in the abstract, because the first is evidence and the second is an assertion nobody can verify.
Common questions
Isn't "critical thinking" basically the same thing as this checking skill?
They point at similar territory, but critical thinking is too broad to practice directly. Checking a specific output against something outside it, a source, a test case, a colleague who knows the domain, is a concrete act you can name, repeat, and get feedback on. That is what makes it practicable rather than aspirational.
Does this mean I shouldn't bother getting fast with AI tools?
No. Speed is still useful; it just stopped being the differentiator, because most people can now get fast with the same tools. The skill worth building on top of speed is knowing when to slow down and check.
How do I demonstrate this skill if I have no work experience yet?
Build a specific example rather than a claim: a moment you generated something with AI, checked it against a real source, and found it was wrong. That concrete story is something an interviewer can question in detail, which is what makes it read as evidence.
Is this the same as being good at prompting?
Different skill, often mistaken for the same one. Prompting well affects what you get out of a generation step; checking well affects whether you trust what you got. A well-crafted prompt can still produce a confident, wrong answer that needs the same scrutiny as a lazy one.
What if my field doesn't have an obvious 'right answer' to check against?
Most fields have something to check against even without a single correct answer: a stated requirement, a stakeholder's actual constraint, a prior version of the work, a colleague with domain knowledge. The skill is picking the right thing to check against, not always having a clean answer key.
References
- 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) economics.mit.edu Supports that AI writing assistance produces large, measured speed and quality gains on short professional writing tasks.
- 2. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot (arXiv:2302.06590) arxiv.org Supports the famous 55.8% speed figure as evidence that producing a first draft got dramatically cheaper on a synthetic task.
- 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the speed and quality gain on tasks inside an assistant's strength, the half of the frontier where confidence is a reliable signal.
- 4. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) mitsloan.mit.edu Supports the accuracy drop on a task outside an assistant's strength, the failure mode the checking skill exists to catch.
- 5. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) arxiv.org Supports that self-reported AI speedups are unreliable even among experienced practitioners, backing the claim that speed alone stopped being a trustworthy signal.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.