Assessment design
What Replaces an Online Assessment ChatGPT Solves in Thirty Seconds?
Replace an online assessment ChatGPT solves in thirty seconds with a task where the model's answer is only the start. Hand campus candidates occupational material carrying a defect (a dataset that will not support the conclusion, a brief missing a constraint, a figure its source does not carry), let them use AI openly, and grade three acts: how the problem was framed before generating, what evidence was demanded, what got checked outside the chat. Keep it to judgment a first-year already brings. Proctoring restores the old test, not the signal.
The takeThe solved assessment has been this thin for years, and this is the year it shows. A task a model finishes in thirty seconds was already measuring something no first-year job asks for, and campus programs kept it because it cost nothing per candidate, not because it predicted anything. That is the bill arriving. Judgment cannot be machine-scored, so somebody reads. Programs unwilling to pay for the reading will buy proctoring instead and call it rigor, and I suspect they end up spending more defending that choice than the reading would have cost them.
Where Olive fits
Open a role and see what the work shows
If you build this yourself, the expensive parts are the answer key and the evidence trail: somebody has to decide what counts as a check and then find it in the submission. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings, each anchored to a timestamped moment in the session rather than to a number.
Rank your shortlistWhat does the replacement task look like?
Occupational material with something wrong in it, plus an assistant the candidate is welcome to use. The model answers confidently in seconds, that answer is wrong or incomplete, and everything worth measuring happens after it arrives. Three acts carry the signal: the problem framed before anything is generated, evidence demanded for the claim the answer rests on, and one check run against something outside the chat.
The defect is the design. A task with a single correct answer sitting in the model's training data measures retrieval, and retrieval got free. A task whose inputs do not support the obvious conclusion measures whether anyone looked. Developers already live in that gap: in Stack Overflow's 2025 survey, 66% named "almost right, but not quite" as their top frustration with AI tools, and 45.2% said debugging AI-generated code costs them more time than writing it 3. The failure mode is not gibberish. It is plausible.
It is worst where the tooling looks most trustworthy. Purpose-built legal research systems, sold to lawyers on the promise of grounded answers, returned incorrect information in 17% to 33% of benchmark queries in a 2024 Stanford evaluation 4. Nothing on the screen separated the usable memo from the sanctionable one. Someone opening the cited case did.
So build the task around that gap instead of around the answer. Hiring for verification rather than production is the same argument applied to the job description, and the three observable moves are the same three whether you are watching them in an interview or reading them off a submission.
Don't reach for proctoring first
Proctoring restores the conditions of a test whose content stopped discriminating. Lock the browser, and strong candidates still spend an hour on a task the job would let them do with an assistant. You have measured tolerance for being watched, not judgment. Detection is worse for a campus funnel, because the tools misfire hardest on writing by people who learned English second 5.
The numbers are not marginal. Seven detectors run over 91 human-written TOEFL essays produced a 61.3% average false-positive rate, and 97.8% of those essays were flagged by at least one tool, while the same detectors were near-perfect on US eighth-grade essays 5. A campus pipeline is exactly where that falls: students on temporary visas earned 36% of US science and engineering master's degrees and 59% of computer science doctorates in 2019 6. A rule that misfires on one group and not another, applied at campus volume, is not a fairness problem you can average away.
There is a plainer objection too. If the job allows AI (and for most roles a campus program feeds, it does), then a banned-AI assessment tests conditions the hire will never work under. Whatever you run instead, the EEOC's ask of any step that decides who advances is the same: apply it identically to every candidate, and be able to say what it measured 7.
Open the assistant, say so in the invitation, and put the difficulty somewhere the assistant cannot reach. The proctoring-versus-AI-open decision turns on one question: is the hard part producing the answer, or noticing the answer is wrong?
How do you build one in an afternoon?
Take a real artifact your team produced last quarter, break one thing in it on purpose, and write the answer key before you write the prompt. The key lists moves rather than a solution: what a candidate has to notice, what they would have to open to notice it, and what an honest submission says about the limit it ran into. Ninety minutes of your time, once per role.
Four decisions do most of the work:
1. Pick the defect from the occupation. A dataset whose sample cannot support the conclusion. A brief missing the one constraint that changes the recommendation. A quoted figure that does not appear in the source cited beside it. A model will sail past all three, fluently. 2. Ask for the intermediate artifact, not only the deliverable. A plan, a criteria list, an outline, made before the deliverable and visibly shaping it. It is the cheapest thing to require and the hardest thing to reconstruct afterwards. 3. Require the transcript. Whatever assistant the candidate uses, ask for the exchange pasted in beside the work. You are not policing it; you are reading it. 4. Say what happens to it. One paragraph in the invitation: how long the task takes, who reads it, what gets graded, and that AI is allowed.
Keep it recognizable as the job. Content validity, under the guidelines an agency would apply, rests on the selection procedure being a representative sample of the behavior of the job, not on the task being hard 8. A puzzle nobody at your company has solved on an ordinary Tuesday is not a work sample, however cleanly it resists ChatGPT.
| Occupation | The defect you seed | What a real check leaves behind |
|---|---|---|
| Data and analytics | A dataset that will not support the requested conclusion | A caveat added, or the question re-scoped |
| Software engineering | A spec that contradicts the tests already in the repo | A failing test named, a requirement queried |
| Finance | A figure in the packet the filing does not support | A number changed, or a claim withdrawn |
| Marketing | An on-message statistic whose source says something narrower | The statistic cut or re-qualified |
| Operations | Three quotes written on three different terms | A comparison rebuilt on one basis |
If you already run a take-home, this is a rewrite rather than a new step, and redesigning the work sample is usually cheaper than defending the old one to a hiring manager who watched a candidate finish it in a minute.
What do you grade when every submission is polished?
Grade the working, not the prose. Fluency has stopped varying between candidates, so the deliverable carries almost no information and the record of how it was made carries nearly all of it. Four observable acts are enough: a constraint named before anything was generated, a source opened for the claim the answer rests on, a piece of model output refused with a reason, and a check run against something outside the conversation.
Write each one as a yes-with-evidence rather than a rating. "Refused something" is a claim; "rejected the second recommendation because the segment size came from a sample of eleven" is evidence, and two graders will agree on it. Count refusals as a rate against the answers actually taken up, so three well-aimed prompts are never beaten by thirty.
Two cautions. Volume of AI use is not a virtue: a candidate who decided the model was the wrong instrument for a step and did it by hand has demonstrated the thing being tested. And the job-specific version beats the generic one: the 2023 re-estimation of selection-method validity put job-knowledge tests at .40 and work samples at .33, and found job-specific measures outperforming general ones 1. A rubric written against your own material is worth more than a bought one written against nobody's.
Every submission will look good. Grading polished take-homes is the same problem one step downstream, and the answer is the same: stop reading the artifact as evidence about the person. See how Olive measures this.
How does this survive campus volume?
Run one task per role rather than one per candidate, cap it at an hour, and put it beside your current assessment for a cycle before it gates anybody. Work samples carry high content and criterion-related validity and candidates react well to them, but they fit only where the competency is expected on entry rather than trained after the hire 2. For campus, that means testing judgment on the material, not craft a first-year analyst has not learned yet.
The arithmetic is the real constraint. A machine-graded assessment costs nothing per candidate, which is why it lasted this long; a read work sample costs someone ten to fifteen minutes. Two ways out, both unglamorous: move it later in the funnel so fewer people take it, or split grading across the hiring team after one calibration session on five submissions everyone reads.
Then hold it steady. Same task, same rubric, same time allowance, same accommodation offer for every candidate on that role. A step that decides who advances is a selection procedure whether or not a machine scores it, and consistency is what makes it explainable a year later 7. Rotating the task per candidate to stop leaks costs you the comparison, which was the whole point of running it.
Pilot it before you gate on it. Running the new assessment beside your existing round gives you the completion rate and the grading spread before anyone is rejected on it, and keeping the loop from getting longer is the constraint campus programs actually fail on.
Common questions
Can't candidates just paste the whole task into ChatGPT?
They can, and that is the design. The model returns a confident answer in seconds; if the inputs are broken, that answer is wrong. What you read afterwards is what happened next: whether anyone opened the dataset, noticed the missing constraint, or checked the quoted figure against its source. A candidate who pastes, copies back and submits produces a clean, wrong deliverable with no working behind it. That is a result, not a hole in the test.
How long should the replacement task take?
Under an hour, on the candidate's own clock. An hour holds the shape that matters (read the material, frame the problem, generate something, check it, say what you found), and it is short enough that a student running six processes at once will finish. Past that you have written a project rather than an assessment, and you are selecting for free time. Give a hard cap in the invitation and honor it.
Does this work for non-technical campus roles?
The shape is identical and only the defect changes. A marketing task hides a statistic whose source says something narrower than the headline. An operations task gives three supplier quotes written on three different terms. A recruiting task supplies a job description with a requirement nobody can justify. In every case an assistant produces a fluent answer built on the flaw, and the assessment is whether the candidate opened the thing that would have caught it.
What about a candidate who doesn't use AI?
Judge the same four acts with no assistant in the story. Someone who framed the problem, opened the source, threw out their own first approach and checked a figure against the filing has demonstrated exactly what the task exists for. Make the assistant available rather than mandatory, and say which it is. How much AI a candidate used is not the measurement, and doing a step by hand because the model was the wrong instrument for it counts in their favor.
How do you stop the task leaking between candidates?
Accept that it will leak and design for the cost. One authored case per role, fixed for the cycle, means a leak costs that case rather than the whole instrument, and everyone on that role met the same material, which is what keeps the comparison meaningful and the step consistent. Rotating per candidate buys a little secrecy and destroys the comparison. Author two or three cases per role and retire one when leakage starts showing up in the submissions.
How does Olive assess this?
Olive is an employer-purchased assessment rather than a task builder. A candidate spends 40 to 60 minutes on an assignment authored for their occupation with an AI assistant available, and a human reviewer writes six findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each attached to a timestamped moment in the session. Outcomes are demonstrated, partly demonstrated or not demonstrated, never a number, and the report carries no hiring recommendation. The candidate is granted the identical document, free, on every tier.
References
- 1. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors ✓ doi.org Revised operational validity estimates: structured interviews .42, job knowledge tests .40, work samples .33, general mental ability .31; job-specific measures outperform general ones.
- 2. Work Samples and Simulations ✓ opm.gov High content and criterion-related validity plus favorable applicant reactions, and appropriate only where competencies are expected on entry rather than trained after selection. Undated guidance, verified 2026-08-24.
- 3. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co 66% name "almost right, but not quite" as their top frustration with AI tools and 45.2% name debugging AI-generated code as more time-consuming than writing it.
- 4. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools ✓ arxiv.org Purpose-built legal research systems returned incorrect information between 17% and 33% of the time despite vendor claims of eliminating hallucination.
- 5. GPT detectors are biased against non-native English writers ✓ pmc.ncbi.nlm.nih.gov Seven detectors over 91 human-written TOEFL essays: 61.3% average false-positive rate, 97.8% flagged by at least one detector, and near-perfect accuracy on US eighth-grade essays.
- 6. Science and Engineering Indicators 2022: International S&E Higher Education ✓ ncses.nsf.gov Temporary-visa students earned 36% of US science and engineering master's degrees in 2019 and 59% of computer science doctorates.
- 7. Employment Tests and Selection Procedures ✓ eeoc.gov A step that decides who advances is a selection procedure and has to be applied consistently to every candidate.
- 8. 29 CFR 1607.14 - Technical standards for validity studies ✓ ecfr.gov Content validity requires a job analysis of important work behaviors and that the behavior demonstrated in the selection procedure be a representative sample of the behavior of the job.
8 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.