Assessment design

Whose AI Account the Candidate Uses Changes What You Can Read

Provide the AI account for an assessment when submissions have to be comparable across candidates, or when the exercise turns on how a specific model behaves. Let candidates bring their own tools when you are hiring for how someone works inside a stack they already know. Either way, name the model and version in the brief, say up front whether the session is kept and for how long, and never let a paid tier become the price of entry.

The take"Candidates may use approved tools" is the sentence that gives this away. Approved by whom, and paid for by whom? A twenty-dollar subscription is nothing to a company and a real calculation for somebody between jobs, so a policy that quietly assumes one has added a means test to a hiring stage without writing it down. Teams that do provide accounts usually have not noticed the other half of the trade: from that moment they are holding a record of how a person thinks, produced before that person worked for them.

Where Olive fits

Open a role and see what the work shows

Olive takes the bring-your-own side of this: the candidate works in the assistant they already use and pastes each exchange into the session transcript, and the provided assistant built into the workspace is not switched on in production yet. Screen capture covers the assessment tab only, and declining it is a supported outcome.

Rank your shortlist

Which questions decide it?

Two, asked in order. Does the work have to be comparable across candidates? Does the exercise turn on how a particular model behaves? Two yeses and you supply the account. Two noes and the candidate's own setup is the better test, because part of what you are hiring for is fluency in a toolchain they already have running.

Comparability binds when several people do the same assignment and the difference between them is the decision. If one candidate is working against a current frontier model and another against whatever the free tier serves this month, the gap between their outputs is partly a gap between products, and no rubric can separate that from the gap between people.

Model behavior binds when the exercise is built around the assistant overreaching in a specific way. Assignments designed to surface judgment usually contain a trap: a plausible citation that does not exist, a calculation the assistant will do confidently and wrongly, a scope the assistant will silently expand. Those traps are model-specific and version-specific. Hand the candidate a different assistant and you have handed them a different exercise.

Bring-your-own wins in the opposite case, and it is more common than the vendor framing suggests. For an engineer with their own editor, their own prompt library and five years in a codebase, taking the setup away measures friction rather than judgment. Ask instead for the work plus a short account of what they checked, and the tooling question stops mattering.

One mixed case comes up constantly: a hybrid where you provide an account and permit anything else the candidate normally uses, then ask them to say what they used. That is fine for a bring-your-own exercise and it quietly destroys comparability, so pick which of the two you are running while writing the assignment brief, rather than after the first submission lands.

Never let a paid tier become the price of entry

Whatever you decide, a candidate's ability to pay for a subscription cannot be part of it. A paid tier runs about twenty dollars a month, which is a rounding error for a company and a genuine decision for somebody out of work. If the assignment cannot be done well on a free tier, either supply accounts or redesign the task, because otherwise the exercise partly measures who is currently employed.

The assumption underneath the usual policy is also weaker than it looks. In nationally representative US surveys run in August and November 2024, nearly 40% of the population aged 18 to 64 had used generative AI, 23% of employed respondents had used it for work at least once in the previous week, and 9% used it every work day 1. Those are self-reported, and "at least once last week" is a very low bar, so the share of candidates arriving with a paid subscription and a settled workflow is smaller still. Adoption figures also age in months rather than years, so date any number you rely on.

The practical version is three lines in the brief. Say whether a tool is provided. If it is not, say plainly that a free tier is enough for this assignment, and make that true. If a paid feature genuinely is required, supply it or drop the requirement.

The posting matters here too, because the candidate decides whether to apply long before they see the brief. Whether tool access belongs in the job ad is a separate call, but the same rule holds: if access is part of the exercise, it is your cost.

Name the model and version in the brief

Write the model name and version next to the tool rule. Assistants differ enough at the task level that "AI is allowed" names no condition at all, and the same assignment run six months apart is not the same assignment. It also gives you something concrete to point at later, when results shift and nobody in the debrief can say what changed.

How much difference a version makes is easy to underrate. In the Boston Consulting Group field experiment, on one task deliberately chosen to sit outside the assistant's capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right, against 60% and 70% in the two AI conditions 2. That is one task in one sample with a 2023 model, and it is not evidence that assistants make people worse in general. What it does show is that people could not tell which side of the capability line a task was sitting on, and that line moves with every release.

So your assignment has a position relative to that line, and the position drifts without anyone touching the brief. An exercise that reliably separated candidates last spring can turn into a task the assistant simply completes. That is the mechanism behind a pass rate that jumps after a model release, and recording the version is what turns the jump from a mystery into a dated event.

Two lines of bookkeeping cover it. Put the model and version in the brief, and keep a note of which version was in force for each round alongside the submissions. When the results move, you will know whether to fix the assignment or the answer key.

Say what happens to the session log

Four things belong in the brief: what gets recorded, who reads it, how long it is kept, and what happens to it for a candidate you do not hire. Supplying an account means you are holding a transcript of how somebody thinks, produced before they worked for you, and the default retention setting on a workspace product is not a decision anybody actually made.

A regulator has already audited what recruitment tools do with candidate data. The UK Information Commissioner's Office audited providers of AI sourcing, screening and selection tools between August 2023 and May 2024 and reported in November 2024, making 296 recommendations, of which 97% were accepted with actions set. It found that some tools gave recruiters search functionality capable of filtering out candidates with certain protected characteristics, and that others inferred gender, ethnicity and other characteristics from an application or from a name alone, which the ICO concluded is not accurate enough to monitor bias and was often processed without a lawful basis and without the candidate's knowledge 3. Those were consensual audits of self-selecting vendors rather than a representative sample, and they are UK data-protection findings rather than a discrimination ruling. The finding still lands where it matters here: some tools processed more about a candidate than the decision needed, and did it without the candidate knowing.

Four commitments cover almost every version of this, and all four fit in a paragraph a candidate can read: the transcript is used only to evaluate this assignment, only named reviewers read it, it is deleted on a stated schedule, and it is not used to train anything. Then honour the schedule, which is the part that decays quietly. The broader question of what you may keep from a candidate's AI session, and for how long deserves a settled answer before the first invite goes out rather than after.

Asking for a transcript from a candidate's own account raises a different question, because you are asking somebody to hand over a working session from a personal tool. Ask only for the part covering the assignment, tell them why, and be specific about what a reviewer is looking for in it. A vague request for chat logs invites people to curate, which leaves you reading an edited document and calling it evidence.

See what gets scored

Common questions

Can we require candidates to use a specific model?

Yes, if you supply access to it. Requiring a named model that the candidate has to buy is the paid-tier problem again, and requiring one they cannot legally use in their country is worse. Where the exercise genuinely turns on one assistant's behavior, provide the account, say so in the brief, and expect a few candidates to be slower with an unfamiliar tool. Build that into the time budget rather than into the rubric.

What if a candidate says they do not use AI at all?

Let them do the assignment without it, and grade the same rubric. An AI-open exercise permits AI rather than requiring it, and treating non-use as a failing is a shortcut to grading a preference instead of the work. The one thing to check is whether the assignment is still passable in the time budget without an assistant. If it is not, the cap was set for AI-assisted work and that needs saying out loud.

Should the assignment be done in a monitored environment?

Only if the environment is what you are testing. Monitoring changes what candidates do, adds an accessibility and consent burden, and buys very little, since anything worth doing dishonestly can be done before or after the session. Prefer an exercise where the interesting evidence is what somebody rejected, checked or reframed, which is visible in the work itself and does not need a watcher.

Does providing an account create legal exposure we did not have?

It creates data obligations rather than new hiring liability. A provided account produces a record about an identifiable person, so retention, access and deletion questions attach to it wherever data-protection law reaches. Which rules bind you depends on where the company and the candidate sit, and that is a question for counsel rather than for a policy page. The safeguard travels anyway: collect only what the decision needs, name the retention period in the brief, delete on schedule, and ask counsel before keeping session transcripts beyond a single hiring round.

How do we compare a candidate who used AI with one who did not?

Grade against the answer key, not against the process. The key defines what a strong submission establishes, which claims it checks, and which errors matter, and none of those lines needs to know how the work was produced. If the rubric cannot be scored without knowing whether an assistant was involved, the rubric is measuring tool use rather than the job, and that is worth fixing before the round rather than during it.

References

  1. 1. The Rapid Adoption of Generative AI (NBER Working Paper 32966) National Bureau of Economic Research, 2025. nber.org Supports the late-2024 adoption baseline behind the claim that a working paid AI setup cannot be assumed of a candidate.
  2. 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the outside-the-frontier result behind the claim that which model a candidate uses, and which version, changes what an assignment measures.
  3. 3. AI tools in recruitment: Audit outcomes report Information Commissioner's Office (UK), 2024. ico.org.uk Supports the audit counts and the finding that some recruitment tools processed more about a candidate than the decision needed, without the candidate knowing.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.