Interviewing
Should You Add an AI-Fluency Round to the Interview Loop?
A separate AI-fluency interview round is rarely worth adding. Two cases justify one: a role that's mostly AI-assisted production, and a loop whose rounds you can't change. Otherwise, fold it into an existing round, or build a work sample if the loop has none. Hand the candidate the AI-assisted work of their first month, leave the assistant open, and record what they framed first, which claim they wanted a source for, what they refused, and what they checked outside the chat. The task changes per function. The rubric doesn't.
The takeAI fluency is a temporary name for something older. Take the assistant out of the brief and the question is still whether a person checks a claim before leaning on it, which is what a case round was always for. That is why a dedicated stage ages so badly: it dates itself to a tool generation, and anything named after a tool generation gets retired with it. Nobody has run this long enough to prove it, but the loops that quietly changed what their existing task asks for will still be running that task in three years, and the AI round will have been renamed twice.
Where Olive fits
Open a role and see what the work shows
Doing this per function means authoring a different task and a different answer key for the marketer, the consultant and the analyst, then re-authoring all three as the work moves. Olive ships twelve authored cases per occupation and returns six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session rather than to a number, and the candidate is granted the same report.
Rank your shortlistShould you add a separate AI-fluency round?
In most loops, no. A new stage costs candidate time, panel time and two weeks of scheduling, and what a standalone AI round usually measures (which tools someone has open and how they phrase a request) is the part of this work that turns over every few months. Changing the content of a round you already run is cheaper and produces better evidence.
A round added because AI is on everyone's mind, scored on impressions of how modern someone sounds, gets hard to explain later. Any step that decides who advances is a selection procedure, and it has to be job-related and consistent with business necessity 3. A round built out of the work does not have that problem.
Two situations do justify a separate stage. The first is a role where AI-assisted production is most of the job rather than a tool inside it: a content operation running a model over a queue, a support team whose first draft is always generated. The second is a loop whose existing rounds genuinely cannot be modified: a licensing exam, a fixed panel format, a shared process you do not own. Everywhere else, the extra round is a scheduling cost buying evidence you could get for free. If loop length is the binding constraint, there are ways to get this signal without adding time.
And if your loop has no task round at all (resume screen, two conversations, an offer), the honest answer is not an AI round. It is a first work sample, which happens to be the thing that makes AI use readable.
What does a generic AI round actually measure?
Familiarity with an interface, and a candidate's own account of their habits. Ask someone to name their stack, describe their prompting approach, or write a prompt for an invented task, and the answer is rehearsed, portable between companies, and impossible to falsify in the room. Nothing in it separates a person who checks what a model claims from one who forwards it.
Self-report is weak evidence here even from experts working on their own material. METR ran a randomized trial with 16 experienced open-source developers on 246 real issues in repositories they had contributed to for years. With AI tools allowed the work took 19% longer, and afterwards the developers still believed the tools had sped them up by 20% 1. If people cannot correctly report their own AI-assisted throughput on code they wrote, a forty-minute conversation about their AI habits is not going to settle it either.
The selection-procedure guidance says the same thing in older language. Content validity holds to the extent a procedure is a representative sample of the content of the job, and a procedure resting on inferences about mental processes cannot be supported by content validity alone 4. "Tell me how you would verify a claim the model gave you" is an inference about a mental process. "Here is the claim, here is the source packet, you have thirty minutes" is a work sample.
The tool inventory is the weakest version of all, and it is the most common opening. A candidate listing six AI tools tells you about their subscriptions, not their judgment. What the round should surface is narrower and more observable: three moves you can actually see someone make while the work is happening.
Build the round out of the task that function does with an assistant
Take the AI-assisted task the person will do in their first month, hand it over intact with the assistant available, and put one thing in the brief the model has no access to. Work sample tests ask applicants to perform tasks that mirror the tasks employees perform on the job 2, and the mirror is what makes the round readable. Three functions in the same company need three different tasks.
Marketing. Give them a research folder and ask for a positioning brief in forty minutes. Inside the folder, put one statistic that is on-message, quotable, and drawn from a vendor's own survey of its customers. The assistant will lift it, because it is the most useful sentence in the packet. Watch whether the candidate opens the source before the number reaches the brief, and what they write once they see the sample it came from.
Consulting. Give them a two-page client situation with a real budget line and ask for a sizing and a recommendation. The assistant will produce a confident chain of multiplications from a market figure nobody has checked. Watch whether one link in that chain gets recomputed by hand, and whether the budget appears in the candidate's first move or only in the final paragraph.
Data and analytics. Give them a query result and a dataset that will not correct a wrong story about it. The assistant explains the result fluently and plausibly. Watch whether the candidate re-runs anything, whether they ask what changed in the period before interpreting it, and whether they name the thing the data cannot settle.
Same company, same week, three unrelated briefs. That is the actual finding, and it is expensive. Microsoft researchers classified 200,000 real Copilot conversations against O*NET work activities and found the pattern differs by occupation, including which occupations tend to hand a task to the model outright and which use it alongside an existing workflow 5. A round that ignores that difference is testing the assistant, not the person. Whether that means one shared exercise or one per function is a build decision worth making explicitly, and it is not only a technical-roles problem, since the same design question lands on recruiting, finance and operations hires.
Where does the round fit, and who runs it?
Inside the stage that already carries a task (the work sample, the case, the technical screen), run by someone who can read the work while it happens. Replace the brief rather than appending a sixth conversation. Say in the invite that an assistant is permitted, name which one, and use the same wording for every candidate, so you are not comparing people who believed you against people who played it safe.
The interviewer's job changes more than the schedule does. Reading an AI-assisted session means watching the order of moves, not grading a deliverable at the end: what was asked first, what the assistant offered, what came back edited. A panel that has only ever graded finished artifacts will grade polish, which is the one variable the model reliably controls. Half an hour of calibration on two recorded sessions is worth more than another page of rubric.
One practical caution about content. The brief has to contain something the model cannot infer from its own wording, and it has to be discoverable inside the time limit: a source the candidate can open, a number they can recompute, a constraint printed on the page. A secret you never let them reach is a trick question, and it produces a confident wrong answer at your end as well as theirs. The full mechanics of putting that constraint into an existing question are worth working through before you rewrite the brief.
What do you write down afterwards?
Six acts, each one a yes or a no that two interviewers would mark the same way: what got framed before anything was generated, which claim had a source demanded for it, what the candidate kept for themselves, what existed between the brief and the deliverable, what they refused and on what grounds, and what they tested against something outside the conversation.
Each row needs an observable rather than an adjective. "Strong AI fluency" is unscoreable and will drift by interviewer. "Opened the cited source before quoting it" is a fact two people agree on. Structured interviews work the same way: every candidate gets the same predetermined questions in the same order, and every response is evaluated against the same rating scale and the same standard for an acceptable answer 6. Fold the AI observations into that structure rather than running them as a separate impression. A worked version of the rows themselves is a short exercise you can do before the first candidate.
Two things stay off the sheet. Do not score how much AI the candidate used, because someone who judged the model was the wrong instrument for a step and did it by hand has demonstrated the thing being tested. And do not score the fluency of the writing, which is exactly what the assistant contributes.
Budget for maintenance, because this is the part people skip. Work samples are costly to develop in both time and money, and they require periodic updating as the job changes 2. Yours will change faster than most: a brief written around a model's blind spot in March may not have one in September. Plan on re-authoring the task roughly every two quarters and on retiring it the week it appears in a forum thread. Building this per function is real work, and often still the right call at low volume. Olive is one of several instruments in this category. Multiple-choice AI literacy tests, code-collaboration graders, live AI interviews and unwatched take-homes all exist, and each of them trades something different away.
Common questions
How long should the AI-assisted part of the round take?
Between thirty and sixty minutes of task time, inside a round you already run. Under thirty, the candidate gets one exchange with the assistant and no chance to catch anything; past an hour you are measuring stamina and losing finalists. If the task genuinely needs longer, move it out of the live round entirely rather than stretching the interview, and ask for the working record along with the deliverable.
Does a prompt-engineering certificate cover this?
No. A certificate shows someone sat through a course on how to phrase requests, which is the part of this work that ages fastest and the part a model will do for them anyway. It says nothing about whether they open a source, refuse an answer, or check a number against something outside the chat. Treat it as evidence of interest, not evidence of judgment, and put the same task in front of certificate holders as everyone else.
Should the round be pass or fail on its own?
No. Make it one input among the others, with the findings written out per behavior rather than reduced to a verdict. A single stage that gates advancement on its own carries more weight than one afternoon of work can support, and it is the version that becomes hard to explain to a candidate who asks. Write what happened, hand it to the panel, and let the decision stay with people who have the rest of the evidence.
What if the role doesn't use AI much yet?
Then don't run the round. A work sample is appropriate when the competency is one applicants are expected to have on entry, not one you will teach after selection. If the function will start using an assistant next quarter, that is a training question rather than a hiring filter. Keep the existing task round, and revisit when the work has actually changed.
Is an AI-open work sample fairer than a live quiz?
It has a better claim to job-relatedness, which is the fairness question a regulator asks. A task that mirrors the job supports its own validity argument; a quiz on tool names does not. That is not the same as being bias-free. Nothing here is audited by default, subgroup effects depend on what the task actually measures, and you still owe candidates the accommodations process you run for every other assessment.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers on 246 real issues took 19% longer with AI tools allowed, and afterwards still believed the tools had sped them up by 20%.
- 2. Assessment and Selection: Work Samples and Simulations ✓ opm.gov Work sample tests require applicants to perform tasks that mirror the tasks employees perform on the job; development costs are high in time and money and the tests require periodic updating.
- 3. Employment Tests and Selection Procedures ✓ eeoc.gov A test or selection procedure used in an employment decision must be job-related and consistent with business necessity.
- 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 ✓ ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job; a procedure based on inferences about mental processes cannot be supported by content validity alone.
- 5. Working with AI: Measuring the Applicability of Generative AI to Occupations ✓ arxiv.org 200,000 anonymized Copilot conversations classified against O*NET work activities; the method distinguishes occupations likely to delegate tasks to AI from those using it to assist existing workflows.
- 6. Assessment and Selection: Structured Interviews ✓ opm.gov In a structured interview all candidates are asked the same predetermined questions in the same order and all responses are evaluated using the same rating scale and standards for acceptable answers.
6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.