Interviewing

Build the AI Question From the Role's Own Tasks in an Afternoon

Derive AI interview questions from the role's tasks, not a borrowed list. Write down the three tasks the role spends most of a week on, mark the step in each where a model now produces the first draft, and put the question at the point where an unchecked draft would reach a customer, a payer or a regulator. That yields two or three questions whose right answer you can state before anyone walks in, and the work is the same afternoon whether the role is nursing or demand generation.

The takePer-function question lists are a content format rather than a hiring method. The five questions under every job title are the same five with the noun swapped, which is exactly why candidates have already rehearsed them and why the answers all sound competent. An afternoon spent deriving one question from your own task list beats the entire list, and the derivation is the part that survives the next model release. Lists go stale in a quarter. A seam in your own workflow does not.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing how they would check a drafted handover; it cannot capture them checking one. Olive puts that in front of the person as work, with an assignment grounded in one occupation, an assistant that will do the whole thing if nobody stops it, and a human reviewer who writes six findings with the moment behind each one.

Rank your shortlist

Why doesn't a generic AI question list work?

Because the list copies the wording and drops the derivation. Every per-function version runs the same handful of questions with the job title changed: describe a time you used AI, what are its limits, how do you check the output. Those have well-known good answers now, so a prepared candidate returns one and you learn what they read last week. A question built from a task only your team performs has no rehearsed answer available.

The practical tell is whether you can say, before the interview, what a good answer contains. Ask a nurse how they use AI at work and forty answers are defensible, which means none of them separates anyone. Ask what they would do when the drafted discharge summary lists a medication the patient stopped two admissions ago, and there is a right answer, a wrong answer, and a small band of partly right answers that show you exactly where the person sits.

The derivation transfers where a copied list does not, and that is the whole difference. Before you write anything, it helps to know what AI actually does in this role day to day, because the answer usually differs from what the job description claims and from what the person who wrote it assumed.

Start from the three tasks the role does most

Open the job description, ignore the responsibilities list, and write down the three things this person will genuinely spend most of a week doing. Then ask the two people already doing that job which step of each task now begins with something a machine produced. The result is short, specific to your team, and the only input the rest of this needs. It usually takes under an hour of conversation.

What comes back will be narrower than the job title suggests. Indeed's GenAI Skill Transformation Index rated almost 2,900 work skills found in US job postings and mapped them onto more than 53.5 million postings, placing 40% of skills in minimal transformation, 19% in assisted, 40% in hybrid and 1% in full, with 46% of the skills in a typical posting falling into the hybrid or full categories 1. Read that carefully before repeating it: large language models produced those ratings by scoring skills, and nobody watched the work being done, so it measures what current models are judged capable of on one job board over twelve months. Hybrid in that scheme explicitly means human oversight is still required.

So the interesting unit is the seam inside a task, not the task itself. A financial analyst still builds the model; the comparable-company pull now arrives pre-drafted. A recruiter still writes the requirement; the first version comes out of a tool. A paralegal still files; the summary of the deposition arrived already written. Write down the step, the artifact it produces, and who receives that artifact next. Three tasks, three seams, three receivers, one page.

Write the question at the seam where a bad draft escapes

Pick the point in each task where an unchecked draft leaves your team: sent to a client, filed with a payer, merged, published, read aloud to a patient. Put the question there. That is where the cost lands, so it is where a candidate's habits become visible and where you can say in advance what a passing answer has to contain. Anything upstream of the seam is preference, and preference has no right answer.

Aim there because people are poor judges of which side of the line a task sits on. In the Boston Consulting Group field experiment, on one task deliberately chosen to sit outside the model's capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right against 60% and 70% in the two AI conditions 2. One task, one sample, a model from 2023, and the group given a prompt-engineering overview did worse than the group given none. What travels is the failure shape, which is a confident wrong answer on work that looked routine.

Two question forms cover most roles once you have the seam:

1. The artifact question. Hand over a real first draft the assistant produced for that task, with one thing wrong in it that your field would catch. Ask what they would not send, and how they would check the part they are unsure about. Fifteen minutes, and it is the closest thing to the job an interview can hold. 2. The boundary question. Ask which part of this task they would not hand to an assistant at all, and what it would cost if they did. The answer turns on accountability, and it is hard to fake without knowing the work.

Grade both against what the artifact actually contained rather than against how the answer sounded. The four lines worth putting on the scorecard covers what to write down while they talk.

What does this look like in a non-technical role?

The derivation does not change; only the artifact does. A demand generation manager checks a claim inside a drafted campaign brief. An accounts payable clerk checks a coded invoice against the contract terms. A charge nurse checks a drafted handover against the chart. In every case the moves are identical: name the seam, produce the artifact, ask what they would not send onward and why.

Most job titles that mention AI now sit outside technology. Indeed Hiring Lab counted job titles carrying AI language in the employer's own raw title and found US AI-touched titles rising from 264 in the first quarter of 2022 to 822 in the first quarter of 2026, with 63% of them outside tech occupations 3. That counts distinct titles with at least five postings in a quarter, so it tracks how far the vocabulary has spread across role names and says nothing about volume of demand, and the count dipped before it rose. The direction is enough for the point here: the manager writing an AI question this month is often not a technical manager.

Three worked seams, to show how little the shape changes:

  • Demand generation. Seam: the brief goes to an agency. Question: here is the drafted brief, one of these three claims about the product is not true, tell me which and how you would confirm it.
  • Accounts payable. Seam: the coded invoice enters the payment run. Question: the assistant coded this to the wrong cost center and the total matches the PO. What would you check before approving it?
  • Ward handover. Seam: the summary is read at shift change. Question: this drafted handover omits an allergy that is in the chart. Walk me through what you do in the next two minutes.

Each one has a right answer, each takes under ten minutes, and none of them appears on a list of interview questions for that job title. If the role sits outside engineering and this still feels like a technical exercise, screening for AI judgment in finance and marketing roles works the same derivation through longer examples.

See how it works

Common questions

How many AI questions should a loop actually carry?

Two or three, once, in the round where the work is discussed. More than that and the loop turns into a survey about tooling rather than an interview about the job. Two questions derived from real seams beat six borrowed ones, because you can state what a passing answer contains for the two and cannot for the six. Ask them the same way of every candidate, in the same order, so the answers can be compared.

What if nobody on the team uses AI for this role yet?

Then the question is premature and the honest move is to leave it out. A requirement nobody in the team can demonstrate has no passing answer, so anything you ask measures how well a candidate performs enthusiasm. Spend the afternoon on the task inventory anyway. If no step of the work currently begins with a machine-written draft, you have learned something useful about the role, and you can revisit when that changes.

Can I reuse one derived question across several roles?

Only where the seam is genuinely shared: the point where an unchecked draft leaves the team. Two roles that both send drafted client-facing copy can share the question with a different draft in each. Two roles that differ in who receives the output cannot, because the cost of a bad draft is what makes the answer scoreable. The cheap test: if you have to change what a passing answer contains, you have changed the question and should write it out separately.

Does a candidate need to have used the exact tool your team uses?

No, and requiring it narrows the pool for no measurable gain. What transfers is the habit of checking a confident draft against something outside the conversation, which someone can carry from any assistant to yours. Tool-specific fluency is a week of onboarding. Judgment about when not to trust a draft is the part that took years, and it shows up in the answer regardless of which product produced the draft they practiced on.

How do I keep the question from leaking to later candidates?

Assume it leaks and design so that leaking costs little. An artifact question with a planted error survives disclosure better than a trivia question, because knowing the format does not tell anyone what is wrong inside this particular draft. Keep two or three drafts in rotation for the same seam and swap them between candidates. If a question stops working the moment it is described out loud, it was testing recall rather than judgment.

References

  1. 1. AI at Work Report 2025: How GenAI is Rewiring the DNA of Jobs Indeed Hiring Lab (Annina Hering and Arcenis Rojas), 2025. hiringlab.indeed.com Supports the claim that AI lands on a step inside a task rather than on the whole task, and that hybrid transformation still assumes human oversight.
  2. 2. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that the failure mode worth interviewing for is a confident wrong answer on a task that looked like one AI handles well.
  3. 3. AI Is No Longer Just a Tech Occupation Story: It's Spreading Across Job Titles in the US and Europe Indeed Hiring Lab (Pawel Adrjan), 2026. hiringlab.indeed.com Supports the claim that AI language has spread into non-technical job titles, which is why the derivation has to work outside engineering.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.