Interviewing

Four Questions for a Manager Whose Team Ships With AI

A manager candidate who will run a team that uses AI daily gets four questions, all about decisions they already made. What did they let the team automate, and what did they refuse. How did they find out somebody was over-relying on a model, and what did they do about it. How do they judge a junior's work now that every artifact arrives clean. What did their review process catch last quarter. A policy statement is a weak answer. A decision with a cost attached is a strong one.

The takeIf your manager loop still carries a generic question about AI strategy, that is the one to cut. It has a correct-sounding answer any manager can assemble in the car, it rewards fluency in change-management vocabulary, and it says nothing about the two calls this job actually turns on. Those are which work the manager refused to hand to a model, and how they keep developing people whose output stopped carrying visible evidence of effort. Ask about those and the rehearsed paragraph has nowhere to go.

Where Olive fits

Open a role and see what the work shows

A manager candidate can describe how a team should work with a model, and the description comes apart from the practice more often than an interview can see. Olive assesses one person against six evidence-anchored dimensions in a single occupational session, and a human writes every word of the six findings it returns.

Rank your shortlist

Ask what the team automated and what they refused

Start with the pair, never the first half alone. Asked what the team automated, most candidates have a rehearsed answer waiting. Asked what they refused to automate, and why, they have to name a decision. A manager who can give both halves, with a reason attached to the refusal, drew that line themselves.

The distinction has a name in the labor data. Employment of 22-to-25-year-olds in the two most AI-exposed occupational quintiles fell about 11% between November 2022 and June 2026 while the same age group in the three least-exposed quintiles grew about 10%, and the declines concentrate in occupations where AI substitutes for human tasks rather than complementing workers 1. Those measures come from occupation-level usage patterns, not from how any single employer deploys anything, and the authors describe the work as descriptive rather than causal. It says nothing about your company.

What it does give you is the right axis for the question. A manager who automated the parts that were substituting for judgment has made a different bet from one who automated the parts that were eating time before the judgment started, and the two produce very different teams within a year. Ask which of those they did.

The follow-up worth asking every time: what did that refusal cost. A manager who refused to automate something and cannot say what it cost has probably refused nothing. If the role is mostly about supervising automated work, assessing candidates who manage agents is a narrower version of this question.

How did they find out somebody was over-relying on a model?

Listen for the moment they noticed, because this is the question with the fewest good fake answers. Over-reliance is invisible in the artifact and visible only in what happens next: an explanation that falls apart on the second question, a figure nobody can trace, a fix that takes longer than the draft saved. A manager who has met it once can describe how they noticed, and what they changed the following week.

The useful research here is about who the tool helps. In a randomized trial with 640 Kenyan small-business owners given a GPT-4 assistant over messaging, there was no detectable average effect on business performance, but high performers at baseline gained just over 15% while low performers did about 8% worse, and the gap came from which advice owners chose to act on rather than from differences in the advice itself 2. Kenyan small businesses in 2023 and 2024 are not a knowledge-work team, the average effect could not be distinguished from zero, and this cuts against rather than refutes the studies where AI lifted weaker performers most.

The transferable part is the mechanism a manager is being hired to handle. The same assistant, given to two people on the same team, can raise one and sink the other, and the difference is whether the person can judge what came back. A manager who understands that will describe an intervention aimed at judgment: pairing, a required check, a rule about tracing figures.

Weak answers here reach for tooling or for a ban. Both are the same move, which is turning a coaching problem into a policy. What to listen for instead is a manager who names the person, the moment and the change, without turning it into a story about somebody being caught out. Structuring the first ninety days of an AI-heavy hire is what a good answer to this question usually looks like written down.

How do they judge a junior when every draft arrives clean?

They read the process. The rough edges in a draft used to show where somebody was struggling, and a clean submission said they had it, and that reading stopped working once first drafts began arriving polished regardless of understanding. Strong answers name what replaced it: questions asked, checks run, what got escalated and when.

Do not accept an answer built on being able to tell. In an ACL 2021 study, non-expert evaluators asked to distinguish GPT-3 text from human writing across stories, news articles and recipes performed at random chance, and three quick training methods lifted accuracy only to about 55%, inconsistently across the three domains 3. That was 2021-era output judged by crowdworkers rather than managers reading work in their own field, so it is not a measurement of your reviewers. It is enough to disqualify a manager whose whole development plan rests on spotting the difference.

The answers that hold up move the coaching upstream of the artifact. A manager who asks a junior to bring the question they started from, or who reviews the checks a junior ran, or who sets a weekly session where the junior explains one decision in detail, has rebuilt a signal that a polished draft cannot fake. So has one who simply changed what they assign: smaller pieces, tighter feedback, more work where the answer is verifiable.

This is also where the manager's own habits show. Hiring managers who do not use these tools themselves tend to grade the artifact because it is the only thing they know how to read, and a candidate for a management role should be able to say what they read instead.

What separates a decision from a policy statement?

A cost. A policy statement describes what should happen and carries no evidence anybody tested it. A decision names what was chosen, what was given up, and how it turned out. Ask what their review process caught last quarter and the difference appears at once: either there is an example with a date attached, or there is a description of a process nobody has run.

Be skeptical of productivity claims, because the honest aggregate numbers are small. A nationally representative US survey asked people how many extra hours the previous week's work would have taken without generative AI: mean self-reported savings were 5.4% of work hours among people who use it at work, which works out to 1.4% of total work hours across all workers 4. That is self-reported time saved rather than measured output, from late-2024 US data, and the authors call it a rough estimate. A manager candidate claiming their team got 40% faster is quoting a feeling.

What a strong answer sounds like, in one sentence each. The team automated first-draft release notes and refused to automate incident write-ups, because a write-up nobody thought through is worse than none. Somebody was accepting generated figures without tracing them, which showed up in a client meeting, so tracing became a required step and reviews got ten minutes longer. Juniors now bring the question they started from, because the draft stopped telling anyone anything.

When a manager's answer is that the capability has to come from outside, hiring for AI skills against training the team you have is the decision sitting behind the requisition.

Ask the same four questions of every manager candidate for the role and write the answers against the same four headings. Without that, the debrief compares how well each of them talks. It is the same discipline as giving each panel seat one AI question and grading each level against its own bar.

See the benchmarks

Common questions

What if the manager candidate has never led a team that used AI?

Ask the same questions about the closest thing they have led through. Every manager has decided what to hand off and what to keep, has met somebody relying on a tool or a template they did not understand, and has judged work whose surface looked better than its substance. A candidate who answers well about outsourced work, contractors or a shared component library is showing the judgment the role needs. What is disqualifying is having no example of any of it.

Should a manager candidate be able to use the tools themselves?

Enough to read the work, which is a lower bar than fluency and a higher one than none. A manager who has never watched a model produce a confident wrong answer tends to over-trust polished output and under-trust the people who question it. Practical version: they should be able to describe what the assistant their team uses does badly. They do not need to be faster with it than the people they manage, and expecting that usually selects for the wrong candidate.

How do I ask this without it turning into a policy debate?

Anchor every question to something that already happened. What did you refuse, when did you notice, what did the review catch last quarter. A question about the past has a right answer only for the person who lived it, while a question about principle has a right answer available to everyone in advance. If a candidate keeps returning to principle, ask for the date and the name of the project, once. Twice is a finding.

Is this a separate round or part of the normal manager loop?

Part of the normal loop, in one seat. These four questions replace whatever generic AI question the loop currently spreads across everybody, and they take about twenty-five minutes with follow-ups. Adding a dedicated round for it signals that AI use is a separate competency from managing, which is the framing that produces rehearsed strategy answers. It is a management question, so it belongs in the management interview.

What if the whole team already uses AI and the manager is the gap?

Then the hiring decision is about whether that gap closes, so ask the training question directly. A manager who has deliberately learned enough to review this work will say how; one who has avoided it will describe delegating the judgment to whoever on the team seems most current, which is a real answer and a real risk. The same choice, at team level, is whether to hire the capability or build it in the people already there.

References

  1. 1. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Stanford Digital Economy Lab (Brynjolfsson, Chandar and Chen), 2026. digitaleconomy.stanford.edu Supports the substitute-versus-complement distinction behind the automate-and-refuse question, including the 11% decline and 10% growth figures for 22-to-25-year-olds.
  2. 2. The Uneven Impact of Generative AI on Entrepreneurial Performance eScholarship, University of California (Berkeley Haas / Harvard Business School), 2024. escholarship.org Supports the claim that the same assistant can raise one user and sink another: high performers gained just over 15%, low performers did about 8% worse, with no detectable average effect.
  3. 3. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL-IJCNLP 2021), 2021. aclanthology.org Supports the instruction to reject any development plan resting on a manager being able to tell generated text apart: untrained evaluators performed at chance, rising to about 55% with brief training.
  4. 4. The Rapid Adoption of Generative AI (NBER Working Paper 32966) National Bureau of Economic Research, 2025. nber.org Supports the skepticism about large claimed productivity gains: mean self-reported savings of 5.4% of work hours among users, 1.4% across all workers.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.