Teams

Name the Human Who Checks It, or Nobody Did

When AI-assisted work turns out wrong, accountability stays with the person who shipped it. That never moved, and saying it plainly ends most of the argument. What has to be built is the check they are accountable for: on each recurring deliverable, name the one claim that must be verified against something outside the chat, and name who verifies it when that is not the author. Treat a miss as a missing check rather than a character finding.

The takeGovernance language has done real damage here. Human in the loop was written to be audited, and it audits well: the sentence is either in the policy or it is not. Teams then adopt it as though writing it down were the work, and what they have bought is the appearance of oversight with none of the substance. Which one they bought is discovered during an incident review, by a person who is about to be held to a check nobody ever assigned them.

Where Olive fits

Open a role and see what the work shows

Every finding in an Olive report is written by a person and carries the excerpt it rests on, so a reader can see what the finding is based on instead of taking a verdict on trust. The report is an input to a decision that stays with the human reading it.

Rank your shortlist

Who owns the mistake?

The person who shipped it. Not the model, not the vendor who sold it, and not whoever happened to be standing nearest when the error surfaced. Saying that out loud before anything goes wrong is worth more than any policy paragraph written after, because most of the argument that follows an incident is really an argument about whether ownership moved. It did not.

One arena where a US agency answered this in writing went the same way. In technical assistance issued in May 2023, the EEOC said an employer administering a selection procedure may be responsible under Title VII even when an outside vendor built the tool, and may be responsible for the acts of agents including software vendors given authority to act on its behalf 1. That document was removed from the agency's site in January 2025 and never had the force of law, it addressed hiring rather than internal work, and it has not been tested as such in a decided case. Cite it for the shape of the answer, not the authority.

Inside a team the basis is simpler still. The author is the last person who could have caught the error and the only one who knew what the work was for. Nobody downstream can supply either. Whatever process gets built on top, that fact is the floor it sits on.

Why doesn't 'human in the loop' settle it?

Because it names no human and no moment. The phrase answers a regulator asking whether oversight existed at all, and for that purpose it is a good answer. The team's question is a different one: who was supposed to have caught this, and against what. A sentence that satisfies the first question leaves the second entirely open, which is how accountability ends up landing on proximity.

There is a harder problem underneath. A review can feel thorough without being thorough, and better presentation makes that worse rather than better. In an online experiment, 410 German-based HR managers compared recruiting dashboards against versions enriched with three styles of explanation; the explanations improved perceived helpfulness and trust among users with moderate or high AI literacy without increasing their objective understanding, and the more complex explanations may even have reduced accurate understanding 2. That is one experiment on simulated dashboards, published as a preprint, and it is about explanation interfaces rather than AI tools in general.

The direction is what carries. The feeling of having understood an output and having understood it came apart, and the interface closed the feeling gap. A reviewer who read a fluent deliverable and felt satisfied has produced exactly the artifact the policy asked for and none of the safety it implied. Oversight that is not attached to a specific claim is a mood.

Name the claim, the source and the checker

For each recurring deliverable, write three lines. The one claim that carries the decision. The thing outside the model that claim gets checked against. The person who does the checking when that is not the author. Three lines is the whole register, and it is what the word accountable attaches to, because it converts a disposition into an event someone either performed or did not.

Worked, it looks like this:

  • Weekly revenue summary. Claim: the headline figure. Source: the warehouse export. Checker: the author, with a monthly spot check by the manager.
  • Client memo citing precedent. Claim: every authority cited. Source: the reporter itself. Checker: the author, who pulls the case and reads the holding rather than the summary.
  • Meeting notes that become a decision record. Claim: any sentence attributing a commitment to a named person. Source: the recording, or that person. Checker: whoever chaired.

The third example is the one teams skip, and it is not hypothetical. An audit of one widely used speech-to-text model found roughly 1% of transcription segments contained entire hallucinated phrases appearing nowhere in the audio, and 38% of those carried explicit harms such as invented violence or implied false authority 3. The same paper found no comparable hallucinations from four other commercial services on the same segments, and it audited an aphasia research corpus rather than meeting recordings. What transfers is not the rate. It is that a fabricated sentence looks exactly like a real one to whoever reads the transcript six weeks later. That is also why training on AI use is not the same as checking the work afterwards.

Treat a miss as a missing check

Run the review on the process, not the person. Four questions: which claim carried the decision, which check would have caught the error, did that check exist, and who owned it. None of the four asks whether AI was used. Asking that one directly is how a team learns to stop mentioning the assist, and a team that hides the assist has lost the only evidence anyone had.

When the answer is that no check existed, the outcome is a new line in the register with a name against it, and the incident is closed. That is not leniency. It is the accurate finding, and it produces a change that survives the week.

When the check existed, was owned, and was skipped, that is a performance conversation, and it is a fair one precisely because the expectation was written down beforehand. The two outcomes feel similar in the room and are completely different in what they should produce.

Early on, almost every miss is a missing check. As the register fills, the balance shifts, and that shift is the clearest evidence the thing is working. Watch it particularly closely for anyone new, whose habits are still forming: the first ninety days set what they think normal looks like, which is the real subject of structuring the start for a hire who runs everything through AI. And keep the register beside the list of tasks the model may not produce at all, since one decides what gets made and the other decides who checks it.

See the benchmarks

Common questions

Is the reviewer accountable too?

Yes, for the check they were named for, and for nothing wider. That narrowness is deliberate: a reviewer asked to be responsible for a whole deliverable will skim all of it, while a reviewer asked to confirm one figure against one source will actually do that. If a deliverable needs three separate checks, name three, with the owner written next to each. Diffuse review responsibility reliably produces diffuse review.

Does a disclosure rule help with accountability?

Only slightly, and it can hurt. Knowing a model was involved does not tell anyone which claim went unchecked, and a rule enforced through suspicion teaches people to describe their process less. Disclosure is most useful when it is specific and low-stakes: a line saying which steps were assisted and what was verified. That is a note about method, which is worth having, rather than a confession, which is not.

What happens when an agent takes an action on its own?

Accountability sits with whoever authorised the agent to act in that scope, which is a decision a person made on a particular day. The useful move is to write the scope down in the same register: what the agent may do without a human, what it may never do, and who reviews the log and how often. An unbounded authorisation is the mistake, and it happens before any error does.

Should an AI-caused error affect someone's performance review?

A skipped check belongs in a review, because that expectation was set in advance and the person had a fair chance to meet it. An error that happened anyway does not, since errors happen in work that was checked properly too. Keeping those separate is what stops the register from turning into a blame instrument, at which point people stop reporting near misses.

Who does the check when the author is the most senior person?

Someone else, chosen for access to the source rather than for rank. The person who can open the warehouse, pull the case, or call the customer is the right checker regardless of level, and a junior reading a figure back against an export is doing real verification work. Seniority is a poor proxy here, and self-checking by the most senior author is where the register most often quietly stops running.

References

  1. 1. Select Issues: Assessing Adverse Impact in Software, Algorithms, and Artificial Intelligence Used in Employment Selection Procedures Under Title VII of the Civil Rights Act of 1964, Question 3 (archived capture, 2025-01-25) U.S. Equal Employment Opportunity Commission, via the Internet Archive Wayback Machine, 2023. web.archive.org Supports the claim that responsibility for a tool's output stays with the party that used it, in one arena where a US agency said so in writing.
  2. 2. Explained, yet misunderstood: How AI Literacy shapes HR Managers' interpretation of User Interfaces in Recruiting Recommender Systems arXiv (Yannick Kalff, Katharina Simbeck), 2025. arxiv.org Supports the claim that a review can feel thorough without being thorough, and that presentation closes the feeling gap rather than the understanding gap.
  3. 3. Careless Whisper: Speech-to-Text Hallucination Harms arXiv (also published at ACM FAccT 2024), 2024. arxiv.org Supports the claim that an automatic transcript can carry sentences nobody said, indistinguishable from real ones to a later reader.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.