Teams
A Veto Is Fine If It Has to Name What It Saw
Give the hiring manager the decision and give anyone on the panel the ability to block, with one condition attached to both: a no has to name the specific thing that was said or done and the round it happened in, and a yes carries the same obligation. A veto that cannot cite an observation is a preference wearing a procedure's clothes. Written that way, the rule survives the meeting it was written for.
The takeThe folklore treats a veto as a status question, which is why it gets argued about by job title and settled by whoever has the most tenure. The clause that decides whether it works is evidentiary. A blocking right with a citation requirement is a good rule, because the cost of blocking becomes having to be specific in front of colleagues. The same right without one is a way for the least accountable opinion in the room to win, and it usually will.
Where Olive fits
Open a role and see what the work shows
A dissent that has to cite an observation needs observations to exist in the first place. Olive returns six findings from one working session, each quoting the timestamped moment it rests on and written by a human reviewer, which is an input to the panel's argument rather than a verdict handed to it.
Rank your shortlistWhat is the rule when a panel splits?
The hiring manager decides, any panel member can block, and every position, for or against, names what was said or done and in which round. Three sentences, written down before the loop opens. That allocates authority in the first two clauses and puts the whole weight on the third, which is the reverse of the usual ordering.
The manager decides because somebody has to own the result, and a majority vote converts the room's composition into a decision rule. The block exists because a manager overriding the structured evidence is not, on average, exercising better private information. Across 15 firms hiring low-skilled service workers, introducing a job test raised completed job tenures by 0.23 log points, just over 25%, and comparing managers at the same location, a one standard deviation higher rate of hiring against the test's recommendation went with 6% to 7% shorter job durations 1. Not a randomised experiment, and not an interview: the instrument is an online questionnaire scored into a green-yellow-red band, the firms chose to buy it, and completed tenure stands in for quality in a high-turnover service setting. The paper treats overrides as routine. It says the average override was worse. It does not say the test was right about any individual person, and nothing in it licenses removing the human from the decision.
Tell the two kinds of disagreement apart
Ask both people what they are looking at before anyone argues about who is right. If they are weighing the same observation differently, that is a real disagreement and the room should spend its time there. If one watched a candidate work and the other read a summary of a round they were not in, the room is holding two different sets of facts, and a show of hands will resolve it in favour of whoever sounds more certain.
When the evidence differs, the answer is never a vote. It is one more observation aimed at the specific claim in dispute, which usually costs twenty minutes instead of another full round: one question, one small task, one reference call with a named question attached. Write down in advance what result would change each person's mind, then go and get exactly that. A panel that cannot state what would change its mind has not been disagreeing about the candidate.
Structure is worth defending here for a reason beyond tidiness. A resume audit of 36,880 applications to 9,220 job advertisements for new US college graduates found callback gaps varying with the task content of the job: in management occupations, callbacks were 28 to 43 percent lower for Black men, Black women, White women and Hispanic men than for otherwise identical White men, and gaps were widest in roles combining high analytical and interpersonal demands with low routine content 2. It is an arXiv preprint under review, it measures callbacks and not offers, the range spans four applicant groups, and the task-content mechanism is a model the authors propose, which the experiment never manipulated. Read at that strength it still points somewhere: the measured gaps ran widest in the jobs the authors read as leaving an evaluator the most room. Requiring every position to cite an observation is one of the few cheap ways to narrow that room at the decision itself.
Which vetoes should not count?
Rule out one class in writing: a block resting on a guess about how the work was produced, or on a suspicion about who or what wrote it. It cannot be checked, it cannot be appealed, and it falls hardest on candidates whose written register sits furthest from the panel's. A concern about the work has to be stated as a claim about the work, in the round where it was observed.
Two measurements retire the confident version of this. Asked to tell GPT-3 output from human writing across stories, news articles and recipes, non-expert evaluators performed at random chance, and three quick training methods lifted accuracy only to about 55%, inconsistently across the three domains 3. That is 2021, short passages, and crowdworkers doing the judging in place of hiring managers. Newer models make unaided reading harder still, though the paper measures none of that.
The tools do not rescue the judgment. Seven widely used detectors run over 91 human-written TOEFL essays by non-native English speakers produced an average false-positive rate of 61.3%, with 97.8% of those essays flagged by at least one tool, while the same detectors were near-perfect on essays written by US eighth-graders 4. One language cohort, seven detector versions as they stood in 2023, and academic essays as the material. The number is not a constant; the direction of the error is the durable part, and it points at people writing in a second language. A block built on that reasoning follows a candidate's language background into the decision, and it will not look like that in the meeting. Whether it also creates legal exposure is a question for counsel. Stopping the move inside a hiring team is its own problem, worked through in how to stop managers rejecting candidates for sounding like AI.
Write the rule before you need it
Write it this week, while no offer is waiting, because the version drafted during a deadlocked loop gets written by whoever is losing. Three sentences on one page: who decides, who can block, what a position has to cite. Circulate it to everyone who interviews, and put the citation clause first, where it reads as the point of the rule.
A stalemate under this rule is informative rather than annoying. If both sides have cited real observations from rounds they sat in and still disagree, the loop failed to cover the claim that decides the hire, and the next step is a targeted round. If one side cannot cite anything, the rule has already answered and nobody has to say so unkindly. If both sides are citing the same moment and reading it differently, that is a calibration question about the bar, and it will return on the next candidate whether or not this one is hired.
Two adjacent decisions travel with this one. Who runs the meeting and whether that person votes is settled separately, in who chairs the debrief and whether the recruiter votes. And a panel that has never done this kind of work with an assistant will produce dissents about AI use that no decision rule can adjudicate, which is the harder problem in getting managers who do not use AI to judge AI-assisted work.
Common questions
Who should hold a veto?
Anyone on the panel, provided the veto has to cite an observation. Restricting it by title makes the rule about seniority, which is the failure the folklore is full of. Making it expensive to use, by requiring a specific claim and the round it came from, controls how often it gets used far better than restricting who holds it. A block that has to survive being said out loud, with the round attached, is one the room can act on.
Is a majority vote ever the right rule?
Rarely, because it converts panel composition into a decision and hands real power to whoever picked the panel. It also encourages people to state a position rather than their evidence, since only the position gets counted. Where a vote does help is as a diagnostic: if the room splits evenly after everyone has cited what they saw, the loop failed to test the claim that matters, and the fix is to go and get another observation.
What if two people cite the same moment and read it differently?
That is the good case and it deserves the meeting's time. Have each of them describe what they expected to see instead, which usually surfaces a difference in the bar rather than in the observation. If the bar is the disagreement, it is a calibration question for the whole team and it will recur on the next candidate, so write down which reading the team is adopting and why, before anyone leaves the room.
How do you stop a block being used as a proxy for something else?
Require the observation and require the round. A concern that cannot be attached to something the candidate said or did has nowhere to hide once the rule is in place, and the person raising it will either find the real observation or withdraw. The requirement also protects genuine dissent: with a citation attached, a junior interviewer's block carries the same weight as anyone else's, which is the point of writing it down.
Should the candidate be told the panel disagreed?
Telling a candidate the vote was split rarely helps in that form. What does help is specific feedback tied to the same evidence a citation requirement already forces the panel to produce. If the loop cannot name a reason precise enough to say out loud, that is worth knowing before the rejection goes out, because it usually means the disagreement was never about the work in the first place.
References
- 1. Discretion in Hiring nber.org Supports the claim that managers who more often hire against a structured signal see shorter completed job tenures (6% to 7% per standard deviation), and that introducing the test raised tenures by 0.23 log points.
- 2. Hiring Discrimination and the Task Content of Jobs: Evidence from a Large-Scale Resume Audit arxiv.org Supports the claim that measured callback gaps varied with the task content of the job and ran widest in high-analytical, high-interpersonal, low-routine roles, cited as a preprint with the callback-not-offer limit stated.
- 3. All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated Text aclanthology.org Supports the claim that untrained evaluators distinguished model-written from human-written text at random chance, and that brief training lifted accuracy only to about 55%.
- 4. GPT detectors are biased against non-native English writers pmc.ncbi.nlm.nih.gov Supports the claim that detector errors fall on second-language writers: an average false-positive rate of 61.3% across seven tools on 91 TOEFL essays, with 97.8% flagged by at least one.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.