Roles

Hire A Growth Experimentation Lead Who Can Kill A Winning Test

The person who decides is a growth experimentation lead: an owner for the whole pipeline rather than a test runner. When AI generates a thousand variants a week, most apparent wins are noise, and the scarce skill is calling which results are real and which ones should change strategy. Hire someone who has killed a test that looked like a winner, who can explain false discovery rate without slides, and who owns the guardrails on automated bid and budget decisions.

The takeThe instinct is to hire whoever can run the most tests, and that is now the cheap half of the job. Generation capacity went to roughly free; decision capacity did not move at all. My bet is that the best version of this hire comes from someone who has been burned by a false positive that shipped, because scar tissue is the only reliable teacher of restraint at volume. Hire for the willingness to say a result is not real when the dashboard is green and the team is excited.

Where Olive fits

Open a role and see what the work shows

The same six dimensions describe what capable AI work looks like in a role like this: framing before generating, demanding a source for the number that matters, keeping the judgment that should not be delegated, and testing a claim against something outside the conversation. Olive reads those from a real working session rather than from a self-assessment.

Rank your shortlist

What Breaks When AI Runs A Thousand Growth Experiments?

On a Monday standup, forty creatives cleared the confidence threshold last week and one of them beat control by eleven percent. Someone in the room says ship it to the whole account. The person you are hiring says no, and then has to explain why, out loud, to a table that can see the green number on the wall.

The reason is arithmetic rather than taste. Forty tests ran in that family, and at a plain threshold two of them clear on noise alone with nobody able to say which two. So the lead asks for the correction and the peeking policy before reading any single result. That is what volume did: it broke the meaning of the word significant, and a team that has not corrected for it spends the quarter chasing variants that were never real.

Which is why the job stops being generation and becomes adjudication. The lead owns four decisions no model makes: which questions are worth spending traffic on, when a result gets acted on, when a winning variant is promoted from a lucky slice to actual strategy, and where the automated bid and budget systems may move money without a human in the loop. Everything upstream of those four can be handed to a machine, and mostly should be.

What makes the no survivable on a Monday is a sentence written the Tuesday before. Ahead of the batch, the lead had written down what result would ship, what would kill, and what would earn a rerun, and the eleven-percent creative landed in the third box because it had run four days against a rule that asked for ten. Ask a candidate to show you the case where their own rule went against what they wanted. A real one flinches at the word significant used alone and asks how many tests were in the family; a performed one recites the number of experiments the last team shipped.

Two more things come out of that same standup. Ask what happens when the creative that won in week one loses in week four, and listen for whether novelty decay is a familiar problem or a new idea. Then ask about the variant that outperforms and still does not go out, because machine-generated creative reaches into territory a human copywriter would never take and somebody has to make a call with no metric behind it. That call is the reason the role exists.

The stakes justify the seniority. McKinsey concentrated roughly 75% of generative AI's estimated $2.6 trillion to $4.4 trillion in annual value potential across four functions, marketing and sales among them 1, and LinkedIn's Work Change Report found 51% of companies adopting generative AI reporting revenue increases of 10% or more 3. Those are estimates about a function, not a promise about your funnel, and the whole point of this hire is having someone who can tell the difference.

The Clinical Trial Coordinator Is A Better Bet Than The Brand Strategist

Four backgrounds produce this person reliably: growth or lifecycle marketers who owned a real A/B program, performance media buyers who ran automated bidding at scale, product analysts who built the experimentation platform's guardrails, and quantitative people from outside marketing who learned the funnel. Each arrives strong on one half of the job and teachable on the other within a quarter.

The growth marketer brings funnel intuition and knows which questions are worth traffic, which is the part that cannot be taught quickly. What they usually lack is statistical hygiene at volume: sequential testing, multiple-comparison correction, and the discipline to fix a stopping rule before launch. The performance media buyer arrives with the opposite kit. Someone who has run a large automated bidding account already knows what it feels like when an optimizer finds a loophole in the objective you gave it, and that instinct transfers directly to agentic budget systems.

The unexpected backgrounds are worth naming because resume screens filter them out. Clinical trial coordinators and biostatisticians have spent careers on pre-registration, interim analysis and stopping rules, which is precisely the discipline that machine-scale marketing lacks. Quality engineers from manufacturing think in process control charts and know the difference between a shift and a wobble. Sports analytics people have argued about small-sample effects in public for years. Anyone who has run a recommendation or ranking system has already lived through the counterfactual problem that makes attribution hard.

What transfers less well than people expect: brand strategists with no measurement practice, and pure data scientists with no exposure to the commercial argument. The second failure is the more common one. A person who can compute the right answer but cannot hold the line in a room with a founder who liked the losing creative will not survive the role. If the gap you actually have is proving that any of this paid for itself, that is a different and complementary hire, closer to an AI value and ROI analyst than to this one.

Screen The Candidate On A Test They Called Wrong

Ask for a specific experiment they shipped that turned out to be a false positive, and what changed in their process afterward. Everyone has one. Candidates who claim otherwise either have not run enough volume to meet the problem or are not telling you about it, and both answers are useful. The follow-up that separates people is whether the fix was a rule or a resolution to be more careful.

The strongest candidates got good at this by pointing AI at their own work and watching where it failed them. Ask what they actually built. The answers that land are concrete: a generation loop that produced two hundred ad variants and a classifier that killed the eighty which were off-brand before a human ever saw them; a script that reads campaign output and drafts the weekly readout; an agent that proposes budget shifts and requires a signature above a threshold. Then ask what it got wrong and how they found out.

Listen for whether they check the model. The useful answer sounds like: the assistant drafts the analysis, and the first move is recomputing one number by hand against the raw export, because a plausible query over a misjoined table gives a plausible wrong lift. A candidate who describes an assistant as reliably right about causal claims has not audited enough of them. This is the same habit an agent quality analyst practices on model output, and the two roles trade candidates more often than their titles suggest.

A working screen takes an hour, not a take-home week. Hand over a real, redacted month of experiment results with a known false positive buried in it, the kind of month that produces an eleven-percent winner on day four, plus an assistant, and ask for three recommendations ranked by confidence. Read for whether they asked how many tests were in the family before they interpreted any single one, whether they said out loud which of their numbers are soft, and whether they refused to recommend something on insufficient evidence. A candidate who declines to call a result is demonstrating the skill, not failing the exercise.

One thing not to screen for: whether the application material was written with AI. It cannot be determined reliably, and it has no bearing on whether the person can hold a decision rule under pressure, which is the only question that matters here.

Where Do AI Growth Experimentation Leads Actually Work?

They are already employed, usually inside a company where paid acquisition is a large line item, and they are not scanning job boards. Find them by the work rather than the title. The people who publish teardowns of their own experiment programs, who argue about attribution and holdout design in public, and who present at growth and analytics meetups are the visible slice; the rest come through referral from anyone who has run a large media account.

Adjacent titles that already contain most of the skill are head of growth, growth product manager, performance marketing lead, marketing analytics manager, and experimentation platform owner. AI-native framings of the same work are appearing as named categories, including agentic operator, AI SEO lead and AI marketing operations roles 2. Search the responsibilities, because the title is new enough that many qualified people hold it without those words anywhere on their profile.

Feeder companies are the ones whose margin depends on acquisition efficiency: consumer subscription businesses, marketplaces, direct-to-consumer brands past their first scaling year, and any product with a self-serve funnel large enough to test on. Ecommerce operators who lived through the measurement disruption of the past few years have a particular kind of skepticism that is hard to teach.

On location, this is remote-friendly work and mostly done that way, since the artifacts are dashboards, creative assets, briefs and a decision log. Two things pull it toward hybrid. The first is the creative review itself, which goes faster in a room and which is where the brand judgment gets exercised. The second is regulated categories. If you sell financial products, health products or anything with mandatory disclosure, the approval loop has legal in it and the practical answer is quarterly presence at minimum. On-premise in the literal sense is rare here; where it appears, it is a data residency constraint on customer records rather than anything about the marketing work, and it is worth naming in the posting because it narrows the pipeline sharply.

What Does This Growth Lead Cost, And What Kills The Offer?

No published salary series covers this title yet, so any point estimate for a growth experimentation lead is a guess dressed as data, and this piece will not add another. Build the band from two comparables you already pay: your senior growth or performance marketing lead, and your senior product analyst or data lead. The right offer sits at or above the higher of the two, because that is who you are competing against for the person.

The pricing mistake that stalls searches is treating this as a marketing manager requisition. The candidates worth hiring are being recruited by companies that price the role against analytics and product bands, and a marketing-band offer reads to them as a signal about how much authority the job actually carries. If your compensation structure cannot reach, say so early rather than discovering it at the offer stage.

What they care about, in the order it comes up: authority to kill work, including the eleven-percent creative the founder liked; access to the raw data rather than a curated dashboard; a budget large enough that decisions have consequences; and a mandate that names quality alongside volume. A candidate who asks whether the experiment target is a number or a direction is asking the right question, and answering that it is a number set before anyone looked at the data will lose the good ones.

Three things kill the offer. Making the role a reporting function, so the first ninety days go to building a dashboard for someone else to act on. Keeping veto power over creative somewhere else, which turns brand risk into a problem the lead can see and cannot stop. And an unbounded always-on relationship with automated spend, where every optimizer anomaly is a page at any hour. Write down before the offer goes out which decisions are theirs alone and which need a second signature. Roles built at this kind of seam, like an AI delivery quality reviewer, fail in exactly the same way when the authority is left implicit.

See the benchmarks

Common questions

How do I become a growth experimentation lead for AI-run experiments?

Start from whichever half you have. From marketing: learn sequential testing, multiple-comparison correction and holdout design well enough to argue about them, and start pre-registering a decision rule for every test you run. From analytics or research: take ownership of a real acquisition budget so the decisions cost something. Then build the loop yourself. Generate variants with a model, ship them, and instrument the point where a false positive would get through. The portfolio piece that lands in an interview is one specific test you killed that everyone else wanted to ship, and the rule that made you kill it.

How many ad variants should a team test with AI?

Fewer than the tool can produce, and the ceiling comes from traffic rather than from generation capacity. Each additional variant splits the same audience, so past a certain point every arm is underpowered and the winner is mostly luck. Decide the number by working backward from the smallest lift worth acting on and the traffic available in the test window. If that math yields four variants, generating four hundred does not help, though generating four hundred and selecting four on human judgment sometimes does.

Does this replace the growth marketer or the marketing analyst?

Neither, in most teams. It changes what both spend time on. The marketer stops producing variants one at a time and starts writing briefs, decision rules and brand guardrails that a generation pipeline runs against. The analyst stops assembling readouts by hand and starts owning the correctness of an automated pipeline that produces hundreds of them. On a small team one person does both jobs, and that person is the growth experimentation lead. On a larger team the lead sets the rules and the two functions execute inside them.

How should a growth team be structured when AI runs the experiments?

One accountable owner for the pipeline, with generation, launch and reporting automated under that owner, and a named second signature for the decisions that move real money. The common failure is splitting generation from adjudication across two teams, which produces volume nobody trusts. Keep the guardrails on automated bid and budget systems in the same person's scope as the creative judgment, since both are answers to the same question about where a machine may act alone.

What should the first 90 days look like?

Trust before volume. Weeks one to four: audit the last quarter of experiments and count how many decisions actually changed as a result, which is usually a smaller number than anyone expects. Weeks five to eight: install a stated decision rule, a correction for multiple comparisons, and a ceiling on what automated systems may move without a signature. Weeks nine to twelve: run the pipeline at full volume against the new rules and publish a readout that names which results are soft. A generation pipeline is an output of this work, not the first deliverable.

Can an agency do this instead of a full-time hire?

An agency can run the generation and the media buying well, and often better than an in-house team at the same cost. What an outside partner will not do is kill a winning test on brand grounds, or argue with a founder about a creative direction that performs. Those decisions require someone with a stake in the company's next three years. Buy the execution outside if that suits the budget; keep adjudication in the building.

References

  1. 1. The economic potential of generative AI: the next productivity frontier McKinsey Global Institute, 2023. mckinsey.com About 75% of generative AI's estimated $2.6 trillion to $4.4 trillion in annual value potential concentrates in four functions, marketing and sales among them.
  2. 2. AI-Native Growth and Marketing Jobs 2026 Gallium, 2026. gallium.ai Names AI-native growth and marketing categories including agentic operator, AI SEO lead and AI marketing operations roles.
  3. 3. Work Change Report LinkedIn Economic Graph, 2025. economicgraph.linkedin.com 51% of companies adopting generative AI reported revenue increases of 10% or more.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.