Roles
An Autonomy Evaluation Operations Manager Owns the Evidence Behind a Release
An Autonomy Evaluation Operations Manager runs the program that produces release evidence: scenario coverage, test fleet scheduling, log triage throughput, and the gate a build has to clear before it carries passengers. The work is operational rather than research. Nuro is hiring a Senior Program Manager for Eval Operations and a Director of Engineering for its Eval Platform; May Mobility is hiring a Tech Lead for Performance Evaluation [1]. Hire someone who can state what was never tested and hold a launch on it.
The takeMost autonomy programs already have people who can build an evaluation. Very few have someone whose job is the throughput and the coverage of it, and that gap is where release dates quietly win arguments against evidence. Evaluation that lives inside the engineering org gets scheduled around the engineering org, so the scenarios that go untested are the expensive ones. Give this role its own headcount, its own fleet time, and the standing authority to hold a build. A gate that has never held anything is not a gate.
Where Olive fits
Open a role and see what the work shows
If you are building this evaluation capability yourself, the hard parts are the answer key and the evidence trail. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings, each anchored to a moment in the session rather than to a score.
Rank your shortlistYour Release Review Has Eleven Thousand Miles and No Answer for the Left Turn
Friday's release review opens on a slide with eleven thousand autonomous miles and a disengagement rate trending the right way. Someone asks how many unprotected left turns across oncoming traffic at dusk are inside that number. Nobody in the room knows. The miles are real, the trend is real, and the question is unanswerable, because miles were counted and scenarios were not.
That gap is the job. An Autonomy Evaluation Operations Manager runs the operation that produces release evidence rather than the research that designs metrics. The postings show the shape: Nuro lists a Senior Program Manager for Eval Operations alongside a Director of Engineering for its Eval Platform, and May Mobility lists a Tech Lead for Performance Evaluation and a Machine Learning Engineer for Autonomous Driving Performance Evaluation 1. Two of those build the platform. The others run the program on top of it.
Three traits separate a real candidate from someone who has attended release meetings. The first is thinking in coverage rather than volume. Ask what they would put on the slide instead of mileage, and listen for a structure: a scenario catalog with counts per bucket, the buckets that are empty, and which of the empty ones matter for the operational design domain you are actually shipping into. A candidate who reaches for a bigger mileage number has told you they will manage the metric instead of the risk.
The second is treating the evaluation pipeline as a production line. The real constraint in these programs is almost never analysis. It is that four hundred hours of drive log sit untriaged, the closed course is booked by the perception team for three weeks, and the regression suite takes eleven hours so nobody runs it before a merge. Ask about the last queue they unblocked and what the cycle time was before and after. The answer is either specific or it is nothing.
The third is a willingness to say the sentence nobody wants. Coverage reports are mostly bad news, and the tell is whether a candidate can describe a build they held and what it cost them. Someone who has never held anything either had no authority or never used it, and both are worth knowing before the offer. That authority is also what marks the boundary with the people watching vehicles in service, such as an autonomous fleet monitoring and response engineer, whose signal feeds the catalog but whose job starts after the gate.
Which Backgrounds Produce Someone Who Can Run an Eval Program?
The strongest feeders are people who have already run a verification program under a schedule that wanted to skip it. Aerospace flight test operations, silicon validation program management, rail signaling assurance and clinical trial operations all qualify. Every one of them has managed scarce test assets, a coverage matrix that nobody wanted to look at, and a stakeholder who needed the result two weeks before it existed.
Flight test operations converts fastest and is under-recruited. That person has spent a career deciding which test points get flown with the aircraft time available, has argued about whether a card can be closed on partial data, and knows exactly what it feels like when the answer is that the envelope was never cleared for the condition somebody just asked about. Swap aircraft for vehicles and test cards for scenarios and most of the craft transfers intact.
Silicon validation is the other high-yield pool, for a less obvious reason. Validation program managers live inside the same tension this role has: the design team owns the fix, the schedule owns the tape-out, and the validation matrix is the only artifact that says what remains unknown. They arrive already fluent in coverage as a governing document rather than a report.
The unexpected feeder is clinical trial operations. Site scheduling, protocol deviations, monitoring visits and a locked analysis plan map more closely onto scenario campaigns than anyone expects, and those managers carry a discipline that AV programs often lack: the analysis is specified before the data arrives, so a disappointing result cannot be reinterpreted afterward. Test operations leads already inside AV companies are the obvious internal pool and should be your first call, and the same is true of a strong autonomous vehicle operations specialist who has been running the fleet side of test campaigns.
Two profiles read well and often disappoint. A machine learning engineer who has built evaluation tooling may want to keep building it and treat scheduling, safety drivers and garage capacity as somebody else's problem, which leaves the operation unowned. And a generalist program manager with no safety-critical background tends to run the calendar competently while missing the thing that matters, which is that a green dashboard with an empty scenario bucket underneath it is worse than a red one.
Ask How They Got Good at Distrusting a Confident Metric
Ask directly how they use AI assistants in their own work, then push past the tooling answer. The useful version of this question is not whether they use a model to cluster drive logs or draft a scenario tag. Nearly everyone does now. It is whether they have been burned by one and changed their process because of it, and whether they can name the moment.
The good answers are concrete. Someone describes an autolabeler that tagged cut-in events with high confidence, went into the coverage count unaudited, and turned out to be wrong often enough that a bucket reported as covered was roughly half something else. What they did next is the part worth hiring: they held out a human-audited slice, measured the labeler against it, and published coverage with the labeler's own error rate attached rather than as a clean integer.
A second good answer describes using a model to triage log clips and then deliberately keeping a random sample outside the triage, because a filter that silently drops a failure class removes that class from the evidence and from the argument at the same time. This is the same habit as checking a claim against something outside the conversation, and it is the one that survives contact with a deadline.
What you are screening against is fluency. This vocabulary is easy to perform: scenario coverage, operational design domain, requeue latency, held-out audit. A candidate can say all of it having run three campaigns or having read one standard, and the interview transcript looks the same either way. Hand them your real coverage matrix with the identifying details stripped, give them ninety minutes, and ask which three buckets they would refuse to sign off on and why. The gap between the people who can do this and the people who can describe it shows up in about twenty minutes.
One more probe worth the time: ask what metric they retired. Every serious evaluation program has killed a number that was being managed rather than measured, and the story of retiring one tells you more about judgment than any answer about which metrics they would add. The same instinct shows up in adjacent safety-critical work such as hiring an autonomous aircraft flight operator, where the operator's log is the evidence and padding it is the failure mode.
Where Do You Find This Person, and What Actually Closes Them?
Start with the AV programs that already staff the function, because the title is new enough that most of the practitioners are inside a handful of companies. Nuro and May Mobility are both hiring into it now 1, which also means their peers have people doing the work under older titles: test operations lead, validation program manager, safety case manager. Search for the responsibility rather than the title, since the title is roughly two years old.
Outside AV, go where verification practitioners gather rather than where autonomy gets announced. SAE International's ground vehicle standards work, the IEEE Intelligent Transportation Systems Conference, the ASAM community around scenario description formats, and the practitioner community around the UL 4600 safety case standard are all real venues where this exact problem is argued about in public. Flight test engineering societies are the adjacent pool nobody in AV recruits from, which is the argument for doing it.
What closes them is authority, stated in writing. Every experienced verification person has a story about a program where the coverage report was advisory and the launch date was not. Name the gate. Name what a hold looks like procedurally, who can override it, and what that override has to be written down as. A role described as producing a weekly coverage dashboard will lose to a role described as owning a gate, even at the same money.
The second closer is resources with numbers attached. How many vehicles, how many hours a week, how much closed course time, how many triage headcount. Candidates who have run these programs know the operation is bounded by fleet hours and triage throughput, and a vague answer reads as a role that will be asked for evidence without being given the means to produce it. The third is honesty about the current state. Telling a strong candidate that log triage is nine days behind and the scenario catalog was last audited in March is not a weakness in the pitch. It is the job description, and the people you want are the ones who lean in at that sentence.
What Does an Autonomy Evaluation Operations Manager Cost, and Can It Run Remote?
No wage series covers this title, and no survey found for this piece prices it, so this stays qualitative on purpose. Price it internally. The role hires against your senior technical program manager band in autonomy, and moves toward the engineering band when the same person also owns the evaluation platform, which is the split visible in the two distinct Nuro postings 1.
Two adjustments are worth making before you set the number. Release-gate authority raises the band, because a person who can hold a launch is doing a different job than a person who reports coverage. And safety-critical verification experience is scarce in a way that AV program management is not, so a flight test or silicon validation candidate will have competing offers from outside your industry. The broader wage pressure is not in dispute: PwC's analysis of roughly one billion job advertisements found an average wage premium of sixty-two percent for roles demanding AI skills 2.
On location, the program half is genuinely remote-friendly and the operation half is not. Scenario catalog work, coverage analysis, triage queue management and release reporting all travel fine. What does not travel is the garage, the closed course, the safety driver briefings, and the habit of standing next to a vehicle when something strange comes back in a log. Most postings in this family sit at the operating sites for that reason, and the honest version of a hybrid offer names how many days are expected at the depot rather than calling the role remote and quietly requiring travel.
There is also a data constraint that decides more than people expect. Drive logs are large, often carry recorded imagery of public roads and the people on them, and frequently cannot leave a controlled environment, so remote review depends on tooling that runs inside your boundary. Scope that before writing the offer, because it determines which candidates can do the work from where they live.
One legal note, offered as a flag rather than as advice. In the United States, driverless testing and deployment approval runs through state-level permitting rather than a single federal license, the requirements differ by state, and they are still changing through 2026. The evidence this role produces is usually what an application or an inquiry ends up resting on. Check with counsel in the states where you plan to operate rather than reasoning from a summary.
Common questions
How do I become an Autonomy Evaluation Operations Manager?
Come from a verification operation and learn the autonomy vocabulary, rather than the reverse. Flight test operations, silicon validation program management, rail signaling assurance and clinical trial operations all produce the core skill, which is running a coverage matrix against scarce test assets and a schedule that wants to skip it. Then learn the domain specifics: scenario description formats, operational design domain definitions, the structure of a safety case, and enough log tooling to read a triage queue. Build one artifact you can show, a scenario catalog with honest empty buckets and a stated audit method, since that is what the interview is actually about.
How is this different from a model evaluations engineer?
A model evaluations engineer builds and improves the measurement: test rigs, metrics, datasets, the platform that runs them. An Autonomy Evaluation Operations Manager runs the program that uses it, which means scenario coverage decisions, test fleet and closed course scheduling, log triage throughput, and the release gate the evidence feeds. Nuro's two postings show the split directly, with a Director of Engineering for the Eval Platform sitting beside a Senior Program Manager for Eval Operations. Small programs combine them and usually get the platform, because the operational half is what slips when a deadline tightens.
Can an existing program manager take this on?
Sometimes, and the deciding factor is not seniority. A program manager who has run verification under schedule pressure in any safety-critical domain will usually convert. One who has only run feature delivery tends to manage the calendar well while missing the failure this role exists to prevent, which is a green dashboard with an empty scenario bucket underneath it. Test it before committing headcount: hand the candidate a real coverage matrix and ask which buckets they would refuse to sign off on. The answer separates the two profiles quickly.
What should this role report on instead of miles driven?
Coverage against a named scenario catalog, with the empty buckets listed rather than summarized. Useful companions are triage cycle time from log capture to classified event, the share of coverage that rests on automated labeling and that labeler's audited error rate, regression suite runtime and how often it actually runs before a merge, and the list of scenarios deferred this cycle with who deferred them. Mileage belongs in the report as context, never as the headline, because it answers a question nobody asked and hides the one that matters.
Is this role only relevant to self-driving cars?
It started there and is spreading to any autonomy program where a release decision needs structured evidence: delivery robots, agricultural and mining equipment, warehouse fleets, and uncrewed aviation. The operational pattern is identical, since all of them face scarce test assets, a scenario space too large to sample randomly, and a gate somebody has to defend. Job titles vary widely across those industries, so search by responsibility. Where the deployment sits under a regulator, expect the evidence requirements to be stricter and the role to sit closer to the safety case owner.
How new is this title, honestly?
New enough that a candidate's title tells you little about their scope. The function has existed inside autonomy programs for years under names like test operations lead or validation program manager, and it is only recently being staffed as its own operations organization with a budget, which is why postings such as Nuro's Eval Operations program manager and May Mobility's performance evaluation tech lead are worth reading as evidence of a category forming rather than a settled market. Ask what a candidate was allowed to block. That answer, not the title, tells you which band applies.
References
- 1. Nuro job board (Greenhouse API) ✓ boards-api.greenhouse.io Supports the claim that Nuro is hiring a Senior Program Manager, Eval Operations and a Director of Engineering, Eval Platform, and that May Mobility lists a Tech Lead, Performance Evaluation and a Machine Learning Engineer II for Autonomous Driving Performance Evaluation. Discovery-sweep evidence collected 2026-09-01.
- 2. PwC AI Jobs Barometer 2026 pwc.com Supports the claim that roles demanding AI skills carry an average wage premium of 62 percent, measured across roughly one billion job advertisements.
2 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.