Roles
What Separates A Real AI Evals Engineer From Someone Who Has Read The Blog Posts
Give candidates the two incidents. Hand over the eval suite that was green when both shipped, plus twenty logged failures from users, and ask for a written diagnosis in ninety minutes. A real AI evals engineer finds why the suite passed: a test set that leaked into the prompt, an LLM judge that rewards length, a metric averaged across cases so a subgroup collapse disappears. Weak candidates rewrite prompts. Strong ones rewrite the measurement, then show the regression firing.
The takeThe instinct after a bad model update is to add tests. That is the wrong repair. Both of those incidents shipped past a green suite, which means the suite was the thing that failed, and hiring someone to add more cases to it buys a longer green bar. The job is measurement design: deciding what counts as wrong, proving the judge agrees with a human on the cases that matter, and keeping the held-out set held out. Hire for that skepticism or keep shipping quietly broken.
Where Olive fits
Open a role and see what the work shows
If you are building this exercise yourself, the hard parts are the answer key and the evidence trail. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings, each anchored to a moment in the session rather than to a score.
Rank your shortlistWhat Did Your Eval Suite Miss When The Model Changed?
Both incidents look the same from the inside. A provider ships a point release, the nightly suite goes green, and support tickets arrive four days later about a behavior nobody wrote a test for. The suite was not lying. It was answering a question that stopped being the right one, and nobody on the team owned that question.
The cause is usually one of three. The test cases sat in the model's training data or in the prompt itself, so the score measured recall rather than capability. The metric was an average over a mixed set, so one slice regressed hard while the mean barely moved. Or the grader was another model, and its preferences shifted with the same update that broke the product. None of those are defects in application code, which is why the people who own the serving path did not catch them either. If your gap is really deployment, rollback and traffic shadowing, that is a different hire with a different test.
The habit worth screening for is a refusal to treat a passing number as evidence. Microsoft's 2026 Work Trend Index found half of surveyed AI users naming quality control of AI output as a skill growing in importance, and 86 percent saying they treat model output as a starting point rather than a finished answer 2. An AI evals engineer industrializes that instinct: writes down what wrong looks like before running anything, then builds the smallest apparatus that can prove it happened.
Which Backgrounds Produce An AI Evals Engineer?
No degree program produces this title yet, so the field is stocked with people who arrived sideways. One 2026 career guide for the title lists the common feeders as ML engineers tired of shipping models nobody measured, data engineers who already own pipelines and statistical tooling, and product engineers who wrote their team's first regression suite by hand 1. The unexpected backgrounds are more interesting, and often stronger.
Psychometricians and survey methodologists have spent careers on the exact problem: writing items that discriminate, measuring whether two raters agree, and knowing that a scale means nothing until somebody checks it against an outside criterion. Clinical trial statisticians bring pre-registration as a reflex, which is the single most useful habit in this job. Test engineers from safety-critical software (avionics, medical devices, payments) arrive already believing that coverage is a claim requiring proof. Linguists and annotation leads know how quickly a label definition rots when six people apply it.
What all of them share is comfort with the idea that the measurement is a product with its own defects. The people who struggle are strong application engineers who treat evals as a CI chore: they will build you a fast, tidy, thoroughly green rig around a question nobody validated. Where the datasets themselves are the bottleneck, the adjacent skill is curating and versioning the corpus, and some teams need that person first.
On the tooling side, expect Python for data work, familiarity with at least one eval framework such as Inspect, Promptfoo, Braintrust or LangSmith, enough statistics to say whether a five-point difference on 200 cases means anything, and a documented opinion about when an LLM judge is trustworthy. That bar overlaps the competencies the same career guide reports teams screening for 1, and framework names are the cheapest thing on it to fake.
Test Candidates Against A Judge That Is Already Wrong
Skip the take-home asking for an eval rig from scratch. Build one exercise out of your own repository instead: a small suite that passes, one seeded defect in the measurement, and thirty real user complaints the suite says nothing about. Ninety minutes, an AI assistant allowed and expected, one written page at the end.
The seeded defect should be a measurement defect, never a code bug. Good ones: an LLM judge prompt that asks for the more helpful answer and therefore rewards length; a rubric where two of five criteria are near-duplicates, double-counting one behavior; a held-out set assembled by sampling the same source the few-shot examples came from; an accuracy number averaged over cases where 80 percent are trivial.
The tells separate quickly. Real evals engineers ask what decision the number feeds before touching anything, because a release gate and a research dashboard need different error tolerances. They hand-label a sample themselves rather than trusting the existing labels. They look at the cases where the judge and a human disagree, not the aggregate agreement rate, and they can say out loud what agreement level would make them stop using the judge. They refuse to declare the fix verified against the same set they tuned on. Performed expertise names three frameworks in the first two minutes, produces a larger suite, and never questions the label.
The AI-use dimension is worth watching directly rather than asking about. Assistants are genuinely good at generating eval scaffolding and genuinely bad at noticing that the scaffolding measures the wrong thing, so the candidate who has practiced this will accept the generated structure, then immediately attack one assumption inside it. Ask them to show a case from their own work where a model handed them a plausible eval and they caught the flaw. Practitioners have that story ready with details; people repeating conference talks have a general answer about staying critical.
Where Do AI Evals Engineers Leave Public Artifacts?
Look for artifacts rather than titles. People doing this work leave behind eval repositories, model cards with a real failure taxonomy, annotation guidelines, and post-incident writeups about a metric that lied. Frontier labs and applied AI companies are the obvious feeders, and the firms selling evaluation tooling are the least obvious and most concentrated pool 1.
The same 2026 career guide sorts the named hiring pools as of mid-2026 into four groups: frontier labs including Anthropic, OpenAI, Google DeepMind, Mistral and xAI; applied AI companies including Cursor, Harvey, Sierra, Perplexity, Decagon and Cognition; enterprise deployers including Stripe, Shopify, Databricks, Atlassian and HubSpot; and evaluation infrastructure firms including Braintrust, LangChain and Arize 1. That is one guide's list rather than a survey of the market, so read it as a place to start calling rather than as a census. That last group matters most for sourcing, because their engineers spend all day inside other companies' broken measurement and have seen more failure modes than any single product team can generate.
Adjacent titles that convert well: trust and safety measurement, ML platform quality, annotation operations lead, applied research engineer, and the QA architect at a regulated software company. In regulated sectors the crossover runs the other way too, since the person who can defend an evaluation to an auditor is close kin to the analyst who writes the policy it has to satisfy.
Practically, the highest-yield sourcing move is reading the eval sections of open model releases and benchmark papers and contacting whoever wrote the methodology appendix. That paragraph is short, unglamorous, and written by exactly one person on the team.
What Does An AI Evals Engineer Cost, And Where Do They Sit?
No government wage series carries this title yet, so treat every figure as a posting-derived band rather than a survey. As of mid-2026, one market guide for the role, and only one, puts mid-level total compensation at $230,000 to $320,000 at applied AI companies and $280,000 to $420,000 at frontier labs, with senior bands at $320,000 to $460,000 and $420,000 to $620,000 respectively 1. Nothing independent corroborates those bands yet.
Those numbers describe a narrow, well-capitalized slice of the market, so anchor them against something broader before writing an offer. Levels.fyi reports an average total compensation of $247,000 for ML and AI software engineers in the United States 3. Most companies hiring their first evals engineer are pricing against that band plus a scarcity premium, not against frontier-lab equity. The same guide notes that staff-level evals engineers often out-earn equivalent product engineers because supply is thin 1, which is a warning about your counteroffer more than about your opening number.
Money is rarely what loses these candidates. What closes them is authority: access to production traffic and real failures, a budget line for human annotation, ownership of the held-out set, and a written rule that a red eval blocks a release. What kills the offer is a reporting line into the team whose bonus depends on shipping, or a job that turns out to be running other people's test requests. Ask any serious candidate what would make them quit in year one, and listen for a version of that.
Remote is the norm and defensible: the work is asynchronous, artifact-heavy, and reviewed through written diagnoses rather than meetings. The exception is data gravity. Where evaluation sets contain patient records, defense material or unreleased model weights, the work happens on-premise or inside a locked enclave, and that constraint belongs in the first recruiter conversation rather than the final round. In those environments the eval set itself is the sensitive asset, which changes who is allowed to look at failures and slows every iteration loop by a factor nobody budgets for.
Common questions
How do I become an AI evals engineer?
Pick a model-backed product you use, write down twenty ways it fails, and turn those into a small eval set with a labeled answer key you made yourself. Then measure whether an LLM judge agrees with your labels, and publish the disagreement rate. That artifact does more than a certificate. From there the fastest routes in are ML engineering, data engineering, annotation operations, or any measurement-heavy science background. Hiring managers screen for whether you have caught a measurement lying, so keep a written example of one.
Do we need a dedicated AI evals engineer, or can the ML team cover it?
If model or prompt changes have shipped past a green suite more than once, the coverage is not working. The structural problem is that the team building the feature also grades it, and graders who want to ship set forgiving thresholds without noticing. A part-time owner is fine below roughly ten evaluated behaviors. Past that, someone needs the held-out set, the labeling budget and the authority to block a release, and those three responsibilities do not split well across people.
What should an AI evals engineer produce in their first 90 days?
A failure taxonomy drawn from real user complaints, a held-out set nobody has tuned against, a documented agreement rate between any LLM judge and human labels, and one release blocked or approved on that evidence. Not a dashboard. Dashboards arrive later and are the easiest deliverable to produce without answering whether the underlying measurement is sound.
How is this different from a QA engineer?
Traditional QA asserts an expected output and fails on mismatch. Model behavior has no single correct string, so the evals engineer has to define what counts as good, prove that definition is applied consistently, and re-prove it whenever the model changes underneath. The overlap is real: test engineers from safety-critical software convert well, because they already treat coverage claims as things requiring evidence rather than assertion.
Should the exercise let candidates use an AI assistant?
Yes, and watch what they do with it. Assistants generate eval scaffolding quickly and miss that the scaffolding measures the wrong thing, so an assistant-enabled exercise shows you exactly the judgment you are hiring for: which generated assumption did the candidate refuse to accept, and how did they check it. Banning the assistant tests a working style nobody on your team will use after week one.
References
- 1. AI Evals Engineer Career Guide 2026 ✓ jobsbyculture.com Total compensation bands by level for applied AI companies and frontier labs; named hiring pools (frontier labs, applied AI startups, enterprise deployers, evaluation tooling firms); the five screened competencies including eval frameworks and LLM-judge reliability; staff-level evals engineers out-earning equivalent product engineers.
- 2. Agents, human agency, and the opportunity for every organization ✓ microsoft.com 50 percent of surveyed AI users name quality control of AI output as a human skill growing in importance; 86 percent treat AI output as a starting point rather than a final answer.
- 3. ML / AI Software Engineer Salary ✓ levels.fyi Average total compensation of $247,000 for ML / AI software engineers in the United States, used as the broader-market anchor beside the role-specific bands.
3 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.