Roles

An Agent Quality Analyst Is Hired to Distrust Your Eval Scores

An Agent Quality Analyst reads what your agent actually did, in transcripts and traces, and decides whether it was correct rather than whether it looked correct. The daily work is classifying failures into named modes, maintaining a golden set of hard cases, keeping rubrics honest, and signing off that production behavior matches what was promised. Hire someone who builds a failure taxonomy from 300 conversations, not someone who reports an average score.

The takeThe reason your agents pass their own checks is that the people who wrote the prompt also wrote the rubric, and a rubric written by the author grades intent instead of outcome. Fixing that is an organizational move before it is a technical one: the person who decides whether a run was good must not report to the person whose release it blocks. That is why this role is worth a headcount rather than a rotation. Give it to a careful reader with independence, and give that reader the authority to hold a launch.

Where Olive fits

Open a role and see what the work shows

If you are building this review capability yourself, the hard parts are the answer key and the evidence trail. Olive ships twelve authored cases per occupation and returns six separately-evidenced findings, each anchored to a moment in the session rather than to a score.

Rank your shortlist

Your Refund Agent Scored 0.94 and Told Eleven Customers a Lie

Last Thursday the refund agent scored 0.94 across the eval suite and closed twelve hundred conversations. Eleven of them ended with a customer being told, politely and with total confidence, that a refund had been issued when the payments API had rejected it. The eval never saw it. The rubric asked whether the reply was helpful and on policy, and the reply was both. It was also false.

That gap is the entire job. An Agent Quality Analyst reads what the agent did, in transcripts and traces, and decides whether it was right instead of whether it read well. Aquent names the title as a distinct hire in its guide to staffing agentic AI orchestration, sitting alongside the engineering roles rather than inside them 1. Three traits separate a real one from a person filling in a scoring sheet.

The first is the instinct to classify rather than count. Ask what a candidate would do with three hundred flagged conversations. A weak answer sorts them by score and reports a distribution. A strong answer builds a taxonomy: named modes such as asserted an action it never performed, answered an adjacent question, used stale tool data, or dropped the escalation path, each with a count and three pinned example transcripts. A distribution gives you a number nobody can act on. A taxonomy gives an engineer a bug list by Monday.

The second is stubbornness about the golden set. Press on how they would build one, and listen for whether the cases come from real production failures or from imagination. The tell is what they do when the agent starts passing everything: a real analyst treats a perfect score as evidence the set has gone stale and goes hunting for new hard cases, because a suite that never fails has stopped measuring.

The third is the willingness to say the agent behaved correctly and the policy is wrong. Plenty of flagged conversations are not model failures at all. They are a refund rule that contradicts the returns page, or a script that was never written for the case customers keep bringing. An analyst who cannot separate those two will file model bugs forever and fix nothing.

The tell that runs through all three is a question about ownership. A real candidate asks who fixes what they find, and keeps asking until you either name the person or admit there is no path from a finding to a change. That instinct also marks the boundary with the AgentOps engineer, who owns the tracing and the deploy gates that the analyst's findings flow into.

Which Backgrounds Produce an Agent Quality Analyst Who Can Read a Trace?

The strongest feeders are people who already graded human work against a written standard: contact center QA analysts, claims and underwriting reviewers, copy editors and fact checkers, clinical documentation specialists. Every one of them has spent years deciding whether an output was defensible, has argued with someone about a borderline case, and has kept a rubric alive through disagreement. That is most of the craft.

Contact center QA converts fastest, and gets overlooked because the title sounds junior. Someone who has scored two hundred calls a month against a quality form knows what calibration means, knows that two reviewers must agree before a score is worth anything, and has learned the hard way that a form which rewards politeness will produce polite agents who solve nothing. Point that person at transcripts and the transfer is nearly immediate.

The unexpected feeders are better than the obvious ones. Teachers who graded essays against a rubric bring inter-rater discipline. Localization QA specialists have spent careers on outputs that are grammatical and wrong, which is exactly the failure surface here. Paralegals and benefits caseworkers bring the habit of checking an assertion against the source document rather than accepting a fluent summary of it, and that habit sits close to the one described in hiring a benefits eligibility caseworker.

What none of them arrive with is the trace. Reading a transcript tells you what the customer saw. Reading a trace tells you why, and the difference between a hallucinated fact and a tool that returned an empty body with a 200 status is invisible without one. This is teachable in weeks and worth budgeting for: a span view, tool inputs and outputs, model and prompt versions. Screen for whether a candidate is curious about the layer under the words, not for whether they have already seen your tooling.

Two profiles read well and often disappoint. Traditional software QA engineers sometimes want a deterministic pass or fail and stall when the same input yields two different acceptable answers. And prompt-focused candidates tend to rewrite the prompt in response to any failure, which quietly makes them the author of the thing they are supposed to be grading. The demand pressure behind all of this is not in dispute: half of workers now name quality control of AI output as an increasingly important skill 2, and forty-seven percent report spending more time managing and directing AI than doing the work themselves 3.

Ask How the Analyst Got Good at Catching a Confident Wrong Answer

Ask directly how they got good at this, and listen for practice rather than coursework. The answers worth hearing describe a specific moment: an assistant produced something plausible, they believed it, it was wrong, and they changed how they work because of it. They can name the claim, name how it fell apart, and name the check they now run every single time.

Good answers share a shape. Someone describes running the same prompt five times to see the spread before trusting any one output. Someone else describes asking for the source of the one claim the whole conclusion rests on, then opening it. A third keeps a running file of cases where an assistant went sideways, which is the closest thing this discipline has to a lab notebook and which usually becomes their first golden set.

The skill underneath all of it is checking an assertion against something outside the conversation. A model says the policy allows a refund after ninety days; the analyst opens the policy. A model summarizes ten transcripts as showing a billing problem; the analyst reads three of them and finds two were about shipping. This habit describes badly in an interview, because talking about verification is easy and performing it under time pressure is not.

Press on rubric writing, which is where the real seniority shows. Ask them to write a two-line pass criterion for a support answer that must be accurate, in scope, and honest about uncertainty. Watch whether they notice that accurate and honest about uncertainty can conflict, and how they resolve it. Watch whether they define what a failure looks like rather than only what success looks like. Rubrics that only describe good outputs produce reviewers who cannot explain a rejection.

One warning about format. A conversation about quality rewards vocabulary, and a candidate saying golden dataset, inter-rater reliability, failure taxonomy may have run three review programs or read one article. The transcript looks identical either way. Hand them twenty real transcripts from your own agent, including four you already know are bad, and ninety minutes. The people who can do this reveal themselves fast, and so do the people who cannot.

Find Agent Quality Analysts Inside Your Own Support Organization First

Look inside before you post. The person who already reviews your human agents' conversations for quality is sitting in your support or operations organization, already knows your policies, already knows which customer situations are genuinely hard, and is usually delighted to be asked. Trust and safety review teams and localization QA teams are the second internal pool, for the same reason: their whole job is judging outputs against a standard.

Outside, go where transcripts get discussed rather than where launches get announced. Conversation design and contact center quality communities are real and long-standing. Data annotation and evaluation vendors employ people who have graded model outputs at volume, and the good ones have opinions about rubric drift that you will want to hear. Adjacent titles worth approaching directly: quality assurance analyst, conversation designer, content reviewer, clinical documentation specialist, claims reviewer.

What they care about, and this decides the offer more than money does, is whether findings go anywhere. Ask any experienced quality person about their last job and you will hear about a report nobody read. The offer dies when the role is described as producing a weekly dashboard. It dies again when they learn the engineering team can dismiss a finding with no written reason, or that the launch date was fixed before the review began.

Three things close the hire. Name the standing meeting where findings become tickets, and name who runs it. Give the role explicit authority to hold a release, even if that authority is exercised twice a year. And be honest about the volume of reading, because this work is genuinely repetitive and the candidates who last are the ones who knew that going in. Where the value of the whole agent program is being questioned, this person's taxonomy is often the only evidence anyone has, which is why they end up working closely with the AI value and ROI analyst.

What Does an Agent Quality Analyst Cost, and Should the Role Be Remote?

No wage series covers this title, and no survey found for this piece prices it, so this paragraph stays qualitative on purpose. Any single dollar figure quoted for the role today is a guess dressed as a benchmark. Price it internally instead: start from your senior quality analyst band, then decide whether the role carries release authority, because a person who can hold a launch is doing a different job than a person who files reports.

Where a candidate can also build and maintain the eval tooling, they will be priced against engineering, and you should expect to compete there.

One honest caution about the market. This title is new enough that levels are inconsistent between companies, so a candidate's current title tells you very little about their scope. Ask what they were allowed to block. That answer, not the title, tells you which band applies.

On location, the reading half of the work is fully remote and always has been. What resists remote is calibration. Two reviewers who never sit together drift apart within a quarter, and drifting reviewers produce a quality signal that moves for reasons nobody can name. Teams that run this well hold a recurring calibration session where several people grade the same conversations independently and then argue about the disagreements, and they treat that hour as the load bearing part of the process rather than overhead.

On-premise requirements show up where the transcripts themselves are the sensitive material: health records, financial detail, anything under a data residency rule. In those environments the constraint is rarely the person's desk and almost always the review tooling, which has to run inside your boundary. Scope that before you write the offer, because it changes which candidates can do the job from where they live, and it often means the analyst works alongside a workflow automation specialist to get a compliant review queue built at all.

One legal note, offered as a flag rather than as advice. Where an agent takes or materially shapes a decision about a person, several jurisdictions now impose notice, explanation and record keeping duties, and those rules differ by jurisdiction and are still changing through 2026. The review records an Agent Quality Analyst produces are frequently the only evidence trail that exists when someone asks what the system did and why. Check with counsel in your jurisdiction rather than reasoning from a summary.

See what gets scored

Common questions

How do I become an Agent Quality Analyst?

Start from any job where you graded work against a written standard: contact center QA, claims review, editing, clinical documentation, teaching. Then do the thing the role is actually made of. Take one hundred real conversations from any AI product you can access, mark the failures, and name each failure mode in language an engineer could act on. Build a small set of hard cases and rerun them monthly to see what changed. Learn to read a trace so you can tell a fabricated fact from a tool that failed quietly. Publish the taxonomy. It does more in a hiring conversation than a certificate.

Can our existing QA team test AI features instead?

Partly, and trying is reasonable. Software QA engineers bring test discipline and often adapt well. The strain is that the same input can produce two different acceptable answers, so pass or fail assertions break down and a single reproduction is not a bug report. What has to be added is grading under ambiguity: a rubric, two reviewers, and a calibration habit. Hire dedicated when the agent talks to customers without review, when the volume of conversations exceeds what anyone can read part time, or when nobody can currently say which failure mode is most common.

What is the difference between an Agent Quality Analyst and an AgentOps engineer?

The analyst judges whether the behavior was correct. The AgentOps engineer makes sure the behavior is visible and the fix ships safely. In practice the analyst reads transcripts and traces, classifies failures, maintains the golden set and writes the rubrics; the engineer owns tracing, eval infrastructure, deploy gates and rollback. Small teams put both in one person and usually get the engineering half, because the reading is what gets dropped when a sprint is tight. That is the argument for separating them once agent volume is real.

How do we QA a chatbot before customers see it?

Assemble a set of cases before you write a rubric, drawn from your real support history rather than invented, and weighted toward the situations that already go badly with humans. Have two people grade the same runs independently and compare, because the disagreements are where your rubric is unclear. Test the failure paths deliberately: missing data, an ambiguous request, a customer who is wrong about their own account, a tool that returns an error. Then keep the set and rerun it after every prompt, model or tool change, since a fix in one place moves behavior in another.

How many conversations should an Agent Quality Analyst review?

There is no standard number, and any figure quoted as one is invented. Set it from the decision you need to make. Sampling for a general quality read needs enough per week to see rare failures at all, which usually means reviewing more when volume is low, not less. Targeted review after a model or prompt change needs a fixed case set rather than a sample, so results are comparable across versions. Expect the analyst to tell you when a number is too small to support the claim being made from it, and treat that pushback as the job working.

References

  1. 1. Who to Hire for Agentic AI Orchestration Aquent, 2026. aquent.com Supports the claim that Agent Quality Analyst is named as a distinct hire for agentic AI orchestration, responsible for reviewing and evaluating agent outputs.
  2. 2. Agents, Human Agency, and the Opportunity for Every Organization Microsoft 2026 Work Trend Index, 2026. microsoft.com Supports the claim that 50 percent of workers identify quality control of AI output as an increasingly important skill.
  3. 3. AI Is Reshaping Jobs Faster Than Companies Are Reshaping Work BCG AI at Work 2026, via PR Newswire, 2026. prnewswire.com Supports the claim that 47 percent of workers report spending more time managing and directing AI than doing the work itself.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.