Roles

Who Is The National Security AI Capability Assessor, And Where Do You Find One?

Governments now do it themselves, and they staff it as a mixed team rather than one title. NIST's Center for AI Standards and Innovation describes its people as software engineers, AI research engineers and scientists, cybersecurity and biological security experts, and measurement scientists, across six teams including Frontier Assessment, Cyber, Chem/Bio and Agent Security, with a remit covering the capabilities of both US and foreign AI systems [1]. The category is real, small, and still forming.

The takeHire the person who has already had to say a capability is real when saying so was inconvenient. Everything else about this seat can be taught. The scarce thing is calibrated judgment under a classified clock: enough technical depth to run the evaluation yourself, enough domain knowledge to know what an uplift actually buys an adversary, and enough spine to write down a number that a program office, a lab, or a policymaker will hate. Most candidates who look qualified have written about capabilities. Very few have measured one and defended the measurement.

Where Olive fits

Open a role and see what the work shows

An assessment is worth what its record is worth, which is the standard this seat already applies to everybody else. Olive returns six separately-evidenced findings on one candidate, each anchored to a moment in the session, and every released report exports with its rubric, scorer and bank versions attached.

Rank your shortlist

Who Actually Assesses Frontier AI For National Security Right Now?

A cable lands saying a foreign lab has released a model that reportedly matches the best domestic system on a set of technical tasks. Somebody has to answer, within days, whether that is true, whether it matters, and what checking it would take. Nobody in the room has run that evaluation, and the answer gets briefed upward whether or not it is well founded. That moment is the job.

The most legible public example of a standing bench is NIST's Center for AI Standards and Innovation, which describes its staff as software engineers, AI research engineers and scientists, cybersecurity and biological security experts, and experienced measurement scientists, organised into six teams: Agent Security, Applied Systems, Chem/Bio, Cyber, Frontier Assessment and Partnerships. Its stated work includes assessing the capabilities of US and foreign AI systems and how those capabilities may evolve, alongside voluntary pre-deployment evaluation with frontier labs 1. Read that staffing list carefully, because it is the clearest available answer to the hiring question. There is no single credentialed profession behind this seat. There is a mixed team, and the composition is the design.

The honest caveat belongs here rather than buried later. The same page notes that most of those teams were not recruiting at the time of writing, with the Applied Systems team pointing at the National Academies' NRC Research Associateship Program as a route in 1. That is what an early category looks like: the org chart is public, the requisitions are intermittent, and the people who end up in the seat mostly arrive from somewhere adjacent. If you are staffing this function in a defence ministry, a national lab, a cleared contractor, or a policy institute, you are not competing for a labour pool that exists. You are converting one.

The distinction worth holding is between this work and product evaluation. A frontier AI safety case assessor audits whether a developer's argument that its own system is safe actually holds. A national security capability assessor asks a different question, often about a system nobody will cooperate on: what can this thing do, for whom, and what does that change. That includes assessing vulnerabilities in models an adversary built, which is closer to weapons assessment than to a benchmark run.

Which Tells Separate A Real Capability Assessor From A Briefing Writer?

The tell is whether the candidate converts a capability claim into an experiment without being asked. Put a paragraph in front of them asserting that a model provides meaningful uplift on some sensitive task. The strong candidate immediately asks what the baseline was, who the human comparison group were, how the task was scaffolded, how many attempts were allowed, and what a negative result would have looked like. The weaker one starts assessing the implications.

Four signals hold up under pressure:

  • They design for the null. Ask what evidence would convince them a feared capability is absent. Somebody who can only describe evidence for the scary direction will produce alarming findings forever, and will be right roughly by chance.
  • They separate capability from access. A model that can in principle assist with a dangerous task is a different finding from a model that does so through a public interface at low cost. Candidates who collapse those two are not yet useful in a policy room.
  • They have shipped a number they later revised. Capability estimates move as scaffolding improves. A candidate who has publicly updated an assessment, and can explain what forced it, has the calibration this seat runs on.
  • They know the limits of their own instrument. Ask what their evaluation could not have detected. Elicitation gaps, contamination, prompt sensitivity, and the distance between a checkpoint and a deployed endpoint should come out unprompted.

The anti-tells are just as sharp. Treat threat fluency without measurement experience as a research seat, not an assessment seat. Be wary of anyone who proposes to identify AI-generated artefacts as part of the method, which is not a reliable technique and would not carry a finding. And discount the candidate whose confidence never varies across domains, because the honest version of this work sounds noticeably less certain about biology than about code.

Which Backgrounds Produce This Person, And How Did They Get Good With AI?

The expected feeders are the ones you would guess: applied ML researchers who have run evaluations, offensive security researchers with real disclosure records, and intelligence analysts who have done technical target assessment. Each brings one third of the seat and is missing another third, which is why the team composition in the public example is mixed rather than uniform 1.

The unexpected feeders are often stronger. Weapons effects and operations research analysts have spent careers estimating what a system does in the hands of an opponent, under uncertainty, for a decision-maker who needs a range rather than a story. Nuclear and chemical safeguards inspectors know how to assess a capability you are not allowed to fully observe. Biosecurity and public health laboratory scientists carry the domain knowledge that decides whether a model output is a genuine step or a textbook paragraph. Metrologists and test-and-evaluation engineers from aerospace bring the discipline that keeps a result reproducible six months later, which is exactly what NIST's own framing of measurement science implies 1. Competitive capture-the-flag players and bug bounty hunters bring elicitation instincts that academic evaluation teams routinely lack.

The part hiring managers underweight is how the good ones learned. The candidates worth having did not read about model capabilities. They built the test rigs, ran the model against their own domain, and got surprised. Ask for that story and listen for specifics: the eval that scored well until they removed the hint in the system prompt; the agent that solved a task by finding the answer in a file it was not supposed to read; the week spent improving scaffolding that moved a score enough to invalidate an earlier briefing. Somebody who has watched an assistant produce a confident, wrong technical claim in their own field will recognise the same failure in a report about an adversary's system. That is the transferable skill, and it is the same instinct that makes a good safeguards enforcement analyst hard to fool.

One practical note on clearance. Requiring an existing clearance shrinks an already thin pool, and the strongest technical candidates often come from labs and open research with no prior government relationship. Decide early whether you are hiring cleared people and teaching them evaluation, or hiring evaluators and sponsoring them.

Source Candidates Where Someone Already Broke A System On The Record

Recruit where technical claims get contested in front of an audience, because that is the only reliable filter for calibration. General job boards return people who follow the field closely and have never been graded on a claim. You want the smaller group whose findings have been argued with by somebody with a stake in the answer, and who kept working afterwards.

The pools worth working, in rough order of yield: security research communities with real review, including established conference programmes and coordinated disclosure records; national laboratory test-and-evaluation and threat assessment groups; academic evaluation and red-teaming labs, whose postdocs are the most convertible population in the category; biosecurity and chemical safety programmes at universities and public health agencies; and the alumni of frontier lab evaluation teams, who are few but exactly on profile. For the US federal route specifically, the National Academies' NRC Research Associateship Program is named on the public careers page as a way into the Applied Systems team, which makes it a concrete pipeline rather than a general suggestion 1.

Screen on artefacts, not on interest in the topic. Ask for an evaluation somebody else had to respond to: a published benchmark with its method and limits stated, a disclosure write-up, a red team report, a technical assessment that carried a range and a confidence statement. Read it before the conversation and check whether the limitations section is real or decorative. In this discipline the caveats are the work product.

One warning about a common sourcing mistake. Hiring an AI engineering manager profile to lead an assessment function usually produces good delivery and weak findings, because the incentives differ: engineering leadership optimises for shipping, and assessment exists to slow a decision down when the evidence is thin.

How Do You Close One, And Where Does This Work Have To Sit?

Close on access and publication before pay. The candidates you want have all seen an assessment softened on its way upward. Name who can edit a finding, what happens when the assessment and the program office disagree, and whether anything the team produces can ever be published in redacted form. Then name the compute, the model access, and the elicitation budget, because an assessor who cannot run the system properly is producing literature review with a security classification on it.

On compensation, resist a point estimate. This category has no established band, and any number quoted for it today is somebody's guess repeated. Say instead which bands you are hiring against and why. In government, the seat is paid on the relevant public service schedule for the technical grade, sometimes with a special hiring or research associateship route rather than a standard requisition, which is what the NRC programme reference on the public careers page reflects 1. Outside government, you are bidding against cleared offensive security and frontier lab evaluation compensation, both of which sit well above general policy analysis, and that comparison is the useful thing to tell a hiring committee. As a directional macro reference only, a 2026 analysis of roughly one billion job advertisements reported an average wage premium of 62 percent for roles requiring AI skills 2. Treat that as evidence about the direction of the market as of 2026, not as a band for this title.

Location is the least negotiable part of the offer, so put it in the posting. Evaluation of open-weight and commercially available systems travels reasonably well. Three things do not. Classified material and any assessment of adversary systems built from sensitive collection happen in a controlled facility, full stop. Access to a developer's pre-deployment model under a partnership agreement is usually granted on the developer's terms and often in person. And the cross-agency coordination this work depends on still runs on people who share a corridor. State the on-site days as a number. Candidates from national laboratory or intelligence backgrounds will find that unremarkable, and the ones who push back on it are telling you which half of the pool they came from.

Read the evidence

Common questions

How do I become a national security AI capability assessor?

Build the measurement half first, in public where you can. Run a real evaluation of a deployed model in a domain you already know, publish the method, the limits and a negative result, and let people argue with it. Pair that with genuine depth in one threat domain: cyber, biology, chemistry, or autonomy. Then enter through a route that exists rather than waiting for a posting, since requisitions in this category are intermittent. Research associateship and postdoctoral programmes attached to government evaluation teams are named publicly and are currently the most reliable door. Expect clearance timelines to be part of the plan, not a formality.

Is this one job or a team of specialists?

A team, in every visible implementation. The public staffing description for the US federal center lists software engineers, AI research engineers and scientists, cybersecurity and biological security experts, and measurement scientists, split across teams for agent security, applied systems, chem and bio, cyber, frontier assessment and partnerships. One person rarely carries evaluation engineering, threat domain expertise and measurement rigour at once. Hire for the missing third of the triangle you already have, and be explicit in the posting about which third that is, or you will interview a stream of candidates who are strong at the part you did not need.

How is this different from AI safety work at a frontier lab?

The question and the cooperation level differ. A lab evaluates its own system with full access, before release, mostly to inform its own deployment decision. A national security assessor often works on systems built by someone who will not cooperate, asks what an adversary gains rather than whether a release is safe, and produces findings that feed policy and acquisition decisions instead of a launch gate. People move between the two, and the lab side is currently the larger training ground. The instincts transfer well. The access assumptions do not.

Do candidates need a clearance before you hire them?

Not usually, and requiring one narrows the pool sharply at exactly the wrong moment. The strongest evaluation engineers tend to come from open research and industry with no prior government relationship. The practical choice is whether to sponsor evaluators through clearance or to teach evaluation to already-cleared analysts. Both work. What fails is assuming the intersection is large enough to recruit from, then leaving the seat open for a year. Decide before you write the posting, and say plainly whether sponsorship is on offer, because that single line changes who applies.

What should the first six months produce?

Method and baselines, not headlines. Expect an inventory of which systems are in scope and what access exists for each, a working evaluation environment with its limits documented, and at least one baseline measurement that later assessments can be compared against. Expect arguments about what the instrument can and cannot show, which are productive. A team that briefs a dramatic capability finding in month two has usually skipped the baseline, and the first serious challenge from a technical reviewer will show it. Hiring managers who promise fast findings are setting up a credibility problem they will own.

References

  1. 1. Careers at CAISI NIST Center for AI Standards and Innovation, 2026. nist.gov Describes staff as software engineers, AI research engineers and scientists, cybersecurity and biological security experts, and experienced measurement scientists across six teams: Agent Security, Applied Systems, Chem/Bio, Cyber, Frontier Assessment and Partnerships. States the work includes assessing capabilities of US and foreign AI systems and how they may evolve, and collaborating with frontier AI labs on pre-deployment evaluations. Notes most teams were not hiring at the time of reading, with Applied Systems pointing to the National Academies NRC Research Associateship Program. Page read 1 September 2026.
  2. 2. PwC AI Jobs Barometer 2026 PwC, 2026. pwc.com Reports an average 62 percent wage premium for roles requiring AI skills across an analysis of around one billion job advertisements. Used here as a directional macro reference as of 2026, not as a compensation band for this title.

2 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.