Assessment design

Can an Internal Candidate Really Move Into an AI-Heavy Role?

An internal candidate's readiness for an AI-heavy role doesn't show in a tool list or in their record in the job they're leaving. Hand them one task built from the destination role's material, an hour on the clock, and read three things back: the frame written before the first prompt, one assistant output refused with a reason, and one check run outside the chat that moved the answer. A light AI footprint isn't a failure. Deciding the assistant was the wrong instrument, and saying why, is the same judgment.

The takeOf the three artifacts, the refusal is the one the building works against. Saying no to a plausible output means slowing something down in front of someone senior, and the people who do it well tend to get called difficult rather than careful. So when the refusal is the artifact missing from an internal file, most likely you are reading the last job's incentives rather than the person's ceiling. Nobody has measured that. It is still the reading I'd carry into the conversation, because a company that never rewarded a no has not earned the conclusion that this person cannot give one.

Where Olive fits

Open a role and see what the work shows

The hard part of running this in-house is the destination role's case and its answer key, because the material has to come from the job the person is moving into rather than the one they are leaving. Olive's item banks are authored per occupation and carry that occupation's SOC code, and a session returns six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer and granted to the person assessed in the same document.

Rank your shortlist

What separates enthusiasm from capability?

Capability leaves artifacts; enthusiasm leaves usage. A person who likes the tools produces long transcripts, a browser full of tabs and a firm opinion about which model is best. A person who can do AI-heavy work leaves something else behind: a written frame before the first prompt, refusals with reasons attached, and a result that changed because it was checked. Read for the second set.

The gap between the two is measurable, and it runs the wrong way from intuition. METR paid 16 experienced open-source developers to work 246 real issues in repositories they had contributed to for years, with AI tools allowed on a random half. They expected the tools to speed them up by 24%. Afterward, having done the work, they still believed the tools had sped them up by 20%. They had in fact taken 19% longer 1.

Those were not novices. Participants came in with dozens to hundreds of hours of prior prompting experience 1, which is to say they were exactly the people an internal mobility conversation describes as the team's AI person. Their self-report survived contact with their own measured performance. So when a manager says a report is great with AI, treat it as a usage report from someone with no counterfactual, not as a measurement.

This is why a tool list settles nothing on an internal application any more than it does on a resume naming six AI tools. Inside the company it gets harder, not easier: internal enthusiasm arrives with a name, a reputation and a sponsor attached, and the reputation was earned in the job the person is leaving. A live demo is the same signal in better clothes: impressive, unfalsifiable, and not what month one looks like.

Ask for three artifacts, not a tool list

Give the person a task built from the destination role's real material and ask for three things back: the intermediate artifact made before the deliverable, the record of what the assistant produced and what was refused, and one verification run outside the conversation. Each is an act rather than a claim, and each is difficult to manufacture after the fact in an hour.

ArtifactWhat to ask forReads as capableReads as enthusiasm
The frameThe plan, criteria list or outline made before the first promptNames what is being decided and what would make an answer wrongThe first message to the assistant is the deliverable
The refusalOne thing the assistant produced that was not used, and whyAn assumption or framing rejected on substance, in a sentenceEverything is kept and the edges tidied
The checkOne claim tested against something outside the chat, and what it changedA figure recomputed, a filing opened or a test run, and a number or a recommendation moves"I verified the output" with nothing behind it

The three are not a checklist to total up. They are three places where borrowed judgment shows, and the task carries more weight than the request: set it on material where the assistant will answer fluently and be wrong, and the frame, the refusal and the check either happened or they did not.

Keep the ask identical for everyone considered for that role, and put it in writing. The same three artifacts do the work on an external candidate's portfolio, where the question is who did which half; the internal version is easier, because you can hand the person the material and watch the clock instead of reading a finished piece and guessing at its history.

Do not add speed as a fourth artifact. Working fast with an assistant and working well with one come apart, and speed is what an internal reputation is usually built on. The person who turns things around by Thursday is the person everyone already calls capable.

Why their current-role record doesn't answer this

Because a move to a different job is a selection decision, and evidence about one job does not carry to another without showing the two jobs share the work. The federal guidelines treat promotion as an employment decision outright, and count transfer as one where it leads to a decision on that list 2. A strong performance review is real evidence, and it is evidence about the role the person is leaving.

The transfer rule is explicit. Validity evidence gathered on one job may be used for another only where incumbents in both "perform substantially the same major work behaviors, as shown by appropriate job analyses" on each job 3. And where the observed work behaviors and the observed work products in two jobs are not the same, the federal enforcement agencies "will presume that the work behavior(s) in each job are different" 4. Difference is the default; sameness is the thing you have to demonstrate.

Take the FP&A analyst who wants the data science role. Both jobs open a model, write in the same document tool and sit in the same review. But the observed work products are a variance commentary and a fitted model with an error analysis, which are not the same product, and the observed behaviors (reconciling a figure to the ledger, choosing a validation split) are not the same behavior. None of that is a criticism of the analyst. It means the existing record has no bearing on the question, and something drawn from the destination role has to.

The same rule sets what the review may look like. A procedure is supported by content validity only "to the extent that it is a representative sample of the content of the job," and where it samples a work behavior, its manner, setting, level and complexity "should closely approximate the work situation" 4. A generic AI exercise (write a prompt, spot the hallucination, summarize this article) approximates nobody's Tuesday, which is also why one company-wide assessment cannot serve every department.

Which matters more: domain context or tool fluency?

The destination role decides, and it is worth deciding before the review rather than after it. Internal candidates almost always arrive long on context and short on tool fluency. Where the work is adjudicable (an audit position, an underwriting file, a diligence memo), context is the scarce half. Where the assistant sits in the production path all day, the destination team's own material is the constraint.

O*NET splits job information along the same seam: descriptors usable across many jobs and industries on one side, and details specific to particular occupations (the tasks, the technology skills) on the other 5. A lateral move carries the first half and not the second. Naming which half your destination role actually runs on is most of the decision, and the answer is not stable across moves: the same person is ready for one lateral step and not another.

There is measured reason to distrust tool speed as evidence in expert work. Across 5,179 customer support agents, access to a generative AI assistant raised issues resolved per hour by 14% on average, with a 34% improvement for novice and low-skilled workers and minimal impact on experienced and highly skilled ones 6. Read that as a caution about what a fast demo proves: the assistant's visible benefit is largest exactly where the person's own judgment is thinnest, so fluency-with-the-tool is weakest evidence in precisely the roles where a wrong answer is expensive.

Two rules fall out. If the destination role is one where a wrong answer costs money, weight the frame and the check, and treat tool fluency as a thing the team teaches in a fortnight. If it is one where the assistant is the medium (engineering, analytics, content operations), weight the refusal, because the failure there is accepting fluent output at volume. Before either, settle whether the role genuinely needs AI skills yet; a great deal of internal mobility pressure traces to a mandate rather than to the work, and the whole-team version of the question is hire for the skill or train the people you have.

How to run the review without stalling the move

Cap it at an hour, run the identical task for everyone applying to that role, and write the reasons down. An internal review that takes three weeks and produces a verbal impression is worse than no review at all: it delays a person who is already employed, and it leaves no record of why the answer was what it was when the next applicant asks.

Four decisions, made once, before the first internal applicant:

  • The case. One task, from the destination role's material, with a constraint that lives in the material rather than in the brief. An hour of work time, on the clock.
  • The outcome words. Demonstrated, partly demonstrated, not demonstrated, per artifact. No total and no ordering of applicants. A number invites an average, and an average across three unlike acts is a claim none of them supports.
  • Who reads it. The manager of the destination role, plus one person who already does that job. Not the current manager, whose read is about the job being left.
  • What the person sees. The notes, in full. An internal candidate stays in the building either way, and nothing kills a mobility program faster than a decision nobody will explain.

A partly-demonstrated result is the common outcome and the useful one. It names the gap (the frame was thin, or nothing was ever refused), and a gap with a name is a development plan and a date, rather than a rejection. That is the real advantage of assessing someone you already employ: you can move a person with a known gap and close it, which is not on offer with an external hire.

Avoid two moves here. Do not accept a certificate in place of the task, because a prompt engineering certificate and a work sample are not the same evidence, and an internal training completion record is the same shape. And do not read a light AI footprint as a disqualification: someone who judged the assistant was the wrong instrument for a step and did it by hand has demonstrated the exact thing the review is looking for, which is why not using AI is not automatically a dealbreaker.

See what gets scored

Common questions

Does an internal AI training certificate count as evidence?

It records attendance, not judgment. A completion record says a person sat through the material; the review is asking whether they framed the problem before generating, refused something on substance, and tested a claim outside the chat. Ask for the three artifacts regardless. If the training ended in a graded task built on the destination role's own material, that task is the evidence and the certificate is its wrapper, so ask to see the task and what the person submitted.

What if the internal candidate is the team's most enthusiastic AI user?

That is the case this review exists for. Enthusiasm and measured benefit come apart even in people with hundreds of hours of prompting behind them, and a reputation as the team's AI person is a usage report from colleagues with no counterfactual. Run the same task you would run for anyone else, on the same material. The enthusiastic user often passes it, and when they do not, the artifact missing is almost always the refusal.

Should internal candidates take the same assessment as external ones?

Same case, same artifacts, same outcome words, for anyone under consideration for that role. Different standards for internal and external applicants are hard to explain and harder to defend, and the comparison worth having is between candidates for one job. What legitimately differs is what happens next: an internal partly-demonstrated result can become a development plan with a date on it, because the person is still in the building.

How long should the review take?

About an hour of the person's work time, plus an hour of reading. Longer costs more than it buys, and internal mobility loses people to delay more often than to a no. A case that cannot be worked in an hour is testing endurance. Put the time on the calendar as work rather than asking for an evening. The person is already being paid for their day, and asking for unpaid hours selects for who has spare ones.

What if the destination manager doesn't use AI themselves?

Pair them with someone who does that job with an assistant daily, and have both read the same artifacts. All three are legible without tool expertise: a plan either exists or it does not, a refusal either carries a reason or it does not, a check either changed something or it did not. What the second reader adds is the one judgment the first cannot make: whether the assistant's output was actually wrong.

Can you decline the move without ending the person's mobility?

Yes, if the result names a gap instead of delivering a verdict. Write down which artifact was missing, what would close it, and when you will look again. Then look again on that date. A program that produces unexplained noes stops receiving applications within two cycles, and the internal candidates you most want to keep moving are the ones with somewhere else to go.

References

  1. 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced open-source developers working 246 real issues expected a 24% speedup from AI tools and still believed afterward that they had been sped up 20%, while measured completion time was 19% longer; participants had dozens to hundreds of hours of prior prompting experience.
  2. 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.2 (Scope) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Employment decisions include hiring, promotion, demotion, membership, referral, retention and licensing; selection for training or transfer also counts where it leads to one of those decisions.
  3. 3. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.7 (Use of other validity studies) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Validity evidence from one job may be used for another only where incumbents perform substantially the same major work behaviors, shown by appropriate job analyses on both jobs.
  4. 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 (Technical standards) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Content validity holds to the extent the procedure is a representative sample of the content of the job, and the manner, setting, level and complexity should closely approximate the work situation; where observed work behaviors and work products differ between two jobs, the agencies presume the behaviors are different.
  5. 5. The O*NET Content Model O*NET Resource Center, U.S. Department of Labor Employment and Training Administration, 2026. onetcenter.org The content model pairs descriptors usable across many jobs and industries with details specific to particular occupations, including occupation-specific tasks and technology skills.
  6. 6. Generative AI at Work National Bureau of Economic Research (Working Paper 31161), 2023. nber.org Across 5,179 customer support agents, access to a generative AI assistant raised issues resolved per hour by 14% on average, with a 34% improvement for novice and low-skilled workers and minimal impact on experienced and highly skilled ones.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.