Pipeline

There Is No Safe Job List. There Is a Task Exposure Test.

There is no reliable list of jobs safe from AI; every one published goes stale within months of the next model release. What holds up is a test you can run on any task: can a model produce a plausible version, does a wrong version cost someone something, does it need context no document contains, and is a specific person accountable for the result. Score your own tasks against those four questions instead of a title.

The takeSafe-jobs lists work as content, not as evidence: one entry per occupation, no mechanism connecting the claim to anything measurable, and a shelf life set in months rather than years. They share a specific error with the doom lists on the opposite side, treating exposure as a property of a job title instead of the tasks a particular person in a particular company actually does. Two people holding the same title at different employers can score completely differently on the same four questions, which is the reason a title was never going to be the right unit to begin with.

Where Olive fits

Open a role and see what the work shows

Olive sits on the employer's side of the funnel as six evidenced findings about one candidate, never a filter across a pool of applicants, and every report it releases is shown to the candidate it describes, free.

Rank your shortlist

Why Safe-Jobs Lists Keep Being Wrong

Safe-jobs lists score the wrong unit. Exposure lives in the tasks a person performs, and a title is only a loose container for them: the same job at two employers can hold a different mix of drafting, judgment, and cleanup. One verdict per occupation averages that away before the reader gets to it, and the average is what makes the list read as authoritative and behave as noise.

The research underneath the lists does not work in titles either. One large index rated roughly 2,900 work skills and judged 19 of them, 0.7% of the set, very likely to be fully replaced by current models 2. A skill is not a job, and no job in that analysis is built only from those 19, so the figure bounds nothing about how many titles vanish. What it shows is the unit the measurement lives in. Separately, in a study of millions of job postings, the occupations experiencing the most automation were the ones experiencing the most augmentation at the same time 4, so a single role absorbs both, which a one-line verdict per title was never built to show.

Titles are not meaningless, though. A title is a decent first filter for the base rate of exposed tasks inside a role. It is a poor substitute for actually running the test on your own list, which is the only version of the answer specific to you.

Run the Four Questions on One Task

Apply four questions to a single task, not a job title: what a model can fake, what a wrong answer costs, what context the work needs, and who answers for the result. Each is a yes or a no, and the pattern across all four is what carries information rather than any single answer. Write them down for the task you actually do, not the one in the job description.

  • Can a model already produce a plausible version?
  • Does a wrong version cost someone something before it's caught?
  • Does it need context that lives outside any document?
  • Is a specific person accountable for the result, by name?

A task that fails the first question and passes the other three is the kind of task that keeps a job alive. Try it on one task from your own week, an expense report you reconcile every Friday. A model can already draft a plausible reconciliation from the receipts. A wrong one costs a finance team an afternoon at most, not a client relationship. It needs almost no context outside the receipts themselves, and a supervisor's name, not yours specifically, sits on the report either way. That task scores low on three of four questions, which is a specific, checkable reason it's a good candidate to hand off rather than a threat to argue with.

The cost question matters more than it looks. In one randomized experiment, consultants using GPT-4 on a single task chosen to sit outside the model's real capability were, averaged across the two AI groups, 19 percentage points less likely to reach the correct answer than a control group working without it 1, and the group given more prompting guidance did worse, not better. The failure mode on a task like that isn't a blank page. It's a confident, wrong one, which is exactly the shape of cost the second question is built to catch.

The third question lines up with a pattern in one analysis of a large payroll panel. Occupations scoring high on codified knowledge, the formal kind taught through schooling and written procedure, show slower entry-level employment growth, while occupations scoring high on tacit knowledge, learned through practice and mentorship, show faster growth for mid-career and senior workers 5. Those are two different age groups and the authors call the gradients raw and descriptive rather than a proven mechanism, so read the direction and not the size. Context nobody wrote down is the part a document-trained model has no way to absorb.

Where the Test Disagrees With the Lists

Run the test on human resources, a field usually filed as safe because a person still signs the offer letter, and it splits. Drafting a routine policy summary or a first-pass job description scores low on cost and context. Mediating a live conflict between two employees scores high on all four questions at once.

BLS projects human resources specialists growing 6.2% from 2024 to 2034 3, against 3.1% for all occupations across the same decade. That is a headcount projection built on a 2024 base, and it does not model generative AI as a driver at all, so it settles nothing about which HR tasks are exposed. It is exactly the shape of number a safe-jobs list runs on, and it cannot answer the question the list claims to answer.

Run it the other way on an entry-level customer support role, usually filed as doomed. Answering a routine billing question scores low; a model can already draft a plausible reply. Recognizing that a specific customer's account sits inside an active fraud case, and escalating it instead of closing the ticket, scores high on all four questions. Employers hiring for exactly this gap are told to test the checking rather than the production, because a model can already produce the plausible version and someone still has to answer for it. That is the first question and the fourth, read from the other side of the table. The title tells you neither of these stories. The task does.

Is This a Guarantee?

No. It's an exposure test, not a safety guarantee, and it can only score the tasks you hand it today. A task that scores durable now can score differently after the next model release, and a task that scores exposed can stay in your job for years if nobody upstream decides to change the workflow around it. Run the test again each time your role changes, not once and filed away.

If your own task list scores exposed on most of what you do, the harder decision is whether to switch fields entirely or move sideways inside the one you already know. The test doesn't answer that question for you. It tells you which tasks are worth building the next stretch of your career around, which is the input that decision actually needs.

Keep the four answers written down somewhere you'll actually reread, not just in your head. A task's score changes quietly, one model release at a time, and the only way to notice the change is to compare this quarter's answer against the one you wrote last time.

See a sample report

Common questions

Which jobs are safest from AI?

None, as a title. Run four questions on your own tasks instead: can a model already fake a plausible version, does a wrong version cost someone something before it's caught, does it need context outside any document, and is a named person accountable for the result. A job that scores durable is usually a bundle of tasks that pass most of those, not a title on a list.

Is there a safe-jobs list I can actually trust?

No. Every list like that is a snapshot against one model generation, and it goes stale within months of the next one. The task-level test in this article doesn't expire the same way, because you can rerun it yourself whenever the tools change instead of waiting for someone to publish an update.

What if most of my tasks score exposed?

That's useful information, not a verdict. It tells you which parts of your role to move away from and which to build on, inside the same job or a different one. It doesn't tell you when, and nobody can honestly give you that part.

Are the classic "safe" jobs like nursing and teaching actually safe by this test?

Mostly, and for a specific reason: they carry a high share of tasks that need context outside any document, a patient's history, a specific kid's day, plus visible personal accountability. That's what the test is built to detect. It doesn't mean every task inside those jobs is untouched.

How often should I rerun the test?

Whenever your tools or your task list change noticeably, which lately means every few months rather than every few years. A task that failed the first question a year ago can pass it now, so the answer you wrote last is the one most likely to be out of date.

References

  1. 1. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports that AI use on a task outside the model's real capability produces confidently wrong answers, the shape of cost the exposure test's second question targets.
  2. 2. AI at Work Report 2025: How GenAI is Rewiring the DNA of Jobs Indeed Hiring Lab (Annina Hering and Arcenis Rojas), 2025. hiringlab.indeed.com Supports that full skill replacement remains a small share of what was assessed, against the fastest 'jobs AI will replace' claims.
  3. 3. Employment Projections: Occupational Projections, 2024-2034 U.S. Bureau of Labor Statistics, 2025. data.bls.gov Supports the HR growth projection used to counter the flattest 'HR is being automated away' claim while noting task content still changes.
  4. 4. Beyond the Binary: How Automation and Augmentation Are Combining to Reshape Work The Burning Glass Institute (Melissa DiMarzio), 2026. burningglassinstitute.org Supports that automation and augmentation rise together inside the same occupations rather than sorting jobs into opposite fates.
  5. 5. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Stanford Digital Economy Lab (Erik Brynjolfsson, Bharat Chandar, Ruyu Chen), 2026. digitaleconomy.stanford.edu Supports the tacit-versus-codified-knowledge pattern, named as the paper's own descriptive gradient rather than a proven mechanism.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.