Roles

Hiring a Scientific Foundation Model Scientist: Who Actually Brings Foundation Models Into Your Science?

You are looking for a scientist who trains and adapts large models on scientific data (protein sequences, molecules, omics, reaction records) rather than on web text, and who turns model output into hypotheses a lab can test. That person is a builder: they make training-set, split and objective decisions and defend them. Interview on those decisions, not on tool fluency. Pharma AI groups now run standing foundation-model teams for exactly this work.

The takeMost candidates who say foundation models mean they call one. The job you are hiring for is upstream of that: deciding what the model is trained on, what a holdout has to exclude before a benchmark means anything, and which of its outputs is a hypothesis rather than an artifact of the corpus. Framework fluency is a month of ramp. Judgment about scientific training data is years, and it is the only part of this role that a strong engineer cannot borrow from a checkpoint someone else released.

Where Olive fits

Open a role and see what the work shows

An interview can capture a candidate describing how they would check a confident claim; it cannot capture them checking one. Olive puts that in front of them as work: an assignment, an assistant that will overreach, and a human reviewer who writes what actually happened at each moment.

Rank your shortlist

When a Protein Model Rewrites the Quarter, Who Owns the Training Run?

A structure prediction lands on a Tuesday and a target your team had queued for two years no longer looks worth the two years. Somebody has to say whether that is a real result or a model recovering a homolog it memorized. In most organizations no one owns that answer, because the people who use the model did not build it and the people who built it work at another company.

That gap is the role. A scientific foundation model scientist trains and adapts large models over scientific data rather than web text, and translates what comes out into hypotheses a bench can falsify. The distinction that matters on day one is builder versus user. A user picks a checkpoint, prompts it, and reports what it said. A builder decides what goes into the corpus, how the split is drawn, what the objective rewards, and therefore what the model is structurally incapable of knowing.

This is now a standing function rather than an experiment. Genentech's Prescient Design has been hiring machine learning scientists into a foundation models team building internal reasoning models for drug discovery, and similar groups exist at Recursion, Insitro and Isomorphic Labs, with traditional pharma standing up internal AI organizations behind them 1. Live postings describe the work in those terms: scientific reasoning models, applied to drug discovery, owned by scientists rather than by a platform team 2.

So the practical question for a hiring manager is not whether to add AI capability. It is whether the person you add can be held accountable for a training decision, or whether they will forward you someone else's model card.

Which Tells Separate a Foundation Model Builder From a Fluent User?

Three tells, all visible inside an hour, all about provenance. A builder talks about the corpus before the architecture. A builder describes a split by what it had to exclude. And a builder can name a model they trained that did not work and say precisely why. Fluent users clear none of these, and they sound excellent until the questions get concrete.

Start with the corpus. Ask what data a protein or molecular model of theirs was trained on and listen for whether the answer has edges: which databases, which release, what was deduplicated, what was deliberately left out, what redundancy remained and how it was measured. The performed version names the dataset and moves on. The real version has opinions about it, usually irritated ones, because they spent a month cleaning it.

Then the split. This is the single most reliable separator in scientific ML, because sequence and structure data leaks in ways that random splits do not catch. A strong candidate volunteers that they split by sequence identity, by scaffold, by time, or by assay campaign, and can say what a naive random split would have inflated. If a candidate quotes a benchmark number without ever mentioning how the holdout was drawn, you have a reader of leaderboards.

Third, failure. Ask for a training run that produced a model they threw away. The answers that matter have a diagnosis attached: the objective rewarded a shortcut, the validation set shared a scaffold with training, the loss looked fine while the embedding collapsed. Vagueness here is not modesty. Anyone who has trained models at this scale has a graveyard, and they remember the headstones.

One trait is easy to miss and worth screening for directly: whether they can state what their model output would look like if the underlying biology or materials chemistry were different. That is the habit that turns a prediction into a hypothesis somebody can go disconfirm, and it is what separates this hire from a research assistant with a GPU allocation. The scrutiny it demands is the same scrutiny that makes a research integrity analyst useful in the room next door.

Which Backgrounds Produce This Person, Including the Ones You Are Not Posting To?

The expected profile is a machine learning PhD who did their thesis on biological or chemical data, or a computational biologist who moved deep into modeling. Both exist, both are being recruited hard by the named foundation-model groups 1, and a search built only on that profile is a bidding war you will probably lose on cash.

The unexpected backgrounds are better than their resumes suggest. Physical chemists and computational materials scientists who ran density functional theory or molecular dynamics for years already think in terms of an approximation with a known domain of validity, which is the correct mental model for a learned potential. Structural biologists carry the same instinct about geometry. Statistical geneticists arrive pre-trained on leakage, population structure and replication, which is exactly the split problem wearing different clothes.

Two more feeders are underused. Assay and screening scientists know what the labels actually are: how noisy the readout is, what the plate effects do, what a reported IC50 means across two labs. A model is only as good as its labels, and almost nobody on a modeling team has touched them. And self-supervised learning researchers from speech or code, who never worked in biology, transfer better than the biology-first path in one specific respect: they have real experience of what a pretraining objective does and does not teach a model. Filter them on one thing, which is whether they will accept that the wet lab is the source of truth rather than a slow evaluation loop.

A note on titles. Candidates arrive as ML Scientist, BioML Scientist, or Foundation Models Researcher, and the same person may be posted under all three. Search on the work described rather than on the noun, and read the papers rather than the affiliations. Teams that get this right tend to have already rethought how they route scientific work generally, which is the same instinct behind hiring a research agent orchestrator.

How This Person Learned to Use AI on Their Own Work

The people who are good at this got good through a repeated loop that is cheap to describe and hard to fake: build something, make it produce a confident claim, then check that claim against a measurement the model never saw, and write down where it broke. Do that a few hundred times and you develop a feel for which outputs are knowledge and which are corpus statistics.

Ask about the loop directly and listen for texture. Strong candidates talk about holding out a whole protein family, or a time slice of a reaction database, or one assay campaign, then watching the metric collapse and understanding why. They can tell you which of their models is reliable on a fold family and useless on a flexible loop, and they learned that from being wrong in front of a chemist rather than from a paper.

Coding assistants are a separate question and worth asking separately, because this role uses them heavily and the failure mode is specific. Assistants are genuinely good at training-loop boilerplate, data-loader plumbing and reading into an unfamiliar experimental literature. They are confidently wrong about evaluation, and an evaluation bug is invisible: it does not crash, it just reports a better number. The candidates you want describe a hard line there. Anything touching the split, the metric or the claim gets checked by hand or by a test they wrote themselves.

Which is why a conversation is a thin instrument for this. An interview can capture a candidate describing how they would verify a suspicious result. It cannot capture them verifying one, and the difference between those two is most of the job. If you can put a real task and a live assistant in front of them and watch which claims they go and check, do that instead of adding a fourth panel.

Find, Pay and Close a Scientific Foundation Model Scientist

Find them where methods get argued rather than where results get announced. Machine learning for structural biology, molecules and materials workshops attached to NeurIPS and ICLR are the densest venue, along with ISMB, MLCB, and the preprint traffic on bioRxiv, arXiv and ChemRxiv. Open-source projects are the most honest signal available, because the code and the benchmark choices are public and other people have criticized them.

On pay, use a sourced range and say what it is. A specialist board aggregating more than one hundred machine learning and AI roles at biotech, pharma and research institutions heads that list with a range of $148,000 to $232,000 as of 2026, with individual senior and principal postings on the same board running well above it 3. That is a posting range across seniorities, geographies and company stages rather than a survey of this title, so treat it as a band to bracket an offer inside, not as a market rate. No public wage series tracks this role, because it is not a standard occupational category. What the band does not capture is equity, which at the AI-native drug discovery companies is often the larger number and the reason a candidate leaves academia.

What closes them is compute and publication. Ask any strong candidate what they need and the first answer is usually a specific GPU allocation they do not have to requisition per project, followed by the right to publish, followed by access to proprietary data that does not exist anywhere else. That last one is your real advantage over a frontier lab: nobody outside your company has your assay history. Say so.

What kills the offer, in order of how often it happens: discovering the role is a service desk that produces predictions for other people's decisions, a data access process that takes a quarter, a publication policy that requires legal review with no time bound, and compute that turns out to be shared with production. Any one of those and a candidate who was excited will take the other offer quietly.

On location, be specific rather than generous. The training work is genuinely remote-capable and many of these teams are distributed. The part that degrades is the standing next to a chemist while somebody decides what to make, which is where the hypotheses come from. Most groups doing this well land on a distributed team with regular onsite time near a bench, and a few run fully remote with an explicit rotation. Name which one in the posting. In this market vagueness about location reads as vagueness about whether the role has a seat at the science, and that is the reading that loses candidates.

See how it works

Common questions

How do I become a scientific foundation model scientist?

Train something end to end on scientific data and be honest in public about how it failed. The modeling half is learnable from open material: protein language models, molecular property prediction, self-supervised pretraining, and the benchmark literature that criticizes all of it. The scarce half is data judgment, so spend real time on a corpus. Build a split that a domain expert cannot break, reproduce a published result and document where it does not hold, and get near an assay so you know what your labels actually mean. Hiring teams read a public record of disconfirmation as evidence in a way they no longer read tool lists.

What is the difference between this role and a computational biologist?

Direction of work. A computational biologist analyzes data and increasingly evaluates model output to decide what earns lab time. A scientific foundation model scientist builds the model: chooses the training corpus, the objective and the split, and owns whether the thing generalizes at all. The roles overlap in vocabulary and are often posted with similar titles, so read the responsibilities rather than the noun. A team that needs someone to triage predictions is hiring the first role. A team that needs a model no vendor sells is hiring the second.

What should the job description ask for instead of a framework list?

Ask for the decisions. Name the data the person will own, the scientific question the model is meant to inform, and the compute they will actually have. Request one written example of a training corpus they assembled and one split they designed, with the leakage it was built to prevent. Keep architectures as context rather than as a filter, since the specific families turn over faster than a hiring cycle. Someone who can defend a split can learn a framework in a month; the reverse is not true.

Do we need to train our own foundation model, or adapt an open one?

Most teams should adapt first, and that does not change who you hire. Adaptation still requires the same judgment: whether the pretraining corpus overlaps with your evaluation, whether the fine-tuning data is large enough to matter, whether the resulting model knows anything your assays did not already tell you. The person who can answer those is the same person who could pretrain, which is why hiring on training-data judgment rather than on scale experience is the safer filter.

How do you interview for this when nobody on the panel builds models?

Bring the panel to the evidence rather than to the architecture. Anyone in your organization who understands your assays can judge whether a candidate's claim about a prediction is falsifiable, and that is most of the signal. Have the candidate present a model they built to your scientists and ask what experiment would prove it wrong. If you need a technical read on the modeling itself, borrow a reviewer from a collaborator or an academic group rather than skipping the question, and be candid with the candidate about who is assessing what.

References

  1. 1. Biotech AI Hiring 2026 KORE1, 2026. kore1.com Reports Genentech's Prescient Design hiring machine learning scientists into a foundation models team building internal reasoning models for drug discovery, alongside AI-native drug discovery companies including Recursion, Insitro and Isomorphic Labs.
  2. 2. Machine Learning Scientist, Scientific Reasoning Models, AI for Drug Discovery jobRxiv, 2026. jobrxiv.org A live posting describing the role in its own terms: scientific reasoning models applied to drug discovery, scoped as a scientist position rather than a platform engineering one.
  3. 3. Machine Learning Jobs CompBioJobs, 2026. compbiojobs.com A computational biology job board aggregating more than one hundred machine learning and AI roles at biotech, pharma and research institutions, heading that list with a range of $148,000 to $232,000 as of 2026, with senior and principal postings above it. A posting range across seniorities and geographies, not a wage survey of this title.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.