Roles
Hire a Research Data Curator for AI Before You Train on Your Archives
The person who gets your archives ready for model training is a Research Data Curator for AI: a domain scientist who writes annotation protocols, checks provenance and licensing on every source, and sets quality gates on synthetic data used to fill gaps. Hire from your own bench scientists, data stewards and lab managers rather than from generic annotation vendors. No published salary series carries the title yet, so budget between senior research associate and senior data scientist bands, and say so out loud.
The takeMost labs hire this backwards. They buy an annotation vendor, hand it PDFs and instrument exports, and get back labels that are internally consistent and scientifically wrong. The expensive judgment is not clicking bounding boxes, it is knowing that two spectrometers in the same building disagree by a calibration constant nobody wrote down. That judgment lives in people who ran the experiments. Pay a scientist to curate the archive, or pay a model to learn your unrecorded errors.
Where Olive fits
Open a role and see what the work shows
Olive is priced per attempt rather than per seat, and an attempt returns six evidenced findings on one candidate: an input to your decision, never a filter. Ten attempts a month are free, so a pilot can run beside your current round and be compared against it.
Rank your shortlistWhat Does a Research Data Curator for AI Actually Fix?
Your fine-tuning run finished last Thursday and the model confidently reports a binding affinity your own team retracted in 2023. The retraction lived in a lab notebook; the archive kept the paper. A Research Data Curator for AI is the person who catches that before training: annotation protocols written with the scientists who ran the assays, provenance traced per record, and hard gates on synthetic data.
The job splits into three pieces that labs usually hand to three people who never speak. The first is the protocol: what counts as a positive result, how a borderline reading gets labeled, what an annotator does with a run whose instrument log is missing. In research and development that document cannot be written by a vendor, because the ambiguities are scientific and so is the arbitration.
The second is provenance and licensing. Consortium data arrives with terms attached. Some instrument contracts claim rights over raw output. A curator who cannot tell you which of your sixty datasets may legally sit in a training corpus has left your counsel a problem to find later, and the answer will be needed under time pressure.
The third is synthetic data. When an assay produced forty usable runs, augmentation is tempting and the failure mode is quiet: the model learns the generator's assumptions and reports them back as evidence. Deloitte's 2026 technology outlook names data quality specialists for synthetic data among the roles AI-native organizations are adding, describing the work as making the training data trustworthy and representative 2. In a lab that means holdouts drawn only from real runs, and a written rule for the ratio.
Which Backgrounds Produce a Good Scientific Data Curator?
The reliable feeders are people who already spent years being accountable for someone else's data: core facility managers, staff scientists on long instrument-heavy programs, biocurators from model-organism databases, clinical data managers, and research software engineers who have written an extraction pipeline for a lab. What they share is having been blamed for a bad record, which teaches a caution no course does.
The unexpected backgrounds deserve a longer interview than the obvious ones. Library and information science graduates who worked in institutional repositories arrive fluent in metadata standards and licensing, and they often cost less than a bench PhD because nobody has told them this job exists. Systematic reviewers from evidence synthesis teams have done little for years except define inclusion criteria and adjudicate disagreement between two readers, which is annotation under a different name. Field ecologists record collection conditions by habit, a discipline most wet-lab scientists never acquire.
What matters more than the degree is how the candidate got good with AI on their own work. The strong ones tell a story with a date in it: they ran an extraction over 900 papers, found the model inventing units on one journal's table format, and built a check for that specific failure instead of abandoning the pipeline. They can say where they stopped trusting output and what test they wrote at that boundary.
Ask what they threw away. A curator who has never deleted a dataset they built has never priced their own errors, and the archive you are handing them will ask for exactly that.
Where Do You Source a Research Data Curator, and How Do You Test One?
Start inside the building. The strongest candidate is usually a staff scientist two years into being annoyed by your archive, already trusted by the principal investigators, already cleared for the restricted data, and already aware of which instrument drifts. Outside, look at data-curation networks and repository staff rather than at annotation marketplaces, where the work is priced as piecework.
Venues that genuinely hold these people: Research Data Alliance and CODATA meetings, the Force11 community around scholarly data, and the curation staff behind Dryad, Zenodo and the Dataverse installations that large universities run. Domain databases are the deepest pool of all. Whoever curates entries at UniProt, the Protein Data Bank or FlyBase has spent a career turning other people's messy submissions into records other people rely on. Adjacent titles worth searching: data steward, biocurator, research data manager, clinical data manager, ontology curator.
The test that separates real from performed takes forty minutes and needs no take-home. Hand the candidate two hundred rows out of your own archive, ugly parts included: a column whose units changed in 2019, three duplicate sample identifiers, a free-text notes field. Ask for a first draft of an annotation guideline and for the questions they would put to you before labeling anything.
Performed expertise returns a tidy schema and no questions. Real expertise returns four questions, two of which you cannot answer, and a guideline with a section on what to do when the answer stays unknown. The same instinct shows up in the support knowledge curator role, where the failure is identical: content that reads clean and teaches a model something false.
What Does a Research Data Curator for AI Cost, and Can the Work Be Remote?
The person who would have caught that 2023 retraction is a scientist, and scientists have a price. No published salary series carries this title yet, so anyone quoting a precise band for it is guessing. Price it against the two roles it borrows from. The levels.fyi data scientist page reported a United States median total compensation near $180,000 as of mid-2026, with the middle half falling roughly between $134,000 and $250,000 4. Lab scientist bands sit lower.
Demand and supply pull in opposite directions here. Demand pushes up: LinkedIn's Jobs on the Rise 2026 put data annotator fourth among the fastest-growing United States roles, describing people who label and review data under detailed guidelines and quality checks so that models train on accurate datasets 1. A second ranking gets quoted in the same breath, the World Economic Forum's 2025 jobs report on big data specialists, but its page did not resolve for this piece, so no claim here rests on it 3. Supply pushes down at the bottom of the market, because generic annotation is genuinely commoditized.
Your candidate is not in the generic market. Someone who can arbitrate a scientific disagreement is a scientist and will not take a pay cut to stop being one. In practice, budget between a senior research associate and a senior data scientist for your region, and expect the top of that when the domain is regulated. Ask two or three candidates what they earn now and what a competing offer looked like; with no survey to cite, your own pipeline is the best series available.
On location the honest answer is split. Protocol writing, provenance work, licensing review and synthetic data audits are all remote-friendly. Anything touching instruments, physical samples, restricted patient records or an air-gapped archive is not. Most teams settle on two or three days on site through the first quarter, while the curator learns which records lie, then mostly remote. Decide which half your archive is before you post the role, because a fully remote posting for on-premise data wastes six weeks.
Common questions
How do I become a Research Data Curator for AI?
Start from a domain you already know. Take one real dataset you have access to, write an annotation guideline for it, get two other people to label a sample against your guideline, and measure where they disagreed. That artifact plus the disagreement analysis is a stronger portfolio than any certificate. Add the second half deliberately: licensing and provenance for the sources you used, and a written rule for how much synthetic data you would allow and why. Repositories, model-organism databases and core facilities hire on that evidence, and those jobs are the usual route into the AI-facing version of the role.
Is a Research Data Curator for AI different from a data annotator?
Yes, and the difference is who writes the rules. An annotator applies a guideline. A curator writes it, decides what happens at every ambiguous case, traces where each dataset came from and under what license, and can refuse to release a dataset for training. In research settings the two often collapse into one hire because the ambiguous cases are scientific, and only someone who has run the experiment can arbitrate them.
Should this role report to IT, to the data science team, or to the lab?
To whoever is accountable for what the model outputs, which is usually the data science or AI function rather than central IT. Reporting into IT tends to strip the curator of scientific standing exactly when they need to overrule a principal investigator about a dataset. Reporting into a single lab tends to strip them of reach across the archive. A dotted line to the research leadership and a seat in data-collection decisions solves more than a title change does.
What does a Research Data Curator for AI cost?
No salary survey publishes the title yet, so treat any exact figure with suspicion, including this one. Anchor on adjacent bands: levels.fyi reported a United States median total compensation near $180,000 for data scientists as of mid-2026, with the middle half roughly $134,000 to $250,000, while senior research associate bands in academic and biotech labs run well below that. Budget in between, pay toward the top in regulated domains, and calibrate against what your own candidates report.
How do you test data judgment in a 40-minute interview?
Give the candidate real, messy rows from your own archive and ask for a draft annotation guideline plus the questions they would ask before labeling. Judge the questions, not the schema. Strong candidates surface unit changes, duplicate identifiers and missing instrument logs unprompted, and they write down what an annotator should do when the answer is unavailable. A candidate who says the dataset is not ready for training, and explains what would make it ready, has just done the job in front of you.
References
- 1. LinkedIn Jobs on the Rise 2026: the 25 fastest-growing roles in the US ✓ linkedin.com Data Annotator ranks fourth on the 2026 list, described as labeling and reviewing data under detailed guidelines and quality checks for training AI and machine learning models.
- 2. Tech Trends 2026: AI and the future of the IT function ✓ deloitte.com Lists 'data quality specialists for synthetic data, ensuring that the training data fueling AI systems is trustworthy and representative' among emerging roles.
- 3. Future of Jobs Report 2025: the fastest-growing and declining jobs weforum.org Places big data specialists at the top of the fastest-growing jobs list through 2030. Page returned HTTP 403 to automated retrieval this session, so the figure behind the ranking is not quoted here.
- 4. Data Scientist salary in the United States ✓ levels.fyi Median total compensation reported near $180,000, with the 25th percentile near $134,000 and the 75th near $250,000, read as of mid-2026.
4 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.