Roles
The Research Agent Orchestrator Decides What Reaches the Bench
A Research Agent Orchestrator runs multi-agent research systems for a lab: frames the research goal precisely enough for a machine to work from, configures and supervises the agents that generate and debate hypotheses, triages what comes back, and decides which machine-proposed hypotheses earn bench time or expert review. The same person owns the audit trail from agent output to experiment, so a result stays attributable. Hire for the quality of the goal statement and the triage, not for tool familiarity.
The takeThe instinct to hire an engineer here is wrong. The bottleneck in an agent-run literature review is not orchestration code, which the vendors now ship. It is judgment about which of forty plausible hypotheses deserves a scarce assay, a scarce animal cohort, or three weeks of a postdoc's attention. That judgment lives in someone who has run experiments and lost months to a bad one, and it does not transfer from a person who has only read about them. Hire a scientist who learned to direct agents, teach the tooling in a month, and give that person standing to send the whole batch back.
Where Olive fits
Open a role and see what the work shows
An interview can capture a candidate describing how they would check a confident machine-generated claim; it cannot capture them checking one. Olive puts that in front of them as work: an assignment, an assistant that will overreach, and a human reviewer who writes what actually happened at each moment.
Rank your shortlistForty Hypotheses Landed on Friday and Nobody Owns the Triage
The co-scientist run finished Friday at four. Forty hypotheses came back, each with a rationale, a confidence and a suggested first experiment. Three look obviously wrong. Six restate work your own group published in 2023. Two are interesting enough to be unsettling, and your bench has room for one this quarter. Nobody in the lab owns the decision about which one.
That gap is the role. Google's AI co-scientist is built as a multi-agent system that generates, debates and evolves hypotheses, and its drug repurposing proposals for acute myeloid leukemia were subsequently confirmed to inhibit tumor cell viability in wet-lab work 1. Systems of that shape now ship from several vendors, spanning hypothesis generation through experiment analysis 3. What none of them ship is the person who says which output is worth a reagent.
Three traits separate a real orchestrator from someone who has watched a demo. The first is the ability to write a research goal a machine can actually work from. Ask a candidate to state, out loud, the goal they would hand the system for a project you are running. A weak version is a topic: mechanisms of resistance in our target. A strong version carries the constraint, the exclusion and the shape of an acceptable answer: mechanisms that are testable in the two cell lines the lab already maintains, excluding anything requiring a knockout mouse, ranked by whether a first result would arrive inside six weeks. The second version is most of the job, and it is written before any agent runs.
The second trait is the reflex to reject in volume. Watch how a candidate talks about the forty. Someone who wants to pursue the top three by the system's own ordering has misunderstood what the ordering is: an internal tournament among the agents, not evidence about the world. Someone worth hiring narrates the cut. Six are prior art, and here is the check they ran. Nine assume an assay the lab does not have. Four cite a review rather than a primary result, which means the underlying claim has not been read. Two survive, and one of those goes to a human expert before it goes anywhere near a bench.
The third trait is the habit of tracing a claim to its source. The tell is small and hard to fake: when a proposal rests on a cited paper, does the candidate open the paper? Ask about the last time an assistant handed them a confident citation that did not support what it was cited for. A real practitioner has that story ready, with the specific claim named, because the experience is common and memorable. A candidate who has never caught one has either not used these systems on real work or has not been checking.
Which Backgrounds Produce a Research Agent Orchestrator Who Can Referee?
The reliable feeder is a working scientist who has already run other people's ideas past a bench and killed most of them: a senior postdoc, a staff scientist, a group leader who did their own literature reviews rather than delegating them. What they bring is calibrated pessimism about how often a plausible mechanism survives contact with an assay, and that calibration is the thing you cannot teach in a quarter.
Systematic review methodologists are the second pool and are usually overlooked. Someone who has run a Cochrane-style review has spent a career on inclusion and exclusion criteria written in advance, on inter-rater disagreement, and on the discipline of recording why each excluded paper was excluded. That is the audit trail the agent version needs, and they already have the habit. Scientific editors and peer review coordinators land close behind, for the same reason: judging a claim against its evidence is their daily work.
The unexpected feeders are worth deliberately sourcing. Research librarians and information specialists build search strategies for a living, and a search strategy is a research goal written for a machine. Patent examiners spend their days deciding whether a claim is genuinely novel against prior art, which is exactly the check that kills a third of any hypothesis batch. Regulatory affairs scientists arrive already fluent in reproducibility documentation. None of them will have the agent tooling on their resume, and that matters less than it looks.
Two profiles read well on paper and disappoint in practice. The first is the pure ML engineer who can wire agents together but has no view on whether a proposed experiment is fundable, and who therefore optimizes throughput of hypotheses instead of quality of the one that ships. The second is the enthusiastic generalist whose evidence is a large volume of runs. Volume is not the constraint here. The same distinction shows up in commercial versions of this work, where a retail agent orchestrator is judged on the decisions the agents were allowed to make rather than on how many ran.
What every one of these backgrounds needs added is the reading of a trace. A hypothesis is not just its text; it is the prompt, the model version, the retrieved sources, the critic agent's objection and the response to it. Screen for curiosity about that layer, not for prior exposure to any specific product, because the products are eighteen months old and the ones your lab uses in 2028 do not exist yet.
Ask How the Candidate Got Good at Refereeing a Machine
Ask directly how they got good at this and listen for practice rather than reading. The answers worth hearing are specific and slightly embarrassing: a system produced something plausible, they believed it, it cost them two weeks or a failed experiment, and they changed how they work. They can name the claim, name how it fell apart, and name the check they now run every time before anything reaches a bench.
Good answers share a shape. Someone describes running the same research goal twice with different framings to see how much of the output was an artifact of their own wording. Someone else describes routinely asking the system for the strongest argument against its own top proposal and treating a weak counterargument as a signal that the proposal was never contested. A third keeps a file of proposals that were pursued and failed, which becomes the lab's negative result set and the only honest measure of whether the pipeline is helping.
The underlying skill is checking a claim against something outside the conversation. The system says a compound has been shown to modulate the pathway; the orchestrator opens the paper and finds the effect was in a different cell type at a concentration nobody could reach in vivo. This habit describes badly in an interview, because describing verification takes thirty seconds and performing it under time pressure takes an hour, and the two sound identical across a table.
So make the candidate do it. Hand over eight machine-generated hypotheses from a real run in your own domain, including two you already know are prior art and one that rests on a misread citation, and give ninety minutes. Do not tell them how many are bad. What you are watching for is the order of operations: whether they check novelty before they check feasibility, whether they open sources or reason from the abstract, and whether they say plainly that a proposal cannot be judged without a domain expert you have not given them. That last sentence is a strong signal, not a hedge.
One caution about vocabulary. This field rewards fluent talk. A candidate saying critic agent, tournament ranking, grounding and provenance may have run a lab pipeline for a year or read three blog posts last week, and the interview transcript looks the same either way. Only the worked sample separates them.
Where Do You Find Research Agent Orchestrators, and What Closes Them?
Look inside your own institution first. The person already doing an unpaid version of this job is the postdoc or staff scientist who quietly runs agent tools for the group and gets asked to check other people's output. They know your assays, your budget cycle and which collaborator to call, and that context takes a year for an outside hire to rebuild. Ask who in the building people already forward machine output to.
Outside, go where the work is discussed rather than where the tools are marketed. Preprint servers are a genuine sourcing channel: read arXiv, bioRxiv and medRxiv for papers where the methods section describes an agent-assisted review honestly, including what it got wrong, and write to the author who wrote that section. Evidence synthesis and systematic review communities, including the Cochrane and Campbell networks, are full of people trained in exactly the judgment this needs. Research software engineering groups and university core facilities are the third pool. Titles worth approaching directly: staff scientist, research librarian, systematic review methodologist, scientific editor, computational biologist.
What they care about decides the offer more than money does. Every candidate worth hiring will ask a version of the same question: do I have authority to say no, and to whom. The offer dies the moment they learn the pipeline reports to whoever is defending its budget, because a person paid to justify a tool cannot also be the person who rejects thirty-eight of its forty outputs. It dies again if the answer to who gets authorship on work the pipeline shaped is unresolved. Resolve it in writing before the offer, not after the first paper.
Three things close the hire. Name the standing meeting where triaged hypotheses become experiments, and name who chairs it. State the reject rate you expect, and say plainly that a high one is the job working. And give the audit trail a real owner, because keeping model versions, prompts and retrieved sources attached to every proposal is engineering work needing a partner with the reliability habits of an AI SRE rather than a scientist's spare afternoons. The co-scientist approach has been reported to compress early hypothesis generation from weeks to days in some cases 2; a lab that generates faster than it triages has moved its bottleneck, not removed it.
What Does the Role Cost, and Should It Sit in the Lab or Remote?
No wage series covers this title, so price it the way your institution already prices scarce judgment: against the senior staff scientist band in the same discipline, because that is the person whose call on a scarce assay this role is taking over. Anything published under the title today describes one employer's scope rather than a market. Adjust from that band once the two questions below are settled.
Two scope questions move the band more than the title does. The first is whether the person allocates bench resources or only recommends. A recommender is priced as a senior individual contributor; someone who can commit a technician's month and an assay budget is doing line management work and should be paid for it. The second is whether they build and maintain the pipeline as well as run it. Where the role includes real engineering, you will be competing with industry machine learning salaries rather than with academic scales, and that comparison is usually the reason an academic search fails twice before anyone revisits the level.
Be careful with current titles as evidence. This role is new enough that scope varies wildly between institutions, and two people called AI co-scientist lead may have had entirely different authority. Ask what they were allowed to stop. That answer, not the title, tells you which band applies, and it is also the question that separates a candidate who ran a pipeline from one who was allowed to watch it.
On location, the orchestration itself is fully remote. Goal framing, agent configuration, triage and the audit trail are reading and writing work, and they travel. What resists remote is the handoff to the bench. The conversation where an orchestrator and an experimentalist decide that a hypothesis is testable in three weeks with existing reagents is faster in a room, and labs that run this well tend to put the orchestrator on site two or three days a week rather than fully remote or fully resident. Manufacturing has landed in a similar place with its agentic operations orchestrator, for the same reason: the decisions are remote, the consequences are physical.
On-premise constraints are real in specific settings and absent in most. If the corpus includes unpublished internal data, patient-level records, or material under a sponsor confidentiality agreement, the constraint is not the person's desk but where the model and the retrieval index run. Scope that before the offer, because it decides which vendors are eligible at all.
Common questions
How do I become a Research Agent Orchestrator?
Start from a research background where you have already judged evidence: bench science, systematic review, scientific editing, research librarianship. Then do the work the role is made of. Take a real question in your field, write a research goal precise enough that a machine could act on it, run it through an agent research system you can access, and triage what comes back in writing. Record why each proposal was cut, open every source a surviving proposal rests on, and note where the system was confidently wrong. Publish that triage log. A hiring conversation about it goes further than any certificate.
Should our lab adopt an AI co-scientist system at all?
Adopt it if your bottleneck is generating candidate directions and you have capacity to evaluate them. Do not adopt it if your bottleneck is experimental throughput, because faster hypothesis generation then only widens a queue you already cannot clear. Before signing anything, decide who triages the output and what their reject authority is. Systems of this kind have produced proposals that held up in wet-lab follow-up, but the published cases involve expert review between the model and the bench rather than direct execution of what the system suggested.
How do you evaluate an AI-generated research hypothesis?
Run four checks in order, and stop at the first failure. Novelty: has this been published, including in your own group's back catalogue. Grounding: open the sources cited and confirm each one supports the specific claim it is attached to, rather than a neighboring one. Feasibility: can it be tested with assays, models and reagents you can actually obtain this year. Discriminating power: would a negative result tell you anything. Most batches lose most of their content at the first two checks, which is normal and is why the triage role exists.
Is this a scientist role or an engineering role?
Primarily a scientist role. The orchestration mechanics are shipped by the vendors and can be learned in weeks; the judgment about which proposal earns scarce experimental time cannot. The pattern that works is a scientist who owns goal framing, triage and the audit trail, paired with an engineer who owns pipeline reliability, versioning and provenance capture. Where a single hire must cover both, expect the engineering half to win when time is short, because triage is the part that quietly gets skipped under deadline.
What does a Research Agent Orchestrator job description need to include?
Three things most drafts omit. State the reject authority explicitly: this person can decline an entire batch and is measured on the quality of what proceeds, not the count. State the reporting line, and keep it out from under whoever defends the tool's budget. State the authorship and attribution policy for work the pipeline shaped, in writing, before anyone accepts. Everything else, including specific vendor names, is detail that will change within two years and should not be written as a requirement.
References
- 1. Accelerating scientific breakthroughs with an AI co-scientist ✓ research.google Supports the description of the co-scientist as a multi-agent system that generates, debates and evolves hypotheses, and the report that its drug repurposing proposals for acute myeloid leukemia inhibited tumor cell viability in wet-lab testing.
- 2. Google AI co-scientist can reduce early hypothesis generation from weeks to days in some cases ✓ rdworldonline.com Supports the hedged claim that early hypothesis generation has been compressed from weeks to days in some cases, not as a general result.
- 3. AI companies introduce agent-based research tools cen.acs.org Supports the claim that multiple vendors now ship agent-based tools for scientific discovery, spanning hypothesis generation through experiment analysis.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.