Teams
How Do You Structure an AI-Heavy New Hire's First 90 Days?
A new hire who runs everything through AI can't be read from the finished work; an assistant produces a fluent deliverable either way. So structure the first ninety days around the steps: three weeks where every task ships with a short framing note, one artifact between brief and answer, and one named check on the claim the work rests on. Allow AI throughout and say so on day one. Then drop the requirement in week four and watch what survives. The answer lands well before month three.
The takeNotice what this ramp actually inspects: your own review path, more than the hire. When no record exists of how any work got made, at any level, someone new is only where that absence becomes visible. Worth asking, too, whether the people who look worst under a scaffolded ramp are the ones being candid about what they generated, while the colleague who quietly tidies it up first reads as fine. So run it on everyone who joins, or run it on nobody. Aimed at one desk, it isn't a ramp; it's a suspicion with a calendar attached.
Where Olive fits
Open a role and see what the work shows
The behaviors this ramp is built to make visible are the six Olive reads from a session: problem framing, evidence sourcing, delegation boundary, working structure, output rejection and verification. Each comes back as a separately evidenced finding written by a human reviewer, read off a recorded piece of occupational work rather than off a self-assessment.
Rank your shortlistWhat do the first three weeks have to produce?
Evidence, not output. The first three weeks have to produce a record of how the work got made: what your hire thought the problem was before generating anything, what they built between the brief and the deliverable, and which claim they actually checked. Assign tasks where those three things are the submission, and you stop guessing at a finished artifact that cannot answer the question.
That is a change to the assignment, not to your attention. A ramp that hands over real work and reads the result is reading the one thing an assistant can produce well no matter who is holding it. The three parts that do carry signal are cheap to ask for and easy to read:
- A framing note, written first. Two or three sentences: what is being decided, and what would make an answer wrong. It goes in before anything is generated.
- One intermediate artifact. A criteria list, an outline, a reproduction, a test that fails. It has to exist before the deliverable, and the deliverable has to be readable against it.
- One named check. Which single claim the work rests on, what it was checked against, and what changed as a result.
Nothing here should arrive as a surprise in week five. Tell your hire in week one that the note, the artifact and the check are part of every submission, that AI is allowed throughout, and that how much of it they use is not being counted. Someone who knows the standard can meet it. Someone who cannot meet it while knowing it has told you something real, and told you in three weeks.
Why won't the finished work tell you what they know?
Because a fluent deliverable is what an assistant produces whether or not the person holding it understood any of it, and because nobody involved can tell from the inside. The two signals worth having are what the work rests on and what got tested against something outside the conversation. Both live in the process, and both are gone by the time the document is finished.
Asking will not close the gap either. METR ran 16 experienced open-source developers through 246 real issues in repositories they already maintained, and found that when they were allowed to use AI tools they took 19% longer to complete issues. Those developers had forecast a 24% speedup going in, and after living through the slowdown they still believed AI had sped them up by 20% 1. That is not a story about carelessness. It is a measurement of how badly first-hand experience reads its own productivity, which is exactly what a month-one check-in asks a new hire to do.
The drift runs the wrong way for someone new. In a Microsoft Research and Carnegie Mellon survey of 319 knowledge workers, who between them gave 936 first-hand examples of using generative AI in work tasks, higher confidence in the tool was associated with less critical thinking, while higher confidence in one's own ability was associated with more 2. A hire in month one has the least of the second kind, in an unfamiliar codebase or an unfamiliar market where nothing yet looks obviously wrong to them. The same paper names what the remaining work becomes: critical thinking shifts toward information verification, response integration and task stewardship 2. Those three are the job now, and they are what the ramp has to look at.
The errors that survive are the ones that read fine. In Stack Overflow's 2025 developer survey the top-reported frustration with AI tools was "AI solutions that are almost right, but not quite" at 66% of respondents, and 45.2% said debugging AI-generated code is more time-consuming 3. That near-miss class is what the ramp has to catch, because it clears a review and shows up later, somewhere a person in month one has no reason to be looking. Why a candidate who demos brilliantly falls apart in month one is the same mechanism read backwards, from the failure instead of from the plan.
Design the ninety days in three stages, not thirty-sixty-ninety
The calendar is not the variable. Stage the work by how much scaffolding you take away: weeks one to three with the note, the artifact and the check required; weeks four to eight with the requirement dropped and the same quality of work expected; weeks nine to thirteen owning something small with a real consequence attached. The answer arrives at the first removal, not at the ninety-day review.
Weeks one to three. Scaffolded, and stated out loud. Three or four pieces of real work, each under a day, each submitted with the note, the artifact and the check. Read those three and not the polish. Write down what you saw in the same words every time, so the fourth task is comparable to the first. That is why a rubric two people can score the same way beats a good instinct here: your instinct will drift toward whoever writes most like you.
Weeks four to eight. The scaffolding comes off. Same shape of task, nothing required alongside it. What you are reading now is whether the framing and the check still happen when nobody asks. A hire who kept the habit has internalized it. A hire whose work goes back to arriving finished and unsourced was doing homework, and week six is a good time to know that.
Weeks nine to thirteen. One thing they own. A recurring report someone actually reads, a small service nobody else is on call for, a client deliverable with their name on it. Ownership is the only stage that tests stewardship, because it is the only one where nobody downstream is quietly catching things on their behalf.
Two details do most of the work. Put one thing wrong in the source material for at least one week-two task, somewhere only checking catches, and say nothing about it: planting a single error is the cheapest verification test available, and it works whether or not an assistant was involved. And keep every task inside a day. A week-long assignment tells you about their calendar. Four one-day assignments tell you whether the second one changed after you talked about the first. See how Olive measures this.
What does week two look like in finance, engineering or marketing?
Different in the material, identical in the shape. The task is small, the source packet settles nothing until somebody opens it, and the check is the part you read. What changes by field is the price of an unchecked claim: a copy edit in one, a recommendation resting on a number nobody opened in another.
- Financial analysis. A one-page read on a small acquisition, from a packet where the deck's revenue split and the audited statement behind it do not add up the same way. Read whether the two ever got reconciled, and what the reconciliation moved.
- Software engineering. A bug in a corner of the repo nobody has touched this year. Ask for the failing test before the fix. A patch that arrives green with no test written first is the thing to ask about, not the assistant that helped write it.
- Marketing. A positioning paragraph whose headline number appears in three secondary write-ups and in no primary source. Read whether anyone chased it back to the original, and what they did on finding a vendor's own customer panel at the end of it.
- Data and analytics. A finished chart with a plausible story attached, drawn from a dataset that does not carry the story. Read whether anything was recomputed, and whether the recomputation changed the claim.
- Legal operations. A short position on one clause, where the standard playbook and the last thing the company actually signed point different ways. Read which one they treated as binding, and whether they said why.
- Healthcare revenue cycle. A batch of denials holding one that is correctly denied, for which an appeal would still read persuasively. Read whether that one got written off out loud rather than drafted.
After each one, ask the same two questions in the same words: which claim is this resting on, and what did you check it against. Two people on your team should be able to write down the answers separately and agree. If they cannot, the task was too big, not the hire.
How do you run this without it reading as an audit?
Put the design in writing in week one, and apply it to everyone who joins the team rather than to the one person whose output made you uneasy. The standard is a description of the job: frame it, build something in between, check the load-bearing claim. Stated and applied evenly, it is a ramp. Unstated and pointed at one person, it is surveillance, and it reads that way.
Three things keep it honest. Your hire sees what you wrote about each task, in the same words you keep for yourself. AI use is allowed and uncounted, because volume of use was never the measure. And the outcome per task is what happened rather than a rating: the check ran or it did not, the artifact existed or it did not. A number standing for a person is worse evidence and a worse conversation.
Do not ask anyone to justify using an assistant. Ask about the work: which claim the recommendation rests on, what they nearly recommended instead, and what stayed unresolved when they shipped. A person who ran the check names the number, the source and the abandoned option without pausing. A person who did not restates the deliverable in different words, and you have your answer with nobody accused of anything.
Then be honest about what a ramp cannot fix. If week six shows the framing and the checking never happen unprompted, that is a real finding, and it is still a teaching question before it is a verdict: whether a new hire who never worked without AI is a training gap or a hiring mistake turns on whether the missing fundamentals are learnable in the time you have. Some are, in weeks, with someone senior reading their work every day. Some are not. A plan that surfaced that in week six cost you six weeks instead of two quarters, which was the whole point of moving the discovery forward.
Common questions
How long should it take to tell whether an AI-heavy hire can do the job?
Weeks, if the ramp is built for it. The first removal of scaffolding is the moment that answers it: three weeks of tasks where a framing note, one intermediate artifact and one named check are required, then the same work with none of them required. What happens in week four or five is the finding. Left undesigned, the same question usually waits for a mistake a customer sees, which is the identical discovery at a much worse price.
Should you ban AI for the first month to see their real skill?
No, unless the job itself bans it. A ban tests something the role never asks for and buys you a month of data about a working style nobody will use in month two. It also makes the tool the subject, which is the wrong subject. Allow it, say so in writing, and make the framing, the intermediate artifact and the named check the graded parts. Those are what actually vary between two people holding the same assistant.
What do you tell the new hire on day one?
Three sentences. AI is allowed on everything, and how much you use is not being counted. For the first three weeks, a short framing note, one intermediate artifact and one named check are part of every submission. You will see the notes I keep on each one, in the same words I keep them in. Said on day one, that is a standard. Discovered in week five, it is a trap, and it costs you the trust the rest of the ramp needs.
What if the work is good but they can't explain it?
Ask about decisions rather than about tools. Which claim is this resting on, what did they nearly do instead, what stayed unresolved when they shipped. Someone who ran the check names the source straight away; someone who did not restates the deliverable in other words. If that happens twice on work with a real consequence attached, treat it as the finding and move to teaching. The gap is usually verification, and verification is the most teachable item on the list.
Does this work for a senior hire?
Yes, with the stages compressed and ownership brought forward. A senior hire should be reading someone else's framing note by week four, which is its own test: a person who cannot say what is missing from a junior's criteria list is telling you what they skip in their own work. Keep the scaffolding for the first two weeks anyway. Seniority is a good predictor of domain knowledge and a poor one of checking habits, and two weeks is cheap.
How is this different from micromanaging?
Scope and an end date. You are reading three named things on small tasks for three weeks, then dropping the requirement on a date your hire knew in advance. Micromanagement reads everything indefinitely and directs the how. This directs only the what: what has to exist alongside the work. The removal in week four is the part that makes it a ramp, because the point of it is that you stop looking.
References
- 1. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced open-source developers across 246 real issues in their own repositories took 19% longer to complete issues when allowed to use AI tools; they had forecast a 24% speedup and still believed afterwards that AI had sped them up by 20%.
- 2. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ microsoft.com Survey of 319 knowledge workers who shared 936 first-hand examples of using GenAI at work: higher confidence in GenAI is associated with less critical thinking, higher self-confidence with more, and GenAI shifts the nature of critical thinking toward information verification, response integration and task stewardship.
- 3. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co The top-reported frustration with AI tools is "AI solutions that are almost right, but not quite" at 66% of respondents, and 45.2% report that debugging AI-generated code is more time-consuming.
3 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.