Teams

Output Ramps in Weeks Now, Judgment Still Takes Months

A new hire's ramp time split into two numbers that used to move together. Time to usable output is now days to weeks across a lot of knowledge work and it tells you almost nothing on its own. Time to unsupervised judgment is the number worth planning against, and the honest definition is a date you can record: the week you stopped reading their work before it went out. That is still months in most roles.

The takeEight months to full productivity, the most repeated ramp figure, survives because it is comfortable for both sides of the table: it excuses a slow first quarter and it justifies a long runway. Its actual provenance is an online survey of HR professionals, sponsored by a moving company, about a quantity most respondents said their company does not measure. A headcount model resting on it is resting on a self-report of a guess. Your own last three hires are better evidence and they cost an afternoon to read.

Where Olive fits

Open a role and see what the work shows

Olive reads six dimensions of AI work from a real work session rather than from a self-rating, which is the same reason a recorded ramp date beats a manager's estimate of one. Each finding comes back as demonstrated, partly demonstrated or not demonstrated, with the excerpt it rests on and never as a number.

Rank your shortlist

How long should a new hire take to get up to speed?

Days to weeks for usable output, and months for unsupervised judgment, which is why one number has stopped working. Usable output means work that reads finished, and it now arrives that fast across a lot of knowledge work. Unsupervised judgment means you no longer read the work before it leaves the building, and that is the figure a headcount plan should be built on. The gap between the two is where the expensive surprises live.

The first number moved because short, self-contained tasks got much faster. In a pre-registered experiment, 444 college-educated professionals did occupation-specific writing tasks; the half given ChatGPT finished 10 minutes faster, 37%, than a control group averaging 27 minutes, and graders scored their output 0.45 standard deviations higher, with the largest gains going to the weakest writers 2. Mid-level professional writing, done once, online, for pay, with no colleagues and no revision cycle. That is a narrow setting, and the narrow setting is the point: it is a good description of a new hire's first assignment.

The second number did not move, because it is made of something else. Knowing which of this team's conventions a generic answer violates is learned from people, from being corrected, and from watching a decision go wrong. Nothing in the tooling shortens that, and a team whose conventions live in people's heads rather than in writing will always ramp slower and will usually blame the hire.

So quote both or quote neither. A single ramp number now averages a thing that takes two weeks with a thing that takes two quarters, and the average describes nobody.

Where did the eight-month number come from?

A March 2012 online survey of 500 US HR professionals, sponsored by a relocation company. It reports eight months on average, with 27% of companies saying a year or more and 25% saying three months or less, and 58% of respondents saying their company does not measure new-hire productivity at all 1. That last figure is the one that decides how much weight the first can carry.

Look at how the number was collected. The sample is 500 HR professionals recruited online across 49 states and the District of Columbia, none of them the hires being described. There is no definition of fully productive anywhere in it. The sponsor is a moving company with a commercial interest in workforce mobility. And it is from March 2012, which predates widespread remote work, the current tooling, and generative assistants entirely 1.

That provenance does not make the figure a lie. It makes it a self-report about a quantity most respondents admitted they do not track. Cite it for what HR leaders believed in 2012, and keep it out of a hiring plan. The companion figure from the same panel, that companies lose on average 23 percent of new hires before the one-year anniversary, is self-report from the same 500 people and carries every one of the same problems 1.

The general lesson is worth more than the specific number. Every widely quoted ramp benchmark traces back to a survey of managers, and vendor pages repeating them treat the figure as a property of roles when it is a property of a measurement year.

Measure the week you stopped checking their work

Open the calendar and the review history for your last three hires and find the week their work started going out without you reading it first. That date is a measurement rather than an impression, it lives in systems you already have, and it is specific to your team, which is the point: ramp time is largely a property of how much of your context is written down.

Where the date actually sits, by function:

  • Engineering. The first week of pull requests approved with no changes requested, by someone other than the person who onboarded them.
  • Client-facing work. The first email or deck that went to a customer without a draft review.
  • Analysis and finance. The first model or memo that reached a decision-maker unedited.
  • Operations. The first week they closed a category of exception without escalating one.

Three hires gives you a spread, and the spread is the useful output. If it is eleven weeks, four weeks and thirty weeks, the variance is telling you about three different managers or three different amounts of written context, and that is a more actionable finding than any industry average.

Do not substitute asking people. In a randomized trial, 16 experienced open-source developers completed 246 real tasks on mature repositories they had worked on for about five years; allowing early-2025 AI tools made them 19% slower, after forecasting a 24% speedup beforehand and still estimating a 20% speedup afterwards 4. Sixteen developers in one setting, on codebases they knew intimately, so the magnitude does not transfer. What transfers is the gap: people misjudge their own productivity badly enough to get the sign wrong. A manager recalling when a new hire became self-sufficient is making that same kind of estimate about their own reviewing habits. Read the date off the calendar.

Why does a fast ramp mislead a headcount model?

Because the thing that ramped is not the thing the model is counting. A plan assuming full contribution at month eight, confronted with finished-looking work in month one, pulls the next requisition forward and books capacity that still needs checking. The hole shows up whenever somebody finally looks, which is later than it used to be, and the correction lands in the quarter the plan already spent.

The underlying failure, confident work that turns out wrong, has been measured in a cleaner setting. In a field experiment, on one task deliberately chosen to sit outside AI capability, consultants using GPT-4 were 19 percentage points less likely to reach the correct answer: 84.5% of the control group got it right, against 60% and 70% in the two AI conditions 3. One task, one sample, a 2023 model, and the frontier moves with every release. The durable part is that consultants at a top firm could not tell which side of the capability line the task sat on, and neither can a new hire in month one, and neither can the manager reading the output.

So the constraint is reviewing capacity. If a team can generate three times the drafts and check the same number, hiring the fourth producer buys less than the model says. Count the second number when you plan the requisition, and read what it means to hire for verification rather than production before you write it.

Two cases deserve their own answer. Early-career hires who have never worked without an assistant ramp fastest on output and slowest on the check, which is a training problem: what to do when new grads never learned to work without AI. And a ramp that looks unusually fast and then stalls is often not a ramp problem at all, which is how to tell which of three failures you are looking at.

See the benchmarks

Common questions

Is there a reliable industry benchmark for time to productivity?

No. The figures in circulation come from asking managers what they believe, which is a different exercise from measuring work. There is no government series for it, no standard definition of fully productive, and the most repeated figure comes from a 2012 online panel of HR professionals, most of whom said their employer does not measure it. A benchmark you cannot trace to an instrument is a number somebody wrote down once. Measure the date inside your own team instead, where the definition is yours and the data already exists.

How long should a senior hire take compared with a junior one?

Seniors reach usable output faster and unsupervised judgment slower than most managers expect, because the second is about your context rather than about their experience. Fifteen years of pattern-matching produces confident work quickly and produces it against the conventions of a previous employer. The specific risk is a senior hire whose output nobody checks because of the title, which is the exact condition under which a judgment gap stays invisible. Set the second date deliberately for a senior hire, the same way you would for anyone else.

Does a longer onboarding programme shorten ramp time?

There is some evidence that longer structured onboarding tracks with staying, and almost none that it shortens time to unsupervised judgment. The two are different outcomes and get measured differently. What plausibly shortens the second is writing your context down, because the conventions a new hire has to learn from colleagues are the ones nobody wrote. A team that documents its decisions is buying ramp time for every future hire at once, which is a better investment than a longer orientation.

What should I tell an executive who wants one ramp number?

Give them the week your last three hires started sending work out without a review. Something like: the last three hires into this role went out unreviewed at weeks four, eleven and thirty, and the spread is about how much of the job is written down. That is more useful than an industry average because it names a lever. If they insist on a single planning figure, use the median of your own hires in that role and attach the range, because a plan built on the median alone will be wrong in exactly the cases that cost the most.

Should the new hire know which date is being measured?

Yes, and telling them changes the conversation for the better. A hire who knows that the milestone is work going out without a review, rather than a volume target, will ask for the review earlier and push for the check. A hire who thinks the milestone is output will optimise for output, which is now the cheap half. Say it in the first week and repeat it at the first checkpoint.

References

  1. 1. 2012 Allied Workforce Mobility Survey: Onboarding and Retention Allied Van Lines / AlliedHRIQ, 2012. allied.com The provenance of the eight-month time-to-productivity figure: an online survey of 500 HR professionals, with 27% saying a year or more, 25% three months or less, 58% saying their company does not measure it, and a companion self-report of 23 percent of new hires lost before the one-year anniversary.
  2. 2. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) MIT Department of Economics, 2023. economics.mit.edu Supports the claim that short self-contained professional writing got much faster: 444 participants, 10 minutes and 37% faster against a 27-minute control, quality up 0.45 standard deviations, largest gains to the weakest writers.
  3. 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) Harvard Business School, 2023. mitsloan.mit.edu Supports the claim that people cannot tell which side of the capability line a task sits on: 19 percentage points worse, 84.5% control against 60% and 70%, on one task chosen to sit outside AI capability.
  4. 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) METR / arXiv, 2025. arxiv.org Supports the claim that self-reported productivity can be wrong in direction: 16 developers, 246 tasks, 19% slower against a forecast 24% speedup and a retrospective 20% speedup.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.