Pipeline

Is It Cheaper to Assess Every Applicant Than to Screen Resumes?

Assessing every campus applicant up front usually costs more than screening resumes. It pays only when a completed assessment costs less than one resume read divided by the extra share who finish once everyone is invited. At a median US HR wage and a three-minute read, that ceiling is near four dollars, near one for a fast screen. Anything a person reads exceeds both; a small track whose screen passes most applicants can afford one. Campus volume needs a cheap machine-marked first step, then the long assessment on its shortlist.

The takeThe cost fight is a distraction from the argument that matters. A completion rate is not a conversion number. It is a filter on who had a free evening, and a free evening is what a term-time job and a caregiving shift take away. Nobody has measured what that costs a campus pool at scale, which is exactly why it keeps getting priced at zero in these models. An assess-first funnel can clear the arithmetic and still hand you the narrowest pool you have ever run. A funnel that saves money by quietly losing the students you built it to reach has not saved anything.

Where Olive fits

Open a role and see what the work shows

Olive is priced per attempt rather than per seat, and ten attempts a month cost nothing, so the assess-first question can be tested on one campus role before it is costed across a whole pool. An attempt returns six evidenced findings on one candidate, each written by a human reviewer, as an input to your decision rather than a decision itself.

Rank your shortlist

What does the break-even arithmetic look like?

Assess-first is cheaper when one completed assessment costs less than a single resume read divided by the extra share of the pool you have just added. That divisor is the trap. Cutting a screen saves recruiter minutes on every applicant, but assessing everyone spends real money on the nine applicants in ten the screen used to stop for free.

Five numbers run the model:

  • t: minutes a recruiter spends on one campus resume.
  • w: that recruiter's cost per minute.
  • s: the share of applicants who reach an assessment under your current process.
  • p: the share who complete one when everybody is invited.
  • c: everything one completed assessment costs, including the time someone spends reading its output.

Assess-first is cheaper when c falls below t times w, divided by p minus s. Nothing downstream of the shortlist enters the comparison, because it happens either way.

Now put money in it. Human resources specialists, the people who read campus resumes, had a median wage of $36.51 an hour in the 2025 federal wage data, or about 61 cents a minute before benefits and overhead 1. A careful three-minute read of one campus resume against a rubric therefore costs roughly $1.83 in wages. Shortlist 10% of applicants today, get 55% of invited students to finish an assessment, and the divisor is 0.45. The ceiling on c is $4.07 per completed assessment.

Speed the screen up and it gets worse, not better. A one-minute triage pass costs 61 cents; at 70% completion the divisor is 0.60 and the ceiling falls to about $1.02. The faster your screen already is, the less there is to save by deleting it.

Neither ceiling accommodates an assessment a human being reads, and most automated ones breach it too once the reading at the other end is counted. On cost alone, assessing every campus applicant loses by a multiple rather than a margin.

How much does the resume screen actually buy?

Less than its cost implies, which is the real argument for inverting the funnel, not the money. The revised operational validity estimates put job-specific measures at the top of the list: structured interviews at .42, job knowledge tests at .40, work samples at .33, general mental ability at .31 2. A campus resume is none of those. It is a document a student wrote about themselves, or asked a model to write about them.

Look at what the screen claims to read for and the mismatch is plain. In a NACE employer survey of 237 respondents, nearly 90% said they look for evidence of problem solving on a new-grad resume and nearly 80% for teamwork 6. Neither is a thing a page can carry. A bullet asserting problem solving describes an event nobody in the room witnessed, in words that are now free to produce.

So the case for assessing early is a signal case, not a cost case, and that distinction changes what you should do about it. If the screen were expensive and useless, the answer would be a cheaper screen. Most campus screens are cheap and useless, which means deleting one saves almost nothing and the money for an assessment has to come from somewhere else.

The screen does buy two real things, and both are worth keeping. It enforces hard requirements (graduation date, work authorization, location, a program the role genuinely needs) at a cost per applicant nothing else matches. And it produces a defensible record of why an applicant went no further. Replacing the screen with a work sample means moving both jobs somewhere else rather than dropping them. Before deciding the screen is noise, test the specific thing you rely on it for: whether GPA still predicts anything at entry level is a separate question with a different answer by field.

What happens to completion when the task comes before a reply?

Completion drops, and it drops unevenly. A thirty-minute task asked before any human has replied is a request for unpaid work from someone with applications open everywhere else. The share who finish, p in the model above, is the single number that decides whether assess-first pays, and it is the one number no published dataset can give you for your own discipline.

Measure it in one cycle instead of arguing about it. Take a single role, invite a random fifth of its applicants straight to the assessment, and run the rest through the normal screen. Record three states rather than two: invited, started, completed. An invitation nobody opens and a session abandoned at minute twelve are different failures with different fixes, and averaging them into one rate hides which one you have.

Then break p out by the things that predict a student's free time. Term-time employment, commute, caregiving, a shared laptop, a course load with a deadline the same week. A completion rate that moves with any of those is a pool selecting on money and availability rather than on the work, and the students it drops are the ones campus recruiting exists to reach.

Discipline matters more than most cost models allow. A computing applicant has done three take-homes this term and treats a fourth as routine. A marketing or humanities applicant has done none, and thirty minutes of unpaid work before a human replies reads as an insult rather than a step. The same c and the same t produce different answers in those two funnels, because p is not the same number.

Pilot the assessment before it becomes a gate, and hold your own pilot to the evidence you would demand from a vendor: a measured completion rate on your applicants, broken out by the groups you already report on.

Which costs appear only after you invert the funnel?

Three, and none of them sit in the price per assessment. Somebody has to read the output. The assessment becomes your only gate, so its adverse impact is no longer diluted by anything else in the process. And an assessment everyone takes is an assessment everyone can request an accommodation for, at a volume a shortlist never produced.

Start with reading, because it is the largest and the most often omitted. Federal assessment guidance is blunt: work samples may be costly to develop in both time and money, may be time consuming and expensive to administer, and require individuals to observe and sometimes rate applicant performance, and they are best used where a limited number of applicants is being tested 3. That is the government describing, in advance, the exact thing assess-first proposes to stop doing. If nobody reads the output, you have not inverted the funnel; you have swapped a human screen for an automated one, which is a different product with a different legal posture.

Then the impact arithmetic. The Uniform Guidelines expect an employer to keep records showing the impact of its selection procedures by sex and by identifiable race and ethnic group, treat a selection rate below four-fifths of the highest group's rate as evidence of adverse impact, and expect individual components to be examined where the total process shows impact 4. When the assessment is the total process, there is no bottom line to shelter behind, and the ratio you owe is computed across every applicant rather than a shortlist.

Work samples generally show little or no performance difference by sex or race depending on the competencies assessed, and applicants tend to perceive them as fair 3. Assessing everyone can be the more defensible design; it is simply not the cheaper one. Budget for the accommodations too, because what an AI assessment owes under the ADA does not scale down as volume goes up.

Where does inverting the funnel actually pay?

In three places: a pool small enough that the screen is noise, a role where the screen is demonstrably wrong, and a two-step shape where the universal step is cheap and the expensive one stays on a shortlist. The last is the only one that survives campus volume, and it is not really assess-first. It is screening on something better than a resume.

The small-pool case is real and undersold. Forty applicants for a specialised track, at $30 a completed assessment, is $1,200, about three days of a recruiter's loaded time at the $36.51 median wage 1 plus benefits and overhead, and it produces evidence a resume cannot. The formula still applies; it simply clears, because s is already high and the pool is small. Run the numbers before assuming your funnel is too big, since plenty of campus tracks are forty people rather than four hundred.

The volume term is moving, and not in your favor. Payroll data through June 2026 puts employment of 22-to-25-year-olds in the most AI-exposed occupations about 19% below where it would have been had it kept pace with less-exposed peers, arriving through reduced hiring rather than separations 5. That measures openings, not applications. The applications side is yours to count, and what to do when applications per opening have tripled decides how much of this model you can afford at all.

So build the two-step. Make the universal step something with a fixed key and no human reader: a short structured question set marked against an answer sheet, or a five-minute exercise where the right answer is checkable. Put the longer assessment on the shortlist that step produces, where c can be a real number because the volume is small. Then add the deeper signal without lengthening the loop by replacing a round rather than appending one: a stage you added is a stage a candidate can drop out of, and p applies to that stage too.

See a sample report

Common questions

What is the break-even cost per assessment?

Divide what one resume read costs by the extra share of the pool you would now be assessing. At a median US HR wage of about 61 cents a minute, a three-minute campus read costs $1.83; shortlist 10% today, get 55% of invited students to finish, and the ceiling is about $4.07 per completed assessment. A one-minute screen at 70% completion puts it near $1.02. Both figures have to cover the time someone spends reading the output, which is where most assess-first models quietly fail.

How do you measure completion before committing to it?

Run one role as a split. Invite a random fifth of applicants straight to the assessment and send the rest through the normal screen. Record invited, started and completed as three separate counts, because an unopened invitation and an abandoned session need different fixes. Then break the completion rate out by the groups you already report on, and by discipline, because a computing cohort and a marketing cohort will not hand you the same number. One cycle replaces the assumption with a measurement.

Does assessing everyone make the process fairer?

It can, and it is not automatic. Work samples generally show little or no performance difference by sex or race depending on what they measure, and candidates tend to perceive them as fair. But an unpaid thirty-minute task before any human replies selects on free time, and free time tracks money. Once the assessment is your only gate, its four-fifths ratio is computed on every applicant with nothing downstream to dilute it. Measure the completion rate by group before claiming the fairness benefit.

Should the resume screen be dropped entirely?

Rarely. The screen still enforces hard requirements (work authorization, graduation date, location, a required program) at a cost per applicant nothing else matches, and it produces a record of why an applicant went no further. Dropping it moves both jobs somewhere else rather than removing them. The stronger move is to stop asking the screen to predict performance, which it does badly, and keep it for the eligibility checks it does cheaply and well.

Is there an assessment cheap enough to give every applicant?

Not one a person reads. A fixed-key exercise marked by machine can sit at the top of the funnel; anything with a human reviewer belongs on a shortlist, where the volume keeps the cost per candidate survivable. Olive is priced per attempt rather than per seat, with ten attempts a month at no cost, so the shape can be tested on one role before it is costed across a campus pool. A human reviewer writes every finding, and the candidate is granted the same report the employer reads.

References

  1. 1. Human Resources Specialists (13-1071.00) O*NET OnLine, U.S. Department of Labor, 2026. onetonline.org Median wages for Human Resources Specialists, whose tasks include recruiting, screening and interviewing: $36.51 hourly and $75,940 annually, from Bureau of Labor Statistics 2025 wage data. Page text verified 2026-08-24.
  2. 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Sackett, Zhang, Berry and Lievens, Industrial and Organizational Psychology (Cambridge University Press), 2023. doi.org Revised operational validity estimates: structured interviews .42, job knowledge tests .40, work samples .33, general mental ability .31; job-specific measures outperform general measures of a candidate's attributes.
  3. 3. Assessment and Selection: Work Samples and Simulations U.S. Office of Personnel Management, 2024. opm.gov Work samples may be costly to develop in both time and money, may be time consuming and expensive to administer, and require individuals to observe and sometimes rate applicant performance; they are best used where a limited number of applicants is being tested; subgroup performance differences are generally little or none depending on competencies assessed, and applicants perceive them as very fair.
  4. 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 - Information on impact U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Users should maintain records disclosing the impact of their selection procedures by sex and by identifiable race and ethnic group; a selection rate below four-fifths of the rate for the highest group is generally regarded as evidence of adverse impact; where the total selection process shows adverse impact, individual components are evaluated.
  5. 5. Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence Brynjolfsson, Chandar and Chen, Stanford Digital Economy Lab, 2026. digitaleconomy.stanford.edu ADP payroll data through June 2026: employment of 22-to-25-year-olds in the most AI-exposed occupations sits about 19% below where it would be had it kept pace with less-exposed peers, arriving through reduced hiring rather than separations.
  6. 6. What Are Employers Looking for When Reviewing College Students' Resumes? National Association of Colleges and Employers (NACE), Job Outlook 2025, 2024. naceweb.org Of 237 employer respondents, nearly 90% look for evidence of problem-solving skills on a new-grad resume and nearly 80% for teamwork.

6 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.