Screening

Skills-Based Hiring Without a Skills Test Is Degree Screening With Extra Steps

Skills-based hiring was always two moves: remove the credential proxy, then add a way to observe the skill it stood for. Most organizations made the first move and skipped the second, and the hiring data shows it. With every applicant now able to assert every skill in the posting's own words, a policy with no observation step hands the decision to whoever optimized their document best. Build the observation step or keep the proxy.

The takeThe movement's problem was never conviction. Firms that removed degree requirements and then changed nothing about how they evaluate people got the result arithmetic predicts, and the ones that changed the hiring practice alongside the posting moved their numbers. Announcing a policy is free. Running a selection step is work, and the work is where the outcome lives. A commitment with no observation step behind it quietly re-routes the decision to whichever proxy is still standing.

Where Olive fits

Open a role and see what the work shows

Olive is one way to run that observation step without building it: a candidate does the same 40-to-60-minute occupational assignment with an AI assistant, and a human reviewer writes six findings with the moment in the session behind each one. The candidate is granted the identical report, free.

Rank your shortlist

What Did Dropping the Degree Requirement Actually Change?

Postings changed a great deal and hiring changed very little. Burning Glass Institute and Harvard Business School researchers tracked 11,300 roles at large firms before and after a degree requirement came out of the posting, and found the share of hires without a bachelor's degree rose about 3.5 percentage points in those roles 1. Spread across how few roles dropped a requirement at all, that is a rounding error.

Two further numbers carry the argument. About 45% of the firms that dropped a requirement changed hiring in name only, with no meaningful difference in who they went on to hire, and nearly all of the real movement came from the 37% of firms the researchers class as leaders, who raised their share of hires without a bachelor's degree by nearly 20% in the roles studied 1. Across the labor market as a whole, the movement produced new opportunity in not even 1 in 700 hires that year 1.

The distinction the report draws is not between committed employers and cynical ones. It is between firms that changed a posting and firms that changed a hiring practice. Removing a requirement is an edit. Deciding what stands in its place is a project, and the outcome data separates the two cleanly.

That separation is the useful part of the finding for anyone holding the commitment now. It names the failure mode as a missing mechanism, and a missing mechanism can be built. Whether GPA still predicts anything at entry level is the same question one credential down.

Why the Claim Layer Stopped Carrying Signal

Because writing a credible claim is now free, and writing a false one costs the same as writing a true one. A posting that lists eight skills is a template for the application that answers it, and the applicant no longer has to compose that answer. What used to take effort, and therefore carried a little signal, now takes a paste.

Three things used to leak through an application by accident: how well someone writes, how much trouble they took, and what vocabulary the work had given them. All three are available to anyone with a browser tab open, calibrated to the exact posting. An applicant who uses those tools is doing what the posting invites. What changed is the instrument: a document has become a poor way to read skill, the way a tape measure is a poor way to read weight.

So a policy that removes the credential and adds nothing does not open the role to skill. It opens the role to whoever presents best on paper, and presentation on paper still tracks the schooling, the networks and the free evenings the credential requirement was already encoding. The proxy went underground, where nobody has to defend it. Screening for AI skill from a resume runs into the same wall for the same structural reason.

Build the Observation Step the Degree Was Standing In For

One job-shaped task, the same for every candidate at that stage, with a written answer key and a person reading against it. That is the whole specification. It does not need to be long: a short exercise drawn from the actual work predicts about as well as the strongest interview evidence available, at an estimated validity of .33 in the current re-analysis of the literature 2.

Three properties do the work and none of them is expensive. The task comes from the job, so anyone can defend why it is asked. Every candidate at that stage gets the same one, so the comparison means something. And somebody wrote down what a good answer contains before the first submission arrived, so the reading is against a standard rather than against the previous candidate.

Two cautions, both from the same body of research. A job-shaped exercise is not a free pass on fairness: in the 2022 re-analysis, work sample tests carry a mean Black-White subgroup difference of .67 against .23 for structured interviews, which is an argument for combining methods rather than crowning one 3. A subgroup difference is also not adverse impact, which depends on how results are used, on the selection ratio, and on the pool.

Combining two or three genuinely different kinds of evidence is what recovers the accuracy the old single-method numbers promised, with a mechanically combined composite estimated around .61 4. Four conversations are not that. A short work sample plus a structured interview is.

The step is still a minority practice, which is most of the explanation for the outcome gap. In SHRM's 2021 benchmarking of member organizations, 12% of organizations used work sample interviews to assess executive candidates, 11% for middle management and 9% for individual contributors, against in-person interviews at 79%, 78% and 76% 5. What a work sample should test now that a model can produce the work sample is the design question underneath it.

How Do You Tell a Real Change From a Renamed One?

Measure what the policy claimed to change, which is who gets hired, not what the posting says. Pull the share of hires without the credential for the twelve months before the change and the twelve after, in the same roles. Then check that the new step reaches everyone at that stage rather than only the candidates somebody already doubted.

The audit is one table. Share of hires without the credential, in the affected roles, before and after, next to the count of candidates who actually reached the new step. If the second number is small, the step is not being applied, it is being reserved, and a step reserved for the people somebody already had doubts about is a different instrument from the one that was approved.

Swapping the test for a conversation does not move the decision outside the rules either. The Uniform Guidelines on Employee Selection Procedures, the 1978 US federal regulation at 29 CFR Part 1607, define a selection procedure as any measure used as a basis for an employment decision, and say the term runs from paper-and-pencil tests through informal or casual interviews and unscored application forms 6. An unwritten judgment about a candidate's skills is a selection procedure with no documentation attached.

Dropping the resume screen and going straight to a work sample is the version of this some teams have already run. It is a larger change than it sounds, and it is the honest end state of the commitment most firms have already announced.

See a sample report

Common questions

Does this mean the degree requirement should go back on?

No. The requirement was a proxy with a known cost and no direct evidence behind it for most roles. The finding is not that removing it was wrong, it is that removing it was incomplete. Put the observation step in first, then the removal produces the access it promised. Firms whose hiring changed, rather than only their postings, raised their share of hires without a degree substantially in the roles studied.

How long should the observation step take?

Long enough to contain a real decision and short enough that a working candidate can do it in one sitting. Forty to sixty minutes is a workable target for most roles, and anything over two hours needs either payment or a very good reason. Length is not what makes an exercise useful; a written answer key and the same task for everyone are.

Can one task be used across every role?

Not usefully. The point of a job-shaped task is that a person can defend why it was asked of this candidate for this role, which is what federal selection law asks of a selection procedure once it shows adverse impact. One task per job family is the realistic unit. Rewriting the answer key when the job changes matters more than rewriting the task.

What if managers say there is no time to grade a task?

Then the answer key is missing, not the time. Reading a forty-minute exercise against a written key takes ten to fifteen minutes per candidate, and it replaces a resume read plus a screening call that were already happening. If grading feels open-ended, it is because nobody decided in advance what a good answer contains, and that decision is the cheapest part of the whole design.

Does adding an assessment step shrink the applicant pool?

It shrinks the number of people who complete an application and raises the information you hold on each one. Whether that trade is good depends on where the pool is currently lost. A step that arrives after a human has expressed interest costs far fewer candidates than one bolted to the front of an open posting, so place it after the first conversation rather than before it.

References

  1. 1. Skills-Based Hiring: The Long Road from Pronouncements to Practice The Burning Glass Institute and the Harvard Business School Project on Managing the Future of Work (Sigelman, Fuller and Martin), 2024. burningglassinstitute.org Supports the claim that about 45% of firms dropping degree requirements changed hiring in name only, that leader firms (37% of those studied) raised their non-degree share by nearly 20%, that the average shift across 11,300 roles was 3.5 percentage points, and that the movement reached not even 1 in 700 hires.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the claim that a job-shaped work sample is estimated at .33 rather than the widely repeated .54, and predicts about as well as the strongest interview evidence.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection, Table 3 (subgroup differences by selection method) Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, Table 3, 2022. gwern.net Supports the claim that work sample tests carry a larger mean Black-White subgroup difference (.67) than structured interviews (.23), so method choice has demographic consequences.
  4. 4. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300, doi 10.1017/iop.2023.24 (Cambridge University Press), 2023. cambridge.org Supports the claim that a mechanically combined composite of two or three different predictors reaches about .61, more than any single method.
  5. 5. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the claim that work sample interviews remain a minority practice at 12%, 11% and 9% of candidates by level, against in-person interviews at 79%, 78% and 76%.
  6. 6. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that an informal interview and an unscored application form are selection procedures under federal law, so replacing a test with a conversation does not move the decision outside the rules.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.