Assessment design
What Should an Early-Career Assessment Measure, and Is Yours an Aptitude Test?
An early-career assessment should measure the work behaviors of the job you're filling: the tasks, the judgment calls, the tools a first-year hire uses this quarter. Not general aptitude. To tell a real instrument from a repainted aptitude test, ask three things in writing. Which occupation the content was built from. When that occupation's task list last moved. What evidence sits behind each rating. A test that answers all three the same way for every role is measuring a trait and selling it as a job.
The takeEvery one of these questions is really a question about you. A vendor can only build from an occupation if somebody on your side can say what a first-year hire does on a Tuesday, and in plenty of campus programs nobody has written that down since the last reorganization. That gap is what a trait battery gets sold into. It arrives with a number, a norm group and no argument you could check even if you wanted to, which is the appeal. The instrument you can still defend a year from now tends to be the one you had to help build.
Where Olive fits
Open a role and see what the work shows
Olive's item banks are authored one occupation at a time, each carrying its SOC code, and a released report is six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a moment in the session. There is no composite figure, a human reviewer writes every word, and the candidate is granted the same document.
Rank your shortlistWhat Should an Early-Career Assessment Measure?
The work behaviors of the specific job, sampled closely enough that doing well on the assessment and doing well in the first year are nearly the same act. The federal standard is blunt about it: a selection procedure is supported by content validity to the extent that it is a representative sample of the content of the job 1. Everything else is a proxy, and early-career hiring is where proxies fail hardest.
They fail hardest here because the candidate has no record to correct them. A senior hire arrives with five years of outcomes you can reference-check. A graduate arrives with a transcript, two internships and a portfolio anyone's assistant could have produced. Whatever the assessment measures becomes most of what you know, so measuring the wrong thing costs more at this level than at any other.
The updated validity evidence points the same way. In the re-analysis that revised the field's headline numbers downward, structured interviews came out highest at .42 while cognitive ability fell to .31, and the authors say plainly that cognitive ability is no longer the stand-out predictor it was in the prior work 2. Job-specific measures beat general ones. The 80% credibility interval on structured interviews still runs from .18 to .66 2, which is the honest caveat: format alone does not rescue a badly built instrument.
So the target is a list of behaviors, not a trait. For a junior analyst that might be choosing a comparable, noticing that two quotes are written on different terms, and deciding which claim in a source packet actually has to be opened. For a junior developer it starts with reading unfamiliar code before changing it. Those lists are field-specific, which is why one instrument sold into every function deserves suspicion. If your current screen is an online assessment a model solves in thirty seconds, the problem is not the difficulty setting.
Why a Repackaged Aptitude Test Can't Clear the Bar
Because it cannot borrow the cheap validation strategy. The Uniform Guidelines state that a content strategy is not appropriate for demonstrating the validity of selection procedures which purport to measure traits or constructs, such as intelligence, aptitude, personality, commonsense, judgment, leadership, and spatial ability 1. An aptitude test therefore owes you criterion or construct evidence, which is expensive, job-specific, and almost never in the deck.
The same section tells you what a real one looks like from the outside, without a psychometrician in the room. The closer the content and the context of the selection procedure are to work samples or work behaviors, the stronger the basis for content validity; as the setting and manner of administration less resemble the work situation, the weaker it gets 1. That is a sliding scale you can apply during a demo. Watch the screen and ask how much of what a first-year hire does on a Tuesday appears on it.
Repackaging shows up in three tells. The content is identical across roles with only the cover page changed. The output is a single number, or a percentile against a norm group of other test-takers rather than against the job. And the report explains what the number means about the person instead of what the candidate did.
All of this leaves aptitude testing lawful. It makes it a general measure carrying a job-relatedness burden the vendor has to discharge with evidence, and the EEOC's stated position is that tests and other selection procedures should be properly validated for the positions and purposes for which they are used 6. The load-bearing word is *positions*. A study run on inbound call-center hiring is not evidence about your rotational finance program. This is also where two documents get confused, because a bias audit and a validation study answer different questions.
Ask the Vendor These Questions
Send them in writing before the demo, because a demo is answered by a salesperson and an email is answered by someone who has to be right. Each question has a version a repainted aptitude test cannot answer without changing the subject. Grade the replies on specificity: an occupation named, a date given, a document attached.
- Which occupation is the content built from, and what is its SOC code? Worrying: "it works for any knowledge-worker role." An assessment with no occupation behind it has no job analysis behind it, and the job analysis method is an essential element of the federal documentation standard 4.
- When was that job analysis done, and when did the content last change? Worrying: no date at all. The Guidelines make the dates and locations of the job analysis an essential part of the validity report 4.
- What does the report show a hiring manager, and what does it show the candidate? Worrying: two different documents, or nothing for the candidate.
- What sits behind each rating: a rubric, a marked artifact, a transcript excerpt? Worrying: "proprietary model." You cannot defend a rating whose evidence you have never seen.
- Who marks it, and how do two markers agree? Worrying: an agreement figure averaged across dimensions instead of reported per dimension. Averaging hides the one row that is noise.
- What is this assessment not measuring? Worrying: a vendor who cannot name a limit. Every real instrument has a short list of things it does not see.
- Which roles like ours has it been validated for, and can we read the report? Worrying: a badge instead of a report. The standard expects a named contact person and a description of the steps taken to assure accuracy and completeness 4.
- What happens to a candidate who needs an accommodation? Worrying: "extra time" as the entire answer, with no alternative format and no named process.
Keep the replies. A procurement file with eight written answers in it is the document you will want the first time a candidate asks how their result was produced. The same evidence question applies to anything you build yourself, including whether a take-home still tells you something when candidates use AI.
How Recently Did the Benchmark Move?
Ask for a date, and treat a missing one as a no. Occupational content does move, and there is a public record of it: O*NET updated an average of 843 occupations a year between 2017 and 2025, with 891 updated year to date through May 2026 5. An assessment built on a task list nobody has revisited in four years is testing a job that has partly stopped existing.
Recency is not a vanity metric, and the Guidelines treat it as substantive. On currency they say there are no absolutes, and that all circumstances concerning the study, including the validation strategy used and changes in the relevant labor market and the job, bear on when a validity study is outdated 3. Read that as an instruction rather than a hedge: if the job changed, the evidence aged, whatever the copyright line says.
Early-career work is the fastest-moving case. The tasks a first-year analyst or junior developer was hired for two years ago are the tasks most likely to be partly generated now, which changes what a strong first year looks like rather than removing it. That shift is why general academic proxies have lost ground for entry-level roles while job-specific measurement has held.
The practical version of this question is to ask for the diff. Which items changed in the last twelve months, and what prompted each change? A vendor with a live job analysis answers in a sentence and offers to show you. A vendor selling a norm-referenced trait battery answers about their norms, because the norms are the only thing that ever gets refreshed.
What Should the Report Show You?
Evidence you can point at, per thing measured, plus a plain statement of what the assessment does not cover. A number with nothing underneath it cannot be defended to a candidate, to a hiring manager, or to counsel. The federal documentation standard is a useful floor: it treats the job analysis method, the dates and a named contact person as essential, and calls for the steps taken to assure accuracy and completeness 4.
Three properties separate a report you can use from a certificate. Each rating names the behavior it is about rather than a trait. Each rating carries the excerpt, artifact or timestamp it rests on. And the dimensions stay separate, because collapsing them into one figure destroys the only information a hiring manager needs, which is the part that was strong and the part that was not.
Ask what the candidate receives. An early-career pipeline runs on reputation, and a process that rates a graduate on something they never see is one screenshot away from becoming your campus story. It is also the cheapest test of a vendor's confidence: a report written to be read by the person it describes contains no sentence the vendor would not defend out loud.
Then pilot it before it gates anyone. Give it to a cohort you have already hired and know well, keep the results out of every live decision, and check whether the findings recognize the people you know. That is not a validation study and should never be described as one. It is the smallest honest check that the instrument measures your job rather than a general trait, and it costs a quarter.
Common questions
Is a cognitive ability test still worth using for entry-level hiring?
Sometimes, as one input, and rarely as the whole instrument. The updated meta-analytic estimates put cognitive ability at .31 and structured interviews at .42, and the authors state that cognitive ability is no longer the stand-out predictor it was in earlier work 2. It also carries the heavier evidentiary burden: content validity is not an available strategy for a procedure that purports to measure aptitude or intelligence, so the vendor owes criterion or construct evidence for your positions 1. If it is the only thing in your funnel, you are resting the decision on the weaker half of the evidence.
What's the difference between a validated assessment and a bias-audited one?
They answer different questions. A bias audit compares selection or scoring rates across demographic groups. A validation study argues that the results relate to performance in a specific job, which is why the federal documentation standard is built around job analysis, dates and the study's own method 4. A vendor can hold a clean audit and no validity evidence at all for the occupation you are hiring into. Ask for both documents separately, and never accept one as a substitute for the other.
How old is too old for a vendor's validity study?
Shelf life is not fixed anywhere, and the Uniform Guidelines say so directly: there are no absolutes in determining currency, and all circumstances, including the validation strategy and changes in the relevant labor market and the job, bear on when a study is outdated 3. The workable test is whether the job moved. O*NET updated an average of 843 occupations a year between 2017 and 2025 5. If the occupation's task list changed and the study did not, treat the study as describing a job you are no longer filling.
Does an early-career assessment need to be different from one for experienced hires?
The measurement target is the same, the work behaviors of the job, but the stakes differ. An experienced hire brings a record that can correct a bad assessment; a graduate does not, so the assessment becomes most of what you know. The content also has to sample the first year specifically rather than the job at large. An instrument built from a senior task list rewards knowledge a new hire is expected to learn on the job, which the Guidelines put outside content validity anyway 1.
What should we ask a vendor to send before signing anything?
The validity report itself, not a summary of it. The federal documentation standard names as essential the method used to analyze the job, the dates and locations of that analysis, a full description of the work behaviors and their measured importance, and a named contact person; it also calls for a description of the steps taken to assure accuracy and completeness 4. Add two of your own: a sample candidate report exactly as the candidate receives it, and the list of items changed in the last twelve months. A refusal is information too, so record it in the procurement file.
References
- 1. 29 CFR 1607.14 - Technical standards for validity studies ecfr.gov Content validity requires a representative sample of the content of the job; it is not appropriate for procedures purporting to measure traits or constructs such as intelligence, aptitude and personality, or for knowledge an employee will learn on the job; closeness to work samples strengthens the case.
- 2. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors doi.org Corrected operational validity estimates: structured interviews .42, cognitive ability .31, work samples .29; cognitive ability is no longer the stand-out predictor, and the 80% credibility interval on structured interviews runs .18 to .66.
- 3. 29 CFR 1607.5 - General standards for validity studies ecfr.gov Section 5K on currency: there are no absolutes, and all circumstances including the validation strategy and changes in the relevant labor market and the job bear on when a validity study is outdated.
- 4. 29 CFR 1607.15 - Documentation of impact and validity evidence ecfr.gov Essential elements of a validity report: dates and locations of the job analysis, the method used to analyze the job, description of work behaviors and their importance, and a named contact person; the report should also describe the steps taken to assure accuracy and completeness, which the section does not mark essential.
- 5. O*NET Occupation Update Summary onetcenter.org An average of 843 O*NET occupations were updated yearly between 2017 and 2025, and 891 occupations were updated year to date through May 2026.
- 6. Employment Tests and Selection Procedures eeoc.gov Employers should ensure that employment tests and other selection procedures are properly validated for the positions and purposes for which they are used.
6 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.