Assessment design

Personality Explains a Few Percent, and Candidates Can Move It

Personality tests do not predict job performance well enough to screen on. Take the figure vendors quote for conscientiousness at face value, around .2 to .3, and square it: that is four to nine percent of the variation in performance, from a questionnaire whose preferred answers are obvious to anyone who wants the job. Personality earns a place after a job-related assessment, as a source of interview questions rather than a source of decisions, and never as the stated reason for a rejection.

The takeThe category is sold as a screen and is worth having as a prompt. That is not a small demotion. A screen sets a cutoff and produces rejections, while a prompt produces better questions in a round you were running anyway. Most of the trouble teams get into with personality testing comes from the first use, and almost none of the value they report comes from it either. Keep the instrument, move it behind the work, and take the cutoff off it.

Where Olive fits

Open a role and see what the work shows

Olive reads what a candidate did on a role-grounded assignment, and a human reviewer writes each of six findings with the timestamped excerpt it rests on. Outcomes are demonstrated, partly demonstrated or not demonstrated, so there is no number for anyone to set a cutoff on.

Rank your shortlist

Should a personality score gate anyone?

No. Use it after a job-related assessment, never as the gate in front of one, and never as the sentence in a rejection. The reason is not that personality is meaningless. It is that the claim a questionnaire supports is much smaller than the claim a cutoff makes. A cutoff says this person cannot do the job. A trait score says this person described themselves a particular way on a Tuesday.

Federal selection law does not exempt the category. The Uniform Guidelines define a selection procedure as any measure used as a basis for an employment decision, covering the full range of assessment techniques through informal or casual interviews and unscored application forms 1. A personality inventory used to cut a list is a selection procedure, carrying the validation burden that follows if it shows adverse impact.

The disability exposure is sharper, because it arrives without any statistics at all. The Justice Department's 2022 ADA guidance states that employers violate the ADA if their hiring technologies unfairly screen out a qualified individual with a disability, and puts that duty on the employer using the tool rather than on the vendor that built it 2. Trait inventories ask about mood, energy, sociability and stress. Those questions do not mean the same thing to every candidate, and a cutoff applied to the answers is where the exposure starts.

That leaves the instrument intact and rules out one use of it: the automatic cut, applied at volume, with the score as the stated reason. Every other use in this article survives.

What does a validity of .2 to .3 actually buy?

A few percent of the variation in performance, on the vendor's own best case. Take the range typically quoted for conscientiousness, square it, and the answer is four to nine percent. That is not nothing across thousands of hires at a large employer. It is close to nothing when the question is which of two finalists gets the offer, which is the question a hiring team actually has in front of it.

There is a discount to apply before using any published figure. The 2022 re-analysis of the selection literature found that prior meta-analyses had systematically overcorrected for range restriction, and cut mean validity estimates by .10 to .20 points for most of the methods that had ranked high in earlier summaries 3. Any coefficient quoted from an older meta-analysis inherits that problem. Ask which paper a number comes from and what year it was estimated in, then check it against the corrected table.

Correcting the number does not fix the instrument. Nothing in a personality questionnaire hides which answer an employer prefers, and the person filling it in wants the job. Nobody is lying. The scores are self-descriptions produced under an obvious incentive, which is a different measurement from the same questions answered by an employee with nothing riding on the outcome.

Self-report and demonstration are simply different instruments. In one study, 288 teachers took both a self-report and a knowledge-based test of AI literacy built on the same framework, and correlations between the objective and self-reported factors ran from 0.07 to 0.24 4. Different construct, different population, and it says nothing about personality specifically. What travels is the shape: asking someone to describe themselves and watching them do something are not two routes to one number.

Where a personality profile earns its place

Behind the work, as a question generator. Run the job-related assessment first, read it, form a view, and only then open the personality report to decide what to probe in the interview. A profile flagging low tolerance for ambiguity becomes a follow-up question about a project that changed direction twice, and the answer to that question is the evidence. The profile is not.

That sequence matters more than it sounds. Read first, the profile becomes a lens colouring everything after it, and interviewers hunt for confirmation. Read last, it is one more input into a round that was already scored on job content. The same document does different work depending on where it sits, and only the second position produces something defensible in a debrief.

Uses that hold up: choosing which two of six behavioural questions to spend real time on, deciding what to check in a reference conversation, and planning the first ninety days of a hire already made. That last one is the most defensible use in the whole category, because nobody is being excluded on it and the person can see their own report.

Uses that do not: a minimum score to reach the interview, a culture-fit composite, a team-shape argument for preferring one finalist, and anything where the profile is what gets quoted back to a candidate. If a rejected applicant would find the reason unrecognisable as a description of their own work, the reason is in the wrong place. That test separates the assessment categories generally, worked through in what each category can claim.

Ask the vendor for a local validation

Ask for a validation study on roles like the ones being filled, in an applicant sample, with the criterion named. Not a technical manual, not a norm table, not a list of client logos. If what comes back is validity evidence from a different occupation on existing employees, what is being bought is the norm group's performance, with a hope that it transfers to a different job in a different labour market.

Four questions do the work in one call. Which specific roles was this validated on, and how close are they to the one being filled? Was the sample applicants or incumbents, given that incumbents already survived the selection being validated? What was the criterion, supervisor ratings or something harder? And has adverse impact been examined for this instrument at the cutoff being proposed?

A validated instrument and a bias-audited one are not the same claim, and vendors sometimes offer the second when asked for the first. The difference between them is worth having straight before procurement starts, because only one of the two speaks to whether the test predicts anything at all.

If the answers come back thin, there is a cheaper path than arguing about them. Keep the inventory out of the decision, spend the same money on the round already being run, and hand the profile to the interviewer as reading material. That costs nothing to defend, because nothing was decided on it, and it preserves the one use of the category that has never been in dispute.

See what gets scored

Common questions

Are personality tests legal to use in hiring?

Generally yes in the US, with two live constraints. A personality inventory is a selection procedure under the Uniform Guidelines, so if it produces adverse impact the employer carries a validation burden. And under the ADA, an instrument that functions as a medical inquiry, or that screens out a qualified individual with a disability, creates separate exposure regardless of validity. State law adds more in places. The register here is public legal fact, not advice: run the specific instrument and the specific cutoff past counsel before it gates anyone.

What about integrity tests, which look similar?

They sit in the same family and behave differently in one respect worth knowing. In the 2022 re-analysis of the selection literature, integrity tests carry a mean Black-White subgroup difference of .10, among the lowest of the methods in that table. A low difference does not make a method fair or good on its own, and unstructured interviews prove it, sitting at .32 with a validity of .19. Treat integrity testing as its own purchase with its own evidence, not as personality testing under another name.

Can candidates really game a personality questionnaire?

They do not need to game it. The items state their own direction, so a candidate answering the way the job appears to want is not cheating, only responding to an obvious incentive. That is why the same person can produce a different profile as an applicant than as an employee, and why validation evidence gathered on incumbents does not automatically describe applicants. Forced-choice formats and consistency indices reduce the effect rather than removing it.

Should candidates be told their personality results?

If the result is going anywhere near a decision, yes. Sharing it is the fastest test of whether the instrument is being used honestly: a finding nobody is willing to show the person it describes is a finding that should not be carrying weight. Sharing also produces useful corrections, since candidates routinely explain a profile in a sentence that no rating scale captured. Practically, share it after the decision rather than during, so it does not become a negotiation.

Does any of this change now that candidates use AI?

Less than for other assessment categories, and not in a good way. A personality questionnaire was already a self-report with obvious preferred answers, so an assistant helping someone answer it changes little about a measurement that was never observing behaviour. What has changed is the comparison: job-related exercises now have to be designed around an assistant, which raises their cost, and the temptation is to lean harder on the cheap self-report instead. That is the wrong direction.

References

  1. 1. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that a personality inventory used to cut a candidate list is a selection procedure under federal selection law.
  2. 2. Algorithms, Artificial Intelligence, and Disability Discrimination in Hiring U.S. Department of Justice, Civil Rights Division (ada.gov), 2022. ada.gov Supports the claim that the employer using a hiring technology, not the vendor, carries the ADA duty not to screen out a qualified individual with a disability.
  3. 3. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .10 to .20 downward correction that any pre-2022 validity figure inherits, and the integrity test subgroup difference in the FAQ.
  4. 4. How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures arXiv (Zhang, Xiao, Botelho, Liao, Chiu, Stamper, Koedinger), 2026. arxiv.org Supports the narrower claim that a self-report instrument and a demonstration instrument do not measure the same thing, cited with its construct and population stated.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.