Assessment design

How Do You Pilot an Assessment Vendor Before Making It a Hiring Gate?

Pilot a hiring assessment vendor in shadow mode. Invite a defined cohort, seal the results from everyone in the hiring loop, and hire exactly as you would have without it. Tell the candidates the step is a trial. Fix the cohort, the number, the criterion measure and the decision rule in writing before the first invite. At pilot scale you get operational answers, not a validity coefficient, and a gate is only one of three endings: input and drop are the others. Whatever you learn holds for that one role.

The takeNobody counts how these pilots die, but the fatal step comes before the first invite. The middling outcome never gets written into the plan, so a quarter's work lands in a room where the only motions on the table are adopt and abandon. A vendor has every reason to sell you a gate, because a gate is the version that renews. At thirty candidates the honest ending is usually a document a hiring manager reads and argues with, and what you're buying is the short list of people your process and the instrument disagreed about. That list is the finding.

Where Olive fits

Open a role and see what the work shows

A shadow pilot is only readable at the end if each result carries evidence rather than a number. Olive returns six separately-evidenced findings per session (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer and anchored to a moment in the work, and the free tier covers ten attempts a month, so a shadow cohort can run beside your existing round before anything is signed.

Rank your shortlist

What does a shadow-mode pilot actually look like?

A shadow pilot runs the assessment on real candidates and lets none of it reach a decision. The cohort gets invited, the vendor returns results, and those results go into a folder nobody on the hiring loop can open. You interview, decide and hire exactly as you would have without it. When the pilot closes, you open the folder and compare.

Three things break it, and all three are procedural rather than technical. A result leaks to someone in the loop, usually on a debrief call where a manager asks how the finalist did. The cohort is picked for convenience, so it fills up with the candidates a recruiter already liked. Or nobody writes down what the pilot was supposed to tell you, and it ends when someone's patience does.

Shadow mode is also the cheapest legal posture you will ever hold with a new instrument. The Uniform Guidelines define a selection procedure as any measure, combination of measures, or procedure used as a basis for an employment decision 1. While the results reach no decision, the assessment is not one. The moment a result changes who advances, it becomes one, and it has to be job-related and consistent with business necessity, a bar the EEOC applies to work samples and simulations by name 4.

Tell the candidates. A pilot people were not told about is a study nobody consented to, and the first candidate who works out that their hour changed nothing will say so in public. One sentence in the invite carries it: this step is being trialed, it will not affect this application, and here is who sees the record.

Log the funnel while you are at it, because the pilot measures your process as much as the instrument. Invite-to-start, start-to-finish, median minutes, and every drop-off with a reason where you can get one. An instrument with a 40% completion rate has already answered the adoption question whatever its results say, which is why testing for AI skills without lengthening the loop is a design constraint and not a nice-to-have.

Decide the sample size before the first invite

Write the number down before anyone is invited, because an open-ended pilot ends the moment the results flatter someone. The Uniform Guidelines put that call on the employer: the number of persons needed for a meaningful criterion-related study is determined by the user, from the selection procedure, the potential sample and the employment situation 2. Pick a number, a window and a stop date, and hold all three.

Then be honest about what that number can carry. A criterion-related validity study is technically feasible only where an adequate sample of persons is available to achieve findings of statistical significance 1, and thirty candidates in one quarter is not that. A pilot at that scale answers operational questions (can two reviewers read the same session and agree, do candidates finish, does the instrument disagree with your process about specific people), and it does not produce a local validity coefficient. Vendors rarely volunteer that distinction. Ask which of the two you are buying.

Six things get fixed in writing before the first invite:

  • The cohort. One occupation, one seniority band, one requisition family. Not "whoever we can get."
  • The number, and the window it has to arrive in. A pilot with no end date is a rollout nobody approved.
  • The criterion measure, and the name of the person who records it.
  • The comparison you will run, written as a sentence, while no result exists to shape it.
  • The decision rule. What result makes this a gate, what makes it an input, what makes it a no.
  • Who may see results during the pilot. A named person, not a team.

Keep the impact counts by race, sex and ethnic group from the first invite; the Guidelines expect a user to hold records disclosing what its selection procedures do to identifiable groups 3. Expect those counts to be uninformative. Where impact evidence rests on numbers too small to be reliable, the Guidelines look instead to impact over a longer period or to the same procedure used in similar circumstances elsewhere 3, which in practice means the vendor's own audit rather than the eighteen people in your cohort. A validated assessment and a bias-audited one are different claims, and a pilot is where you find out which one the vendor actually has.

What criterion do you compare the results against?

Your own performance data, on the people you hired anyway. Choose the measure before any result is visible, and choose one that represents work: the Guidelines say the criteria used should represent important or critical work behaviors or work outcomes 2. Ramp time to independent output, first-quarter deliverables accepted without rework, or a structured manager rating on named behaviors: one of those, defined in writing, with a name attached to it.

Every available criterion is contaminated, and the skill is picking the contamination you can name. An overall manager rating is the easiest to collect and the weakest signal: that manager interviewed the candidate, remembers the debrief, and is rating a person they already decided to hire. Retention measures compensation, commute and a reorg. A structured rating on the same behaviors the assessment claims to measure is the best of a bad set at this size, and it only works if the anchors are written before anyone is rated.

The ceiling on any shadow pilot is that only hires have criterion data. You never observe how the rejected candidates would have done, so the comparison runs on the survivors of the screen the new instrument was supposed to improve. That does not make the pilot worthless; it makes one particular result unavailable. Say it out loud in the plan, because someone will otherwise present a correlation as though it covered everyone who applied. Whether assessment results predict performance at all is not a question a thirty-person cohort settles.

So read the disagreements one by one. The most useful output of a shadow pilot is the short list of candidates where the instrument and your process reached opposite conclusions: the strong hire whose session was thin, the candidate you passed on whose work held up. Ten of those, read by the hiring manager against the evidence behind each one, are worth more than an aggregate you cannot compute at this sample size. That reading is only possible if a result carries evidence rather than a number. Check which you are getting before you sign, not after.

Why doesn't one role's pilot license a rollout?

Because the criterion was one occupation's performance data, and evidence does not travel with the contract. The Guidelines expect a study sample to be representative of the candidates normally available in the relevant labor market for the job in question 2, and the EEOC's own best practice is that a procedure be properly validated for the positions and purposes it is used for 4. A pilot on analysts says something about analysts.

Purposes is the half buyers skip. An instrument piloted as a late-stage input and then deployed as an early-stage cut is a different procedure with a different impact profile, on identical content from the same vendor. Change the stage, the pass rule or the population and you are back to an untested configuration, whatever the first pilot showed.

The rollout question is also a maintenance question. Every role added is another case, another set of rating anchors and another cohort someone has to watch, and one instrument across every department is a rubric question rather than a purchase. Two or three roles, sequenced by where a bad hire costs the most, is a program. Twelve at once is a rollout carrying one role's evidence.

Run the second role as its own pilot, with its own cohort, criterion and stop date. It costs less than the first because the protocol exists and the reviewers are calibrated, and it is the only thing that puts real ground under the third. Where two roles genuinely do the same work on the same material (a pricing analyst and a sizing analyst reading the same kind of packet), write that down as a job analysis and reuse the case. Where they do not, do not.

How do you decide whether to make it a gate?

Apply the rule you wrote at the start, and notice that gate is not the only option. Three outcomes are available: make it a gate, keep it as an input a human reads alongside everything else, or drop it. The middle one is usually right and almost nobody writes it into the plan, so the decision defaults to adopt-or-abandon on evidence that supports neither.

Gating changes what you owe. A result that decides who advances is a selection procedure 1, and it has to be job-related and consistent with business necessity, the bar the EEOC applies to work samples and simulations alike 4. Keeping the result as an input a hiring manager reads is a smaller claim, a smaller exposure, and for most instruments at pilot scale the only description the evidence supports.

Name the failure conditions in advance and let the pilot fail. Completion below a stated rate. Two reviewers reaching different conclusions on the same session. Results that track exactly what your resume screen already tracked, which means you bought a slower version of it. A candidate complaint you cannot answer. Stop when one of those lands, rather than extending the window until it clears.

The vendor's documentation is not your defense either. The EEOC is explicit that the employer is responsible for ensuring its selection procedures are properly validated for the positions and purposes they are used for, whatever the publisher supplied 4. Ask for the technical report, the sample it was built on and the audit before the pilot rather than after, because the questions to put to an assessment vendor are the ones a regulator asks later, in a worse room.

Keep the pilot's own file: the plan as written, the cohort, the criterion, the results, the disagreements, the decision, with dates. It runs to two pages and it is the only thing that will answer, two years from now, why this step exists and what it was tested on. The evidence a vendor should hand you belongs in the same folder.

See what gets scored

Common questions

How many candidates does a pilot need?

More than one role produces in a quarter, if the goal is statistical significance. The Guidelines make an adequate sample for significance a condition of technical feasibility. Set a number anyway, because the number is what stops the pilot drifting. Thirty to fifty sessions in one occupation will tell you whether reviewers agree, whether candidates finish, and where the instrument disagrees with your process about specific people. It will not produce a local validity coefficient, and a vendor who says it will has answered a different question.

Can you run a pilot without telling candidates?

Yes, and the silence costs more than it saves. A candidate who spends an hour on an assessment that changed nothing, learns that afterwards, and posts about it is a worse outcome than the small drop-off disclosure causes. One sentence in the invite does the work: this step is being trialed, it will not affect this application, here is who sees the record. Disclosure also keeps the pilot honest, because a step nobody is allowed to mention is a step that will quietly shape a debrief.

What if the pilot results contradict the hires you made?

That is the pilot working. Take the disagreements one at a time and read the evidence behind each: the strong hire whose session was thin, the candidate you passed on whose work held up. Sometimes the instrument is catching something your process misses; sometimes it is catching test-taking. Ten cases read closely by the hiring manager separate those two better than any aggregate available at this sample size. Record which way each one went before you decide anything about the instrument.

Does a shadow pilot create legal exposure?

Less than the alternative, while it stays in shadow. A selection procedure is a measure used as a basis for an employment decision, and a result nobody in the loop can see is not that. Exposure attaches when results start changing who advances, at which point the step has to be job-related and consistent with business necessity. Keep the impact counts by group from the first invite regardless, keep the plan and the decision in writing, and expect those counts to be too small to be reliable.

Can you skip the pilot if the vendor has published validation?

No, but a published study changes what the pilot is for. Read the technical report and check the sample it was built on against the job you are hiring for and the market you hire from; the employer, not the publisher, carries responsibility for whether a procedure fits the positions and purposes it is used for. Then run the shadow cohort to answer what no published study can: whether your candidates finish, whether your reviewers agree, and whether the results disagree with your process about your own people.

Who should run the pilot?

One owner in people ops, with a named hiring manager as the reader of last resort. The owner holds the plan, the cohort list and the sealed results; the manager reads the disagreements at the end and says which side was right. Recruiters running the live loop should not see results during the pilot, which is easier to arrange than it sounds and is the single control the whole design rests on.

References

  1. 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.16 (Definitions) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov A selection procedure is any measure, combination of measures, or procedure used as a basis for an employment decision; technical feasibility requires an adequate sample of persons available for the study to achieve findings of statistical significance.
  2. 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 (Technical standards for validity studies) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Section 14B(1): the number of persons necessary for a meaningful criterion-related study is determined by the user from the selection procedure, the potential sample and the employment situation. 14B(3): criteria used should represent important or critical work behaviors or work outcomes. 14B(4): the sample should be representative of candidates normally available in the relevant labor market for the job in question.
  3. 3. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.4 (Information on impact) U.S. Equal Employment Opportunity Commission (eCFR), 1978. ecfr.gov Users should maintain records disclosing the impact of their selection procedures by identifiable race, sex or ethnic group; where impact evidence rests on numbers too small to be reliable, impact over a longer period or the same procedure used in similar circumstances elsewhere may be considered.
  4. 4. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2007. eeoc.gov Work samples and simulations are listed as employment tests and selection procedures; a procedure with disparate impact must be job-related and consistent with business necessity; employers should ensure selection procedures are properly validated for the positions and purposes for which they are used, and remain responsible even where a vendor supplies documentation.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.