Assessment design

An Assessment Does Not Need to Live in Your ATS

For most teams, an assessment does not need to integrate with the ATS. At ten candidates a week, native integration saves a few minutes of copying a link and attaching a report, which is not worth narrowing your choice of instrument. Past a few thousand a quarter, or with several coordinators working one queue, it turns into a real requirement. Either way, ask what it deposits in the candidate record: a number that lands in a file is part of the decision trail whether or not anybody read it.

The takeIntegration is a procurement habit dressed as a requirement, and it settles purchases before anyone examines what the instrument measures. The question underneath it is whether this creates manual work for coordinators, which is a workflow problem with a workflow-sized answer. Ask a vendor to demo the integration, then ask what it writes. If the honest answer is a number in a field, you have bought a filing convenience and taken on a record you may have to account for.

Where Olive fits

Open a role and see what the work shows

If you are building this in-house, the hard parts are the answer key and the evidence trail. Olive ships twelve item banks, each grounded in one occupation, and returns six separately-evidenced findings anchored to moments in the session rather than to a number, delivered by link with nothing to install.

Rank your shortlist

What would integration actually save?

About two minutes per candidate, which is what a link-out workflow costs a coordinator: paste the invite link into the outreach, attach the returned report to the record. At ten candidates a week that is twenty minutes, and across two hundred candidates a quarter it is under seven hours. Count the steps it removes, then multiply, and weigh the total against the instruments an integration requirement takes off the table.

What native integration usually buys is narrower than the word suggests. Status sync, so the stage updates when a candidate finishes. A link on the record. Sometimes a completion event, and sometimes a result field. What it does not carry is the reasoning, because the reasoning lives in a document a person still has to open and read.

There are volumes where the arithmetic flips. Thousands of candidates a quarter, several coordinators working the same queue, a regulated process that requires everything to sit in one system of record, or a service-level commitment on how fast a candidate moves between stages. If one of those describes you, integration is a real requirement and it is worth paying for. If none of them does, it is a preference wearing a requirement's clothes.

The cost of treating it as a bar is predictable, if not measured: an integration catalogue tends to follow the volume. SHRM's benchmarking found work sample interviews used to assess 12% of executive, 11% of middle management and 9% of individual contributor candidates, against structured interviews at 34%, 37% and 36% 3. Those are self-reported figures from 2021, from a random sample of SHRM member organizations rather than of US employers, and employers routinely claim more structure than they run, so read them as an upper bound. Either way, screening vendors by integration support steers you toward what most employers already do rather than toward the evidence you are missing.

What does an integration deposit in the record?

Usually a status, a link, and sometimes a number. The status and the link are harmless. The number is the one to think about, because a value written into a candidate record becomes part of the decision trail whether or not the recruiter opened it, and an unused score sitting in a file is harder to account for afterwards than no score at all.

Nothing theoretical about this, and nothing about how documents get read. In Case C-634/21, decided 7 December 2023, the Court of Justice of the EU held that a credit agency's automated establishment of a probability value about a person is itself automated individual decision-making under GDPR Article 22(1), where a third party receiving that value draws strongly on it to establish or terminate a contractual relationship 1. That is a credit-scoring case, no court has applied it to a candidate result, and the phrase draws strongly on is deliberately fact-specific. What travels is the structure: the party producing the number can be the party deciding, even when somebody else formally makes the call.

Records also outlive the requisition. California's amended FEHA regulations, effective 1 October 2025, extend the employment-records retention period from two years to four and state expressly that automated-decision system data is included, and they make evidence, or the lack of evidence, of anti-bias testing relevant to a discrimination claim and to any defence, including the response to the results 2. If you hire in California, a number deposited automatically is a number you keep for four years, and a number you keep is a number somebody can read back to you. Retention periods differ by state, so check with counsel on the schedule that applies to you.

Write things down; just choose what gets written. A finding that carries the excerpt it rests on still explains itself four years later, and the person who has to explain a bare integer will rarely be the person who configured the field.

Ask these five questions instead

Replace the integration checkbox with five questions that separate vendors on things you will still care about in a year. Each has a short answer, and a vendor who cannot give it quickly has told you where the product is thin. Send them in the same email you were going to use to ask whether they support your applicant tracking system.

1. What exactly gets written back, field by field? Ask for the list. Status and link are fine. A result field is a decision to make deliberately, not a default to inherit. 2. Who writes the report, and can I read a real one before buying? A redacted sample from an actual session tells you more than any feature list. If none can be shared, ask why. 3. What does the candidate receive? The answer separates instruments that treat the candidate as a participant from ones that treat them as an input. 4. What versions are attached to a released result? Rubric version, item bank version, reviewer. Without those, a result from March cannot be compared with one from October. 5. Can the result be read by someone who has never used the product? Export one and hand it to a manager cold. If it needs the vendor's dashboard to make sense, it will not survive a hiring debrief.

Those five are also the spine of a wider vendor review, and the rest of how to evaluate an assessment before buying one covers the evidence questions this list assumes. Whatever the answers, run it beside your existing process before it gates anybody, because a pilot that runs alongside the round is the only way to see what the instrument adds over what you already do.

Pick the instrument first, then solve the workflow

Order matters more than either decision. Choose the assessment on what it measures and what evidence it returns, then spend twenty minutes designing the handoff. Reversing that order lets a coordinator's convenience choose your evidence, and the evidence is the part that has to survive a question from a candidate, a manager or a lawyer. The handoff itself is a link, a due date, and a place to file what comes back.

Choose for difference rather than for fit. Sackett and colleagues show that combining predictors recovers most of what the old single-method numbers promised: a composite reaches about .61, and removing cognitive ability from it costs .05, dropping to .56, while structured interviews sit top of the single-method list at .42 with an 80% credibility interval running from .18 to .66 4. Those are corrected correlations that assume predictors are combined mechanically with sensible weights, not what four unscored conversations produce. The design lesson is the useful part: value comes from adding a kind of evidence you do not already have, so an instrument that duplicates your interview and plugs in neatly is the worse buy.

The workflow, once the instrument is settled, is three lines in a runbook. Who sends the link and at which stage. Where the returned report is filed. Who reads it before the debrief, and what they are asked to write down. A coordinator can run that at any volume a small team hits, and it leaves you free to change instruments later without unpicking a data connection.

One consistency check before you sign. An assessment that returns evidence a person reads sits on the safe side of the line where automation stops being logistics and starts deciding who advances; a result that auto-advances or auto-rejects has crossed it, and integration is exactly the mechanism that makes crossing it easy. If you are still weighing whether to run something of your own instead, the trade-offs in building versus buying the exercise turn on the same two things: the answer key, and who writes the reasoning.

See what gets scored

Common questions

What does a link-out workflow look like day to day?

A coordinator opens the assessment tool, mints an invite link for one candidate, and pastes it into the email or message they were already sending. When the result comes back they attach the report to the candidate record and move the stage by hand. That is two touches per candidate. Teams running this at moderate volume usually add a shared inbox rule and a saved email template, which is the entire integration most of them ever needed.

At what volume does integration genuinely pay for itself?

When the manual touches stop fitting in the gaps of a coordinator's day, which for most teams means thousands of candidates a quarter through one instrument, several people working the same queue, or a service-level commitment on stage movement. Two hundred candidates a quarter is about seven hours of copying, so below that the saving is minutes while the constraint on instrument choice is permanent. Do the multiplication with your own numbers rather than adopting somebody else's threshold.

Should the assessment result be stored in the ATS at all?

Store what a person can read and defend, and be deliberate about anything numeric. A report with the reasoning in it is a record that explains a decision later. A lone score in a field is a record that has to be explained by someone who did not create it. Decide which fields the integration writes before it is switched on, since the default configuration is a decision made by the vendor rather than by you.

The vendor offers a webhook rather than a native integration. Is that enough?

For most teams, yes. A webhook that posts a completion event and a link gives you the status sync that makes integration feel necessary, without depositing a result value in the record. It also needs someone who can receive it, so confirm your system supports inbound events and that a person owns the configuration. If nobody owns it, a manual link-out is more reliable than an integration that silently stops firing.

Does skipping integration make the process worse for candidates?

Not by itself. Candidates experience the invitation, the assessment and what they hear back, none of which depends on how the systems talk. What does affect them is delay, so the risk in a manual workflow is a report sitting unread in an inbox for a week. Name the owner and put a turnaround commitment on it, and the candidate-facing experience is identical to the integrated version.

References

  1. 1. Judgment of the Court (First Chamber) of 7 December 2023, Case C-634/21, OQ v Land Hessen (SCHUFA Holding AG intervening) Court of Justice of the European Union / Publications Office of the EU, 2023. publications.europa.eu Supports the point that producing a probability value about a person can itself be automated decision-making where a third party draws strongly on it.
  2. 2. Final Unmodified Text of Proposed Employment Regulations Regarding Automated-Decision Systems (Attachment B), 2 CCR sections 11009, 11013 California Civil Rights Department, Civil Rights Council, 2025. calcivilrights.ca.gov Supports the four-year retention period covering automated-decision system data and anti-bias testing evidence being relevant to a claim or defence.
  3. 3. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the adoption figures for work sample interviews and structured interviews across executive, middle management and individual contributor hiring.
  4. 4. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors Industrial and Organizational Psychology, 16(3), 283-300, doi 10.1017/iop.2023.24 (Cambridge University Press), 2023. cambridge.org Supports the composite validity figures and the argument for combining different kinds of evidence rather than duplicating one.

4 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.