Assessment design
Should You Build Your Own AI Exercise or Buy an Assessment?
Build your own AI exercise if you hire into one or two occupations and someone here can write and defend the answer key. Buy an assessment at three or more, because every occupation repeats the authoring cost. Two overrides: where the judgment lives in your systems and material, add an in-house round beside a bought case, and if nobody is named to read what candidates submit, neither build nor buy is ready. Buying removes the authoring, not the job analysis, the reviewer hours, or your responsibility for the tool.
The takeEveryone argues this as a purchase, and the thing you are actually buying is a maintenance schedule. A case decays on the model release cycle and on the rate at which candidates pass it around; in-house, that refresh work lands on your strongest practitioner between sprints, the same person whose time you were trying to protect. So ask a vendor how many sessions their reviewers have marked, how often the case has been replaced, and who reads a session on a Friday. Those answers describe an operating habit, and the operating habit is the product. A vendor who reaches for the brochure instead is selling you a PDF.
Where Olive fits
Open a role and see what the work shows
Buying moves the authoring and the reading hours off your team; it does not move the job analysis or the decision. Olive sits on that side of the line: twelve authored cases per occupation, six findings a human reviewer writes against timestamped moments in the session, no hiring recommendation attached, and the identical report granted to the candidate.
Rank your shortlistWhich is cheaper, building or buying?
Building is cheaper for one occupation and more expensive by the third. The cost tracks the number of kinds of work you hire into rather than the number of people you hire, because a task only samples the content of the job it was drawn from 1. One case, one answer key and one calibrated reviewer is about two weeks of somebody's attention. Six of each is a standing program with an owner.
Those are two different slopes, and the usual mistake is comparing them at a single point. A vendor charges per candidate or per seat, so the bought cost rises with hiring volume and falls to nothing when a role closes. Your build cost is close to flat in volume and steep in variety: authoring the fourth case costs roughly what the first did, and no role's answer key is reusable on the next one.
One line item sits in both columns and gets left out of both. Work sample tests are administered by people who observe and, sometimes, rate what the applicant actually did 3. If the assessment you buy does not come with a reviewer, buying removed the authoring and left the recurring cost exactly where it was, which is why the quote and your spreadsheet so often disagree by a factor of two.
The arithmetic changes again if the plan was one instrument for the whole company. That plan does not survive contact with the job-relatedness standard: the rubric can travel and the task cannot. So count occupations before pricing anything, and count the ones you will hire into this year rather than the ones on the org chart.
What does an in-house AI exercise actually cost?
Five line items, and only three of them are paid once. The task, the seeded errors and the answer key cost you once per occupation. The rubric anchors cost you once, then again every time two reviewers disagree. Reading a session and writing the findings costs you on every candidate, permanently. Federal guidance on work samples names the same shape: costly to develop in both time and money, needing periodic updating, and expensive to administer 3.
| Line item | Paid | Who can actually do it |
|---|---|---|
| The case, drawn from that occupation's own material | Once per occupation | Someone who does the job |
| Seeded errors: the confident wrong answer the assistant will produce | Once per occupation, revisited as models change | Someone who does the job |
| The answer key: what a good response settles, and what it cannot | Once per occupation | The same person, plus a second reader |
| Rubric anchors: what each row looks like satisfied in this job | Once, then whenever two reviewers diverge | The reviewers, together |
| Reading a session and writing the findings | Every candidate | A trained reviewer |
| Rewriting a case that leaked | Whenever it leaks | Back to row one |
The seeded errors are the expensive part and the part nobody budgets. A case earns its keep only if the assistant will produce something confidently wrong inside it, and finding that takes several runs through the assistant with the packet in front of you. It also has a shelf life measured in model releases: the plausible-but-wrong answer that carried your case in March may be handled correctly by the summer, at which point the case tests nothing.
Anchors decay quietly. Two reviewers who agreed in March are half a rubric apart by September unless somebody double-marks a sample and reconciles the differences, and the drift surfaces as the same performance drawing different outcomes in different months. Reviewer agreement is a maintenance cost rather than a property the rubric has once and keeps.
Price all of this in hours of the person who would have to write it, not in dollars. That person is usually the strongest practitioner on the team, their hours are the scarcest thing you own, and a build plan that quietly spends four of their weeks is a real cost even though it never appears on an invoice.
What does buying remove, and what does it leave with you?
Buying removes the authoring, the seeded errors, the answer key, the rewrite after a leak, and, where the vendor reviews sessions, the reading hours. It leaves you the job analysis, the decision, and the record. The Uniform Guidelines are blunt about the last one: an employer supporting a procedure with a publisher's validity evidence is cautioned that users "are responsible for compliance with these guidelines" 2.
A vendor's evidence is also not a transfer. Validity established on one job carries to another only where incumbents perform substantially the same major work behaviors, shown by job analyses on both jobs 2. So the question to put to a vendor is not whether the assessment is validated, but validated on which occupation, against what work behaviors, and how those compare to the ones in your job description.
The Guidelines rule out the things a sales process supplies most readily. General reputation, promotional literature, testimonial statements and the credentials of the seller are specifically named as unacceptable substitutes for evidence of validity 4. That makes the useful questions documentary ones. What to ask a vendor before paying is mostly a list of documents, and a vendor who cannot produce them has sold you the authoring and none of the defense.
What buying genuinely buys is somebody else's maintenance schedule. The case gets rewritten when it leaks and refreshed when the models move, by a person whose job that is rather than by your best engineer between sprints. Ask how many sessions their reviewers have marked, how disagreements between two reviewers get reconciled, and how often the case has been replaced. Those three answers tell you whether you are buying a maintained instrument or a PDF.
When is building the right call?
Build when three things are true together: you hire into one or two occupations, the work is genuinely local (your systems, your material, a constraint that lives in your own files), and someone here can write an answer key and defend it under questioning. The third condition fails most often, because the person who could write it is usually the person whose time you were trying to protect.
There is work no bought case reaches. A constraint that only exists in your data, an internal-mobility population being assessed against your own systems, a regulated process with house-specific steps. If the judgment you need to see is judgment against your material, nobody can author that for you. A bought case is grounded in an occupation, which is what makes it defensible and also what stops it from knowing anything about your company.
The honest risk in building is concentration. A case written by the hiring manager who also reads the sessions and also makes the hire is one person's opinion applied three times and called a process. Split the roles even in a small team: one author, one second reader on the answer key, and a bar written down before the first candidate sees the task rather than inferred from the first three who took it.
And build small. One case, one occupation, a fixed time cap and a written protocol beats a six-role plan that ships next year. If the first case teaches you that nobody has time to read the sessions, that is the finding. You have learned the price of the recurring column for the cost of one occupation instead of six.
Decide with a four-question test
Four questions, in order. How many occupations will you hire into in the next twelve months? Can someone here write an answer key and defend it under questioning? Who reads the sessions, and whose calendar do those hours come out of? What happens the third time a candidate arrives already knowing the case? Two answers pointing at nobody means buy.
| Answer | What it implies |
|---|---|
| One or two occupations, an author available | Build, and keep it to one case |
| Three or more occupations | Buy: the authoring cost repeats in full for each |
| No named author with occupational depth | Buy; a case written by someone outside the job measures the wrong thing |
| No named reader for the sessions | Neither yet: an unread assessment is a delay you charge candidates for |
| Work is specific to your systems and material | Build, or add a short in-house round beside a bought case |
Then run whichever you chose beside your existing process without letting it decide anything, for a full round. Shadow mode before it gates anyone is how you find out that the case is guessable, that the anchors read differently to two people, or that nobody has the hours. All of them are cheap discoveries before the assessment starts turning people away and expensive after.
Write down three things whichever way this goes: which decision the exercise informs, what you do when it disagrees with the interview, and who owns it next quarter. Buying does not answer any of the three. It answers who writes the case and who reads the session, which is most of the cost and none of the judgment.
Common questions
How much does it cost to build one AI exercise in-house?
Price it in hours rather than dollars, because every hour is somebody senior's. One occupation runs to a case, a set of seeded errors, an answer key, anchors for each rubric row, and a dry run with two people who do the job. Realistically that is a couple of weeks of part-time attention from your strongest practitioner, plus a rewrite after the first few candidates. Then add the permanent line: someone reads every session and writes the findings. Past the first dozen candidates, that reading is the dominant cost and it never ends.
Does buying an assessment move the legal risk to the vendor?
No. The federal Uniform Guidelines on Employee Selection Procedures place responsibility for compliance on the employer using the selection procedure, not on the publisher that sold it. Validity evidence from another job or another employer carries over only where job analyses show substantially the same major work behaviors. Buying gets you a maintained case and, sometimes, a reviewer. The job analysis, the records, and the defense of the decision stay on your side of the contract.
How many candidates does it take before buying is cheaper?
It is the wrong axis. Per-candidate pricing means the bought cost rises with volume, so high headcount favors building if you already have the case and the reviewer. The build cost repeats per occupation, so hiring into four job families beats you regardless of how few people you hire into each. Count occupations first; only then check whether the per-attempt price at your volume is worse than the reading hours you would spend anyway.
Can we buy now and build later?
That is usually the right order. Two cycles of a bought assessment teach you what your rubric anchors should say, what reviewer disagreement looks like on real sessions, and how many hours reading actually takes. Those are the three numbers that make an in-house build estimable instead of hopeful. Build afterward for the one occupation where your material genuinely matters, and keep buying for the rest. What you cannot do is copy a vendor's case; it is theirs, and a stolen case has no answer key behind it.
What if our work is too specialized for an off-the-shelf case?
Check whether the specialization is in the occupation or in your company. A bought case is authored against an occupation's real work, which covers more than buyers expect: the analysis, the source packet, the deliverable. What it cannot carry is your data, your internal systems, or a constraint that exists only in your files. Where that is the thing you need to see, put a short in-house round beside the bought assessment rather than replacing it, and write the answer key for that round before you use it.
References
- 1. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 (Technical standards for validity studies) ✓ ecfr.gov Section 14C: a selection procedure is supported by content validity to the extent that it is a representative sample of the content of the job, on the basis of a job analysis of that job's critical work behaviors; the procedure may be developed from the job in question or previously developed by another user or a test publisher.
- 2. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.7 (Use of other validity studies) ✓ ecfr.gov Section 7A: users relying on validity studies conducted by other users or described in test publishers' manuals are cautioned that they are responsible for compliance with the Guidelines. Section 7B(2): borrowed validity evidence applies only where incumbents in both jobs perform substantially the same major work behaviors, shown by appropriate job analyses on each.
- 3. Assessment and Selection: Work Samples and Simulations ✓ opm.gov Development costs: may be costly to develop in both time and money, and may require periodic updating when the work changes. Administration costs: may be time consuming and expensive to administer, and requires individuals to observe and sometimes rate applicant performance. Undated guidance, page text verified 2026-08-24.
- 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.9 (No assumption of validity) ✓ ecfr.gov Section 9A specifically rules out, in lieu of evidence of validity, the general reputation of a procedure or its publisher, all forms of promotional literature, and testimonial statements and credentials of sellers, users or consultants.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.