Validation & compliance
Show your working
An assessment that cannot show why it decided something is a liability, not a product. Here is what we measure, what binds us, and where every number came from.
Scoring
Today a person writes every word
There is no automated scorer in the product. Not a weak one — none. Every finding on every report is typed by a human reviewer reading the session against the bank's answer key. The gates below are what an automated scorer would have to pass before it is allowed to propose anything, and none of it is running yet.
Human review is the product today, not a fallback
Release is structural rather than promised: the database refuses every client write to a result and every read of a result that is not released, and the one code path that can mark a result released is gated on a staff allowlist checked against a verified sign-in. There is no scheduler, no auto-promote, and no release-on-submit.
A golden set, double-marked by humans
At least 150 double-human-marked sessions per occupation family, built before any scorer is trusted with anything.
Agreement measured per dimension
Any scorer must reach a quadratic-weighted kappa of 0.6 or better against human consensus — measured per dimension, not averaged across them. A construct that works for five dimensions and fails on one is not shipped as working.
Distributions monitored by slice
Per-dimension outcome distributions monitored across slices where it is lawful to do so, with four-fifths ratios published in an annual bias audit.
Versioned, immutably — and exportable
Every released result carries its rubricVersion, scorerVersion and bankVersion and cannot be re-saved after release. All of them can be exported in one archival file that stays readable without this application, which is what a four-year recordkeeping obligation actually needs.
Human review does not step down on a schedule
If an automated scorer is ever introduced, the share of results it touches moves only against measured agreement, never against a date or a volume target. Today the share is zero.
Regulatory
What binds us, and the answer
Reviewed quarterly. Where a regime is repealed or not yet in force, the row says so rather than claiming compliance with something that does not bind.
| Regime | Status | What binds | Our answer |
|---|---|---|---|
| NYC LL144 | In force | Bias audit plus notice if the tool substantially assists or replaces the decision | Human review before release and per-dimension evidence with no composite, both structural today. The bias audit has not been performed — there is not yet enough volume for a four-fifths ratio to mean anything. The 10-day candidate notice ships in the employer kit. |
| Illinois HB 3773 + AIVIA | In force | Notice when AI is used in hiring decisions; consent for AI analysis of interview video | Disclosure-first flow doubles as the notice artifact; debrief consent is explicit and separate |
| California FEHA ADS | In force | Anti-bias testing evidence; four-year recordkeeping | Immutable versioned results, and a full archival export of every released report carrying its rubric, scorer and bank versions — a record whose instrument cannot be identified is not a record. Deletion is a scheduled request rather than an instant wipe, so a retention obligation can be honored before data goes. The anti-bias testing evidence is still outstanding. |
| Texas TRAIGA | In force | Deployer impact assessments and notice | Impact-assessment template available on request. Not yet packaged as a tier feature. |
| Colorado SB 24-205 | Repealed before it applied | Nothing. SB 26-189 repealed and reenacted it, signed 2026-05-14, and the replacement takes effect 2027-01-01 | No compliance is claimed against a regime that never bound anyone. The 2027-01-01 replacement sits on the quarterly review list, and this row stays published so the wrong 2026 date is corrected rather than dropped |
| EU AI Act, Annex III | Obligations from Dec 2027 | Risk management, technical documentation, conformity, post-market monitoring | Technical file begins in 2027 planning; not treated as a 2026 build gate |
| ADA / Title VII | Statutes bind | Accessible, equally-valid paths; vendor-level exposure is real | Typed think-aloud and text debrief built in the same sprint as capture — not afterwards as an accommodation |
Vendor liability is not hypothetical in this category: litigation has established that an assessment tool's provider can be on the hook alongside the employer. We build as though that is settled, because for our purposes it is — which is also why this table marks what is not built in bold rather than leaving a reader to assume the row is satisfied.
Accessibility
The equal path is not a lesser one.
Speech recognition error rates for deaf and disordered speech have been measured around 78% against roughly 18% for typical speech. Any product that makes voice the only road to a score has built a discriminatory instrument, whatever it intended.
So the typed think-aloud and the text debrief are built in the same sprint as the voice ones. There is scale evidence that chat-based assessment actually completes better than video — the accessible path is not a degraded product, and we do not treat it as one.
78% vs 18%
Word error rate for disordered speech against typical speech. A voice-only instrument does not measure judgment; it measures diction.
Where the numbers come from
Every market figure quoted on this site traces to a published source with a date and a verification status, or to the olive.jobs posting corpus, whose figures are re-derivable from published queries. We keep the distinction visible because it matters:
- Corpus figures — posting prevalence, employer counts, occupation concentration — come from our own measurement and can be re-derived. They carry the caveats of any posting corpus: they measure what employers write, not what they do.
- Published research — the bad-AI-hire rate, candidate-experience figures, take-home completion rates, salary premia — is third-party, cited with publisher and date, and flagged where the publisher is a vendor with an interest in the result.
- Vendor claims about competitors are flagged as vendor claims. Marketing copy in this category ages in weeks; we re-verify before quoting it anywhere it matters.
What we refuse to cite
We do not publish a total-addressable-market figure. Third-party sizings of "talent assessment" span roughly $1.3B to $27B on inconsistent scope definitions — a range wide enough that quoting any point in it is a decision dressed as a measurement. The corpus numbers are load-bearing; the market-size numbers are not, so they are not here.
The honest limit of our own trend
A posting-prevalence trend can move because employers changed their minds, or because which employers are hiring changed. We run the trend within-industry and against a fixed panel of employers present in both periods, and publish both. If they disagree, the headline number is the one that is wrong.
The work itself
Read the source, not the summary
The measurement, the construct and the product are all inspectable. We would rather you checked than took this page's word for it. (The repositories themselves are private; ask if you need to see one.)
The measurement essay
The full scroll-driven analysis of the posting corpus, with every figure carrying the query that produced it and every honesty beat left in.
The six dimensions in full
What each dimension probes, what counts as evidence for it, and what a pass and a fail actually look like on a report.
The assessment itself
The employer dashboard, the candidate workspace and the review console. Open an account and run one against a live bank.
Still open
Questions we have not answered here?
The FAQ covers integrity, proctoring, comparison and procurement. Beyond that, ask us directly.