Press

Every claim on this page is checkable before you print.

Each figure names its source and its method, and the gaps are listed here as plainly as the facts.

Research

Speech is not a flat string of words

Olive builds voice models to improve understanding, not to imitate a human speaker. A speech system normally collapses what you said into flat orthographic text at the earliest stage. Timing, pauses, emphasis, lengthening and overlap are discarded before any model reads the transcript.

Olive does not treat speech as a transient input that collapses into a string. It outputs a linguistically structured transcript that preserves that interactional information. Keeping it matters for usability, since turn-taking and intent resolution depend on it. It is also a fairness safeguard: socially meaningful cues in speech, including dialect, can become a covert risk surface.

The fairness motive here is measurement rather than screening. A transcript that retains timing and emphasis allows systematic tests of downstream model behavior. Hold the meaning constant, change the dialect or the prosody, and see what moves.

Measured

The number that cost a feature

Olive measured its own notation instrument against human ground truth and got a macro F1 of 0.203 at best sample, 0.17 to 0.20 across the range.

That instrument was demoted from a scored dimension to a proctoring signal, and no prosody-derived field may reach a candidate's result.

Provenance

Who measured what

Two of these are Olive's own measurements. The rest belong to their authors, are cited as theirs, and were not discovered here.

Cited

The transcription gap is measured

Koenecke et al. found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers, across five commercial speech recognition systems. PNAS, 2020.

Derived

The 84% figure is ours

That relative gap — (0.35 − 0.19) / 0.19 — is arithmetic Olive did on two numbers the paper published. Printing it beneath someone else's citation is the error this whole discipline was built to stop.

Cited

Safety training misses dialect

Hofmann et al. report that dialect prejudice survives human-feedback training, which "obscures the racism on the surface, but the racial stereotypes remain unaffected on a deeper level." Nature, 2024. Direction and mechanism only; no magnitude is claimed here.

Cited

The stakes are not theoretical

Omiye et al. found language models propagating debunked race-based medicine. npj Digital Medicine, 2023. Referenced for the stakes in an adjacent high-stakes domain, with no figure quoted from it.

Measured

Prosody added 12.2 points

Stance recovery reached .434 with prosody against .312 from words alone. The corpus was synthetic and the judge shared a model family with the generators, so the absolute rates are uncalibrated.

Measured

Distillation scored NO-GO

Production distillation returned +1.1 points against a +4-point gate, so it did not ship. A result that misses its gate is published here at the same size as one that clears it.

Press Desk

No agency, no embargo process

Write to press@olive.is and the reply comes from someone who built the thing. Nobody sits between a reporter and the answer, and there is no embargo calendar to negotiate.

Three of the four pieces above are about the personal AI agent, and one is about the voice research. None of them covers this assessment product, its six dimensions or its human-written findings. Please do not read them as validation of it.

Any figure published here can be re-derived with you on the call. Ask for the query, the corpus or the arithmetic behind a number and you get all three. A figure that cannot survive that is one Olive should not have printed.

Standing offer

Re-derive any number

Bring a figure from this site and Olive will show its source, its method, and whether it was measured, computed or cited.

Boilerplate

One paragraph, usable as-is

Olive builds an assessment of how candidates actually work with AI. A candidate does a real piece of the job — a diligence memo, a repo to extend — with an AI assistant, inside a workspace scoped to one browser tab and with no camera anywhere in the product. Six things are watched for in the record of how the work was actually done: how the problem was framed, what evidence was demanded, what was kept rather than handed over, what existed between the brief and the answer, what was refused, and what was tested against something outside the conversation.

A human reviewer then writes a finding on each of the six with the evidence excerpt attached. There is no automated scorer and no composite number. Nothing reaches the employer until a person has written all six, and the candidate is granted the identical report.

Short version

25 words

Olive tests whether candidates catch what AI gets wrong. Six judgment dimensions, evidence excerpts, every finding written by a human reviewer, and no composite score.

The facts

What is true today

Each of these is structural — enforced in code or in database rules — rather than a policy we are asking you to take on trust.

Review

A person writes every word

There is no automated scorer in the product. Database rules refuse every client write to a result and every read of one that is not released; a single staff-gated path releases it, and only after all six findings are written.

Capture

No camera exists

Not disabled — absent. video: true appears nowhere in the codebase. Screen capture is requested for the assessment tab only, and declining it is a supported outcome rather than a failure.

Fairness

The candidate gets the same report

Not a summary and not a softened version — the identical document, free, granted by the same rule that grants it to the employer.

Scoring

No composite, anywhere

Six findings and the excerpts they rest on. No overall number is stored or rendered, because a number you cannot interrogate is a number you cannot defend.

Coverage

12 item banks are live

Every occupation listed on the site is selectable. Each bank cleared the same author review before it shipped, and none was announced ahead of it.

Research

The requirement is real; its size is not settled

Two instruments bracket the AI-usage requirement between 1.4% and 8.1% of the 998,166 postings dated 2026-07, measured on 2026-08-11. Both ends publish, not the flattering one.

The mark

Logo and usage

The mark is a two-color SVG at /assets/logo.svg. The disc inherits the ink of whatever it sits on and the face is a separate fill, so on a dark background it inverts — white disc, dark face — rather than disappearing.

Please do not recolor it into a brand palette that is not ours, add effects to it, or set it below 24px, where the eyes and nose close up and it becomes a dark blob. It needs a light or dark ground; on a mid-tone it loses the two values that make it readable.

Name

It is "Olive"

Sentence case, never all-caps, never "OLIVE". The assessment product is the Judgment Gate; the company is Olive.

Contact

One address, a person answers

There is no press agency and no embargo process. Write and you will get a reply from someone who built the thing.

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.