The field

We are not the only ones trying

Anyone who tells you this category is empty has not looked at it recently. Here is who is doing what, and where we are actually different.

Honest comparison

Where each approach stops

ApproachWhat it does wellWhere it stops
Multiple-choice AI-literacy testsCheap, fast, scales to any volume, fine as a top-of-funnel filterTests whether a candidate can define a hallucination, not whether they catch one. Vocabulary is not behavior.
Scored hallucination-spotting simsGenuinely tests the catch — this is real and it worksOne dimension inside a short simulation, no candidate's-own-AI workflow, and no timestamped evidence trail across a real deliverable.
Code-based AI collaboration gradersSession replay against real engineering work; strong for the roles they coverEngineering only, inside a sandboxed editor with a provided model. Some score "AI reliance", which measures volume rather than judgment.
Unwatched take-homesEcologically valid, cheap, familiar to candidatesTrivially outsourceable, and produces an artifact with no record of how it was made — which is precisely the part worth scoring.
Live AI interviewsHuman present, conversational, quick to scheduleCan only capture a candidate describing verification. 38% of US candidates have abandoned a process over one.
In-house roundsThe best of the lot, when you can staff itRequires an interview-engineering team, and sees only its own funnel — there is no view of what the occupation is asking for outside it.

Every row here is a real approach with real merit. The claim is not that they do not work — it is that three things together are the gap: six dimensions rather than one, timestamped evidence rather than a score, and the candidate's own AI rather than a sandbox. A fourth used to be claimed here — a live corpus benchmark rather than a panel's opinion — and it is removed: the corpus grounds what the banks test, but nothing in the product compares a candidate against it, which is what /benchmarks has said in print all along.

The four-part difference

What is actually ours

Depth

Six, not one

Six evidence-anchored dimensions, each probed and scored separately, rather than one composite judgment about "AI collaboration".

Evidence

The moment attached

Every finding carries a capture timestamp, a transcript turn, an artifact diff or a debrief answer. You can open the thing you are being told about.

Workflow

Their assistant, not ours

Candidates work with whatever they already use rather than a sandboxed model, so the session is not partly a test of adapting to an unfamiliar tool.

Reach

Beyond code

Analyst, consultant, product and marketing banks face the same six failure modes, and no code-based instrument reaches them.

The strongest validation of this construct is that the most capable employers built it themselves.

Named engineering and design organizations independently converged on scoring validation and error-identification in their own AI-assisted rounds.

Which is also the answer to "who is this for"

Everyone who cannot staff an interview-engineering team — which is almost everyone.

The teams that can build it have already told you the construct is right.

Side by side

Compare it against your current round

Run one live bank against a role you are hiring on now and put the report next to what your current round produces.

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.