The field
We are not the only ones trying
Anyone who tells you this category is empty has not looked at it recently. Here is who is doing what, and where we are actually different.
Honest comparison
Where each approach stops
| Approach | What it does well | Where it stops |
|---|---|---|
| Multiple-choice AI-literacy tests | Cheap, fast, scales to any volume, fine as a top-of-funnel filter | Tests whether a candidate can define a hallucination, not whether they catch one. Vocabulary is not behavior. |
| Scored hallucination-spotting sims | Genuinely tests the catch — this is real and it works | One dimension inside a short simulation, no candidate's-own-AI workflow, and no timestamped evidence trail across a real deliverable. |
| Code-based AI collaboration graders | Session replay against real engineering work; strong for the roles they cover | Engineering only, inside a sandboxed editor with a provided model. Some score "AI reliance", which measures volume rather than judgment. |
| Unwatched take-homes | Ecologically valid, cheap, familiar to candidates | Trivially outsourceable, and produces an artifact with no record of how it was made — which is precisely the part worth scoring. |
| Live AI interviews | Human present, conversational, quick to schedule | Can only capture a candidate describing verification. 38% of US candidates have abandoned a process over one. |
| In-house rounds | The best of the lot, when you can staff it | Requires an interview-engineering team, and sees only its own funnel — there is no view of what the occupation is asking for outside it. |
Every row here is a real approach with real merit. The claim is not that they do not work — it is that three things together are the gap: six dimensions rather than one, timestamped evidence rather than a score, and the candidate's own AI rather than a sandbox. A fourth used to be claimed here — a live corpus benchmark rather than a panel's opinion — and it is removed: the corpus grounds what the banks test, but nothing in the product compares a candidate against it, which is what /benchmarks has said in print all along.
The four-part difference
What is actually ours
Six, not one
Six evidence-anchored dimensions, each probed and scored separately, rather than one composite judgment about "AI collaboration".
The moment attached
Every finding carries a capture timestamp, a transcript turn, an artifact diff or a debrief answer. You can open the thing you are being told about.
Their assistant, not ours
Candidates work with whatever they already use rather than a sandboxed model, so the session is not partly a test of adapting to an unfamiliar tool.
Beyond code
Analyst, consultant, product and marketing banks face the same six failure modes, and no code-based instrument reaches them.
The strongest validation of this construct is that the most capable employers built it themselves.
Named engineering and design organizations independently converged on scoring validation and error-identification in their own AI-assisted rounds.
Which is also the answer to "who is this for"
Everyone who cannot staff an interview-engineering team — which is almost everyone.
The teams that can build it have already told you the construct is right.
Side by side
Compare it against your current round
Run one live bank against a role you are hiring on now and put the report next to what your current round produces.