Item banks
Built for the actual job
A generic reasoning test measures a generic thing. Every bank here is grounded in one occupation and in what its postings asked for this quarter.
Where we ship
Twelve live banks, one bar.
Twelve shallow banks would have been easier to announce and worse to buy. Every live bank cleared the same author review before it shipped; a bank still in build is listed, marked, and disabled until it does.
Software engineering
Extend a small unfamiliar repo with an AI assistant that will write the whole change if nobody stops it.
Financial analysis
Write a diligence memo over a source packet whose claims are only settled by opening them.
Management consulting
Size a market and recommend an entry approach from evidence that settles nothing until someone opens it.
Data & analytics
Interpret a result an assistant will explain fluently, from a dataset that will not correct it.
Product management
Turn a vague request into a spec, where the vagueness is the test and the AI will happily resolve it for you.
Marketing
Build a positioning brief from research where the most on-message statistic is the one least likely to be opened.
Journalism
File a story with an AI assistant, from sources whose claims only settle when someone checks them.
Healthcare revenue cycle
Work a queue of denials with an assistant that will draft a persuasive appeal for the one claim you should be conceding.
Legal operations
Take a position on a customer's contract from a packet where the playbook, the system record and the signed precedent do not agree.
Accounting & audit
Decide what an audit will require a client to correct, from a file that reads settled until someone re-adds it.
Underwriting & claims
Bind, decline or counter on a submission whose application and its survey do not describe the same building.
Supply chain
Award a supply contract on a landed cost that three quotes, written on three different terms, do not agree on.
Beyond code
The roles a coding test never reaches.
The AI-usage requirement is not concentrated in engineering. It is in the memo, the brief, the spec and the model — work whose failure mode is a confident, fluent, wrong paragraph that no test suite will ever catch.
Those roles have had no work-sample instrument at all. They got a conversation and a reference check. This is the first bank built for them.
https://olive.is/benchmarks · Edition 2026 Q3Where the requirement landed
79.7% of the 34,743 postings carrying an AI-usage requirement, measured on 2026-08-11, sit in Technology. The fastest growth in the corpus is in finance, consulting and marketing postings.
Not listed?
Don't see your role?
Twelve banks are live. Tell us which occupation you hire for and we will tell you honestly where it sits in the queue.