SOC 13-2053 · Insurance Underwriters
Underwriting & claims, tested on the work
Bind, decline or counter on a submission whose application and its survey do not describe the same building.
The assignment
What the candidate is asked to do
The candidate produces an underwriting verdict with terms from a submission packet — application, survey, loss runs, guideline, forms and a rating worksheet — with an AI assistant available throughout. The appetite and the binding authority are in the packet; nothing inside the conversation settles which of two documents describes the risk.
Time: 40 to 60 minutes, hard-capped 15 to 20 minutes past whatever the case states and never above 75, with a short written debrief after. Paused time does not count against either.
SOC 13-2053
Insurance Underwriters. The benchmark is derived from live postings mapped to that code and re-derived each quarter.
What the reviewer watches for
The underwriting & claims failure modes
Each maps to a scored dimension, so the reviewer's finding arrives with the moment attached.
Is the decision the verdict has to settle named before any terms are drafted?
Whether the opening move asks what is true, what is wrong, or what the options are — and whether a constraint or criterion is named in the same breath.
Is the survey opened, or is the application taken as the account?
Whether sources are demanded for a specific claim rather than requested in general — and whether at least one is actually opened.
Which figures are rebuilt from the loss runs, and which arrive already totalled?
Whether the split between the assistant's work and the candidate's own acts is deliberate rather than convenient, read from the record of both.
Why this bank
Why we built the underwriting & claims bank
Insurance is the one occupation in the set whose use of AI is already regulated by name — New York and Colorado both require an insurer to show why a model-produced underwriting decision is not unfairly discriminatory — so what an employer buys here is not a nice-to-have signal but the record an examiner will eventually ask them to produce.
This bank is live and has passed author review at launch quality.
https://olive.is/benchmarks · Edition 2026 Q3Named up front
Bind, decline or counter on a submission whose application and its survey do not describe the same building.
The construct
Six dimensions, scored separately
- D1 · Problem framing Did the first move go after understanding, or straight for output?
- D2 · Evidence sourcing Did they demand evidence for the claim that mattered?
- D3 · Delegation boundary What did they keep, and what did they hand over?
- D4 · Working structure Did anything exist between the brief and the answer?
- D5 · Output rejection Was anything the assistant produced refused, and on what grounds?
- D6 · Verification Was anything tested against the world, and did the result change something?
SOC 13-2053
Run one underwriting & claims role, free
Ten attempts a month against this bank, with the full six-dimension report on every one. No card, no salesperson.