SOC 15-2051 · Data Scientists
Data & analytics, tested on the work
Interpret a result an assistant will explain fluently, from a dataset that will not correct it.
The assignment
What the candidate is asked to do
An analysis and a written interpretation, with an AI assistant that will confidently offer the wrong causal reading if not pushed back on.
Time: 40 to 60 minutes, hard-capped 15 to 20 minutes past whatever the case states and never above 75, with a short written debrief after. Paused time does not count against either.
SOC 15-2051
Data Scientists. The benchmark is derived from live postings mapped to that code and re-derived each quarter.
What the reviewer watches for
The data & analytics failure modes
Each maps to a scored dimension, so the reviewer's finding arrives with the moment attached.
Is what would make the reading wrong written down before the analysis runs?
Whether the opening move asks what is true, what is wrong, or what the options are — and whether a constraint or criterion is named in the same breath.
Is the statistical claim traced back to the data, or to the model that said it?
Whether sources are demanded for a specific claim rather than requested in general — and whether at least one is actually opened.
Is the interpretation the candidate's, or the assistant's with the wording changed?
Whether the split between the assistant's work and the candidate's own acts is deliberate rather than convenient, read from the record of both.
Why this bank
Why we built the data & analytics bank
The failure mode here is the most expensive in the set: a fluent, wrong interpretation is indistinguishable from a right one until it reaches a decision.
This bank is live and has passed author review at launch quality.
https://olive.is/benchmarks · Edition 2026 Q3Named up front
Interpret a result an assistant will explain fluently, from a dataset that will not correct it.
The construct
Six dimensions, scored separately
- D1 · Problem framing Did the first move go after understanding, or straight for output?
- D2 · Evidence sourcing Did they demand evidence for the claim that mattered?
- D3 · Delegation boundary What did they keep, and what did they hand over?
- D4 · Working structure Did anything exist between the brief and the answer?
- D5 · Output rejection Was anything the assistant produced refused, and on what grounds?
- D6 · Verification Was anything tested against the world, and did the result change something?
SOC 15-2051
Run one data & analytics role, free
Ten attempts a month against this bank, with the full six-dimension report on every one. No card, no salesperson.