SOC 15-1252 · Software Developers
Software engineering, tested on the work
Extend a small unfamiliar repo with an AI assistant that will write the whole change if nobody stops it.
The assignment
What the candidate is asked to do
The candidate is given a working repository, a feature request, and a short context packet. The brief is open, the assistant is capable, and nothing inside the conversation can settle whether a library behaves the way it is described.
Time: 40 to 60 minutes, hard-capped 15 to 20 minutes past whatever the case states and never above 75, with a short written debrief after. Paused time does not count against either.
SOC 15-1252
Software Developers. The benchmark is derived from live postings mapped to that code and re-derived each quarter.
What the reviewer watches for
The software engineering failure modes
Each maps to a scored dimension, so the reviewer's finding arrives with the moment attached.
Does the first move go after what the repo already does, or straight at the diff?
Whether the opening move asks what is true, what is wrong, or what the options are — and whether a constraint or criterion is named in the same breath.
Is the library's behavior demanded from a source and read, or taken as described?
Whether sources are demanded for a specific claim rather than requested in general — and whether at least one is actually opened.
Is anything run or decided by hand, or is the whole change the assistant's?
Whether the split between the assistant's work and the candidate's own acts is deliberate rather than convenient, read from the record of both.
Why this bank
Why we built the software engineering bank
This is the largest measured concentration in the corpus — 79.7% of the cohort carrying an AI-usage requirement sits in Technology — and it is the only role where code-focused incumbents already compete, so the bank is built to their depth and past it.
This bank is live and has passed author review at launch quality.
https://olive.is/benchmarks · Edition 2026 Q3Named up front
Extend a small unfamiliar repo with an AI assistant that will write the whole change if nobody stops it.
The construct
Six dimensions, scored separately
- D1 · Problem framing Did the first move go after understanding, or straight for output?
- D2 · Evidence sourcing Did they demand evidence for the claim that mattered?
- D3 · Delegation boundary What did they keep, and what did they hand over?
- D4 · Working structure Did anything exist between the brief and the answer?
- D5 · Output rejection Was anything the assistant produced refused, and on what grounds?
- D6 · Verification Was anything tested against the world, and did the result change something?
SOC 15-1252
Run one software engineering role, free
Ten attempts a month against this bank, with the full six-dimension report on every one. No card, no salesperson.