Dimension D1 of 6
Problem framing
Did the first move go after understanding, or straight for output?
The probe
What the session puts in front of them
An open brief, an assistant, and an unscripted first move.
Each live bank ships twelve authored cases and a role opens on one of them, so a leak costs that case rather than the bank. Rotation within a role is the next integrity work; everyone invited to one role still meets the same case, and we would rather name that than imply otherwise.
Enacted, not narrated
Whether the opening move asks what is true, what is wrong, or what the options are — and whether a constraint or criterion is named in the same breath.
The finding is anchored to a moment you can open: a capture timestamp, an AI transcript turn, an artifact diff, or a debrief answer.
The two outcomes
What a pass and a fail look like
No composite. Each dimension resolves on its own evidence, and you can read the moment it turned.
When it goes wrong
The first thing asked of the assistant is the deliverable, and the framing goes with it.
When it goes right
The problem is stated in the candidate's own terms — what is being decided, and what would make an answer wrong — before anything is generated.
What we ignore either way
Prose quality, prompt syntax, tool trivia, speed, tone of voice, and how much AI the candidate used. Volume of usage is not a virtue; judgment about the output is the test.
How it is scored
From session to D1 finding
The same four steps for every dimension, so a result means the same thing across roles.
The situation is set up
An open brief, an assistant, and an unscripted first move. Each live bank ships twelve authored cases and a role opens on one of them; rotation within a role is roadmap, so until it lands everyone invited to one role meets the same case.
The session is captured
Workspace events, the AI exchange, the deliverable's history and — where the candidate allows it — tab-scoped screen capture and microphone, recorded together so a reviewer can read them against each other.
A reviewer writes the finding
A person reads the session against the bank's answer key and writes the finding themselves, with the evidence excerpt attached. There is no automated scorer in the product — not a weak one, none — so this step is the scoring rather than a check on it.
And releases it deliberately
A release is refused until all six dimensions are written and the reviewer has confirmed the report says nothing the candidate should not read about themselves. It is enforced in the database rules rather than by policy: no client can write a result, and no client can read one that is not released.
The rest of the construct
The other five dimensions
- D2 · Evidence sourcing Did they demand evidence for the claim that mattered?
- D3 · Delegation boundary What did they keep, and what did they hand over?
- D4 · Working structure Did anything exist between the brief and the answer?
- D5 · Output rejection Was anything the assistant produced refused, and on what grounds?
- D6 · Verification Was anything tested against the world, and did the result change something?
Previous: Verification · Next: Evidence sourcing
See D1 scored on a real session
The sample report shows all six dimensions with the evidence excerpts attached, exactly as an employer receives them.