The construct
Six things worth watching
Judgment is not one thing, so it is not one score. Each dimension is probed, scored and evidenced on its own.
The six dimensions
What gets probed, and what counts
Every finding arrives anchored to a moment: a capture timestamp, a transcript turn, an artifact diff or a debrief answer.
Problem framing
Did the first move go after understanding, or straight for output?
An open brief, an assistant, and an unscripted first move.
Evidence sourcing
Did they demand evidence for the claim that mattered?
The assistant asserts claims that matter to the answer, with nothing behind them until someone asks.
Delegation boundary
What did they keep, and what did they hand over?
Real work and a capable assistant, so every act is a choice about which of them does it.
Working structure
Did anything exist between the brief and the answer?
A deliverable that could be asked for in one prompt and shipped as it came back.
Output rejection
Was anything the assistant produced refused, and on what grounds?
Fluent, plausible answers arrive, and every answer the candidate goes on to use is a chance to push back.
Verification
Was anything tested against the world, and did the result change something?
The assistant asserts things confidently, and nothing inside the conversation can witness whether they are true.
The wall
What is never scored.
Prose quality. Prompt syntax. Tool trivia. Speed. Tone of voice. And how much AI the candidate used.
That last one matters. "AI reliance" as a metric scores the wrong thing — it rewards volume and punishes the candidate who correctly decided the model was not the right instrument for a step. Judgment about the output is the test; usage is not a virtue.
Prosody is captured and used only as an integrity signal. It never reaches a score, and it never reaches the employer as one.
Six things we throw away
Writing polish · prompt syntax · tool trivia · time to complete · vocal tone · number of AI turns.
Each one correlates with something other than judgment, and every one of them is easy to game.
No composite
A number you can't interrogate is a number you can't defend.
There is no single blended score by default. Six findings arrive separately because they mean different things: a candidate who demands a source for every claim and never refuses one of the answers is a specific, useful profile, and averaging it into a 74 destroys the only information in it.
It is also the defensible posture. Under NYC LL144 and the Illinois rules, "the model gave them a 74" is not an explanation. "At 07:40 the rate went into the memo unchecked, here is the capture" is.
Six findings, six excerpts
Demonstrated · partly demonstrated · not demonstrated — each with the moment attached and a human's confirmation on it.
See it scored
All six, on a real session
The sample report is a complete financial-analysis session: what was caught, what was missed, and the excerpt behind each call.