The construct

Six things worth watching

Judgment is not one thing, so it is not one score. Each dimension is probed, scored and evidenced on its own.

The wall

What is never scored.

Prose quality. Prompt syntax. Tool trivia. Speed. Tone of voice. And how much AI the candidate used.

That last one matters. "AI reliance" as a metric scores the wrong thing — it rewards volume and punishes the candidate who correctly decided the model was not the right instrument for a step. Judgment about the output is the test; usage is not a virtue.

Prosody is captured and used only as an integrity signal. It never reaches a score, and it never reaches the employer as one.

Never scored

Six things we throw away

Writing polish · prompt syntax · tool trivia · time to complete · vocal tone · number of AI turns.

Each one correlates with something other than judgment, and every one of them is easy to game.

No composite

A number you can't interrogate is a number you can't defend.

There is no single blended score by default. Six findings arrive separately because they mean different things: a candidate who demands a source for every claim and never refuses one of the answers is a specific, useful profile, and averaging it into a 74 destroys the only information in it.

It is also the defensible posture. Under NYC LL144 and the Illinois rules, "the model gave them a 74" is not an explanation. "At 07:40 the rate went into the memo unchecked, here is the capture" is.

What you get instead

Six findings, six excerpts

Demonstrated · partly demonstrated · not demonstrated — each with the moment attached and a human's confirmation on it.

See it scored

All six, on a real session

The sample report is a complete financial-analysis session: what was caught, what was missed, and the excerpt behind each call.

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.