Assessment design

How Do You Screen for AI Judgment in Finance or Marketing?

Screening for AI judgment in finance or marketing takes the same six behaviors as engineering, read off a memo or model instead of a diff. Give the candidate an hour on a task from the function's material, an assistant open, and one number contradicted by a source you must open yourself. Grade the moves, not the prose, and grade the six separately: what got framed first, which claim got a source demanded, what was kept, what was built between the brief and the deliverable, what was refused, what was checked.

The takeConfidence is the variable nobody screens for, and it runs the wrong way. Trust the assistant more and you check less; trust your own skill more and you check more. The functions with no automatic check behind them are, as far as I can see, both the desks where a fluent draft gets taken on authority and the desks now advertising for fluency with the tool. If both of those hold, the trust curve and the safety curve point in opposite directions, and a requisition that screens for fluency is selecting against the habit that keeps a memo true.

Where Olive fits

Open a role and see what the work shows

If you build this per function yourself, the recurring costs are the answer key and the evidence trail, and both have to be written again for every occupation you add. Olive ships authored cases for financial analysis, marketing, management consulting, product management and eight further occupations, returning six separately-evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each written by a human reviewer, anchored to a moment in the session, and granted to the candidate in the same document.

Rank your shortlist

What does AI judgment look like outside engineering?

The same as inside it. Framing the problem before generating, demanding a source for the claim that matters, keeping the judgment you should not hand over, building something between the brief and the deliverable, refusing an answer on stated grounds, and testing a claim against something outside the conversation. None of those is about code. All six are visible in a memo, a model, a brief or a spec.

The coding-adjacent answer is everywhere because engineering hands you three free graders. A diff shows what the assistant wrote and what the candidate changed. A test suite says whether the thing works without anyone's opinion. A compiler refuses to be talked around. Finance, marketing, consulting and product have none of that: a memo compiles no matter what it says, and a deck passes every test in the world.

So the instinct to skip the assessment for those requisitions is understandable and backwards. Those functions are where a confident wrong answer travels furthest, because nothing downstream catches it. A growth figure lands in a board deck. A competitor claim lands in positioning. An assumption lands in a spec that three teams then build against. The check happens at the desk or it never happens.

Decide first which roles genuinely need this, because not every opening does. Then keep one rubric across all of them and let only the task change per function. That split is the thing worth building around, and it is what lets a marketing result and a finance result mean the same words.

What do the six behaviors look like in a memo or a model?

Name the artifact first, then the observable. Marketing's deliverable is a report translating complex findings into written text, a forecast of sales and marketing trends, an analysis of competitors' prices and methods 4. Finance's is a memo or a model over a source packet. Consulting's is a sizing and a recommendation. Product's is a spec. Each has one moment where a claim either gets opened or gets copied.

BehaviorFinanceMarketingConsulting or product
Framed before generatingNamed what the memo decides and what would make it wrongWrote the positioning question and the audience before any copyStated the decision the deck serves before sizing anything
Demanded a sourceAsked which filing the growth number came from, then opened itAsked where the market-size figure came from, then opened itAsked which of three quotes the estimate rests on
Kept the judgmentRecomputed the multiple by hand instead of taking the model'sChose the segment after reading the raw surveyWrote the recommendation sentence unaided
Built something in betweenAn argument outline before any paragraphA criteria list for the message before any line of copyA one-page issue tree before slides
Refused an outputCut a paragraph because the evidence behind it was thinDropped the most on-message statistic once its source was readRejected a framing that assumed the answer
Checked outside the chatRe-added the column and found the total wrongOpened the vendor report the number was pulled fromTested the assumption against a public filing

That is a translation, not a menu. The row is identical everywhere; the anchor sentence under it is local, and it has to be written by someone who does that job. Get the anchors wrong and you have four rubrics wearing one name.

Two things get scored by accident in non-technical functions, and neither should. Prose quality, because the assistant is good at prose and the candidate's sentence rhythm is not what is being tested. And volume of AI use, because a candidate who decided the model was the wrong instrument for the pricing step and did it by hand has demonstrated the exact thing you are looking for. Grading a polished take-home breaks on precisely that confusion.

How do you build the task without a test suite?

Put the failure in the material rather than in the brief. Build a source packet where one number the deliverable depends on is contradicted by something you have to open: a footnote, a second tab, a methodology page. The assistant will summarize the packet fluently and carry the wrong number with it. Whether the candidate leaves the chat and opens the file is the assessment.

The packet has to come from the job. A selection procedure is supported by content validity only to the extent it is a representative sample of the content of the job, and where it samples a work behavior, its manner, setting, level and complexity should closely approximate the work situation 3. A marketing candidate sizing a semiconductor market has been tested on nothing they will ever do, and you would struggle to say otherwise if asked.

Forty-five to sixty minutes is enough, and the packet should be big enough that nobody reads all of it. Scarcity of time is what forces the delegation decision you want to watch: something has to be handed over, and which thing tells you more than the finished memo does. Cap it, say the cap in the brief, and let the assistant be genuinely willing to write the whole deliverable.

The same design sits behind hiring for verification rather than production. Production is now free, so the exercise has to be built around the step that is not, and the step that is not free is the one where somebody leaves the conversation to find out whether a sentence is true.

Which behavior carries the most signal?

Verification, by a distance. It is the behavior most often skipped, the one that separates two candidates whose deliverables read identically, and the one that gets rarer as trust in the assistant grows: a survey of 319 knowledge workers found that higher confidence in generative AI went with less critical thinking, while higher confidence in one's own skill went with more 1.

That same study describes the shift underneath the hiring problem. Generative AI moves the nature of critical thinking toward information verification, response integration and task stewardship 1, which is to say that the part of the job that survives is the part most assessments never look at, and the part a fluent draft makes easy to skip.

Retrieval does not fix it, and buying a specialist tool does not either. The first preregistered evaluation of the AI legal research products sold to a document profession (systems built on retrieval-augmented generation and marketed as avoiding or eliminating hallucination) found they hallucinated between 17% and 33% of the time 2. Those tools have a curated corpus behind them. A general assistant reading your source packet has less.

Verification is an act, never a claim. "I would double-check that" in a debrief is worth nothing; a recomputed figure, an opened page, or a number that changed because of what was found is worth everything. If you plant one error, plant it where the correct answer changes the recommendation, and watch whether the candidate catches it rather than whether they say they would.

How do you grade a memo two reviewers score the same?

Write the anchor before the first candidate and keep the outcome words coarse. Three words per behavior (demonstrated, partly demonstrated, not demonstrated), plus one sentence per function saying what satisfying it looks like there. Two reviewers reading the same session should land on the same three words. Where they do not, the anchor is ambiguous, and the anchor is what gets rewritten.

Resist the total. Six behaviors averaged into one figure is a claim about a person that none of the six supports, and it is the first thing anyone challenging a rejection will ask about. Report the six separately, each with the moment it rests on: the timestamp, the transcript turn, the version of the memo where the number changed. That record is also the only honest answer to why one candidate was preferred.

The reviewer has to be someone who does the work. A marketing lead cannot tell whether an analyst's multiple is wrong, and a finance lead cannot tell whether a positioning claim is unsupported by the survey it cites. Pair each function's case with a reviewer from that function, and give them the answer key rather than asking them to improvise one at 9pm.

Agreement is the part that decays, and it decays faster with every function added. Two reviewers scoring the same rubric the same way needs a double-marked sample and a short conversation about the disagreements, monthly. When two finalists hand in work of the same quality, the record of how each of them got there is the only thing left to read.

See what gets scored

Common questions

Doesn't a finance or marketing candidate just need to know the tools?

Tool familiarity is the cheapest thing to acquire and the least predictive thing to test. Naming four assistants proves exposure; it says nothing about whether a claim gets opened before it lands in a memo. Test the behaviors instead, and treat the tool list as trivia. A candidate who used one assistant carefully will outperform one who used six and checked nothing, and only a work sample shows you which you have.

How long should a non-technical AI work sample be?

Forty-five to sixty minutes, with a hard cap stated in the brief. Long enough for the delegation decision to bite, short enough that completion holds. Longer tasks shed candidates, and the ones who decline are not a random sample of your pipeline. Give the packet more material than anyone can read in the time, because triage under scarcity is part of what you are watching.

Can one case cover both a finance and a marketing role?

Only if a job analysis shows both roles perform substantially the same work behaviors on substantially the same material. Usually they do not: one reads filings and models, the other reads survey data and competitor pricing. Share the rubric, the outcome words, the time cap and the review protocol. Author the case, the source packet and the answer key per occupation, and budget about a day each plus a rewrite after the first few candidates.

What if the candidate barely uses AI during the exercise?

That is a result, not a failure. Someone who judged the assistant was the wrong instrument for the pricing step and did it by hand has demonstrated delegation judgment. Volume of use is not a virtue and should not appear in the rubric. What matters is whether the split between handed-over work and kept work was deliberate, and whether the candidate can say why at the point they made it.

Is an AI-open task fair to candidates whose employers ban AI?

Fair, if the brief is explicit and the task does not assume fluency with one product. Say which assistants are permitted, say that using one is expected rather than penalized, and keep the interface out of the rubric. Somebody meeting an assistant for the first hour will look slower; they can still frame a problem, demand a source and check a number. Score those, not the keystrokes.

Should this replace the interview or sit beside it?

Beside it, and behind it in the loop. Run the work sample after a first conversation so you are not asking an hour of unpaid work from everyone who applies, and run it before the panel so the panel has something real to ask about. The strongest follow-up questions in a final round come from a moment in the recorded session, not from a resume line about tools.

References

  1. 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Microsoft Research and Carnegie Mellon University (CHI 2025), 2025. microsoft.com Survey of 319 knowledge workers and 936 first-hand examples of generative AI use: higher confidence in the tool predicts less critical thinking, higher self-confidence predicts more, and the nature of critical thinking shifts toward information verification, response integration and task stewardship.
  2. 2. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Stanford RegLab and Stanford Institute for Human-Centered AI (arXiv:2405.20362), 2024. arxiv.org First preregistered evaluation of retrieval-grounded legal research products: tools from LexisNexis and Thomson Reuters, marketed as avoiding or eliminating hallucination, hallucinated between 17% and 33% of the time.
  3. 3. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 U.S. Equal Employment Opportunity Commission (Legal Information Institute, Cornell Law School), 1978. law.cornell.edu Content validity holds only to the extent the procedure is a representative sample of the content of the job, and where a work behavior is sampled, the manner, setting, level and complexity should closely approximate the work situation.
  4. 4. Market Research Analysts and Marketing Specialists (13-1161.00) O*NET OnLine, U.S. Department of Labor Employment and Training Administration, 2026. onetonline.org Occupation tasks include preparing reports that translate complex findings into written text, forecasting and tracking marketing and sales trends, and gathering data on competitors' prices, sales and methods of distribution.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.