Assessment design

What Should You Look for in a Candidate's AI Chat Log?

Read a candidate's AI chat log for acts, not volume. Four carry the signal: where the first message aimed, at understanding the problem or at the finished deliverable; which single claim the candidate stopped and demanded a source for, and whether they opened it; what the assistant produced that got thrown out, and on what grounds; and what got checked against something outside the conversation. Turn count and fluent phrasing are noise. A pasted log is a curated log, so confirm the load-bearing parts in a conversation about the work.

The takeMost teams ask for the log because they want to know who really wrote it, and that anxiety is the thing most likely to ruin the instrument. Read as a confession, it invites curation, and a curated log is thinner evidence than the deliverable you already had. Read as a record of acts, it is one of the few hiring artifacts that show process rather than product. Whether any of it predicts performance is unmeasured, and I would not claim otherwise. What it reliably does is make the conversation afterwards specific, and specific is where hiring stops being taste.

Where Olive fits

Open a role and see what the work shows

A pasted log is one evidence source, and it only ever shows what the candidate chose to paste. Olive reads the same behaviors out of a whole session, where the transcript sits beside a captured assessment workspace, a spoken or typed think-aloud and a written debrief, and returns six findings a human reviewer writes by hand, each anchored to a timestamped excerpt, with the candidate granted the identical report.

Rank your shortlist

What does a chat log show that the deliverable can't?

Acts, in order. The deliverable is a result; the log is the sequence of decisions that produced it. Four of those decisions carry almost all the signal: where the first move aimed, which claim got a source demanded for it, what the assistant produced that got thrown out, and what got checked against something outside the chat. The rest is texture.

Volume is the trap, because it is the one thing a log makes easy to count. Forty turns can be one person asking the same question forty ways and pasting the last answer into the document. Six turns can be a plan, a sourced correction, and half an hour of work done outside the window. Counting prompts ranks the first candidate higher, and so does reading for eloquence, since fluent phrasing is what the assistant supplies for free.

The research on knowledge work points at the same handful of moments. A survey of 319 knowledge workers who supplied 936 first-hand examples of using generative AI at work reported that it shifts the nature of critical thinking toward information verification, response integration and task stewardship, and that higher confidence in the tool went with less critical thinking while higher confidence in themselves went with more 1. Those are acts a transcript records and a finished document hides.

Both working shapes turn up in candidate logs. Across the Claude.ai conversations sampled for Anthropic's Economic Index, 57% of tasks were augmented and 43% automated 2. One candidate's log usually contains both on different subtasks, and the useful question is never which one they reached for. It is whether the split was decided or defaulted into. That is also the reason to ask for the log at all: a take-home that comes back polished carries none of this on its face.

Read the first three turns: framing or fetching?

Start there, because nothing has been prompted into the candidate yet. Look at whether the opening message states the problem, a constraint or a criterion in the candidate's own words, or whether it pastes the brief and asks for the finished thing. One of those is a person deciding what the job is. The other is an order being placed.

The same diligence take-home, opened two ways:

  • Fetching. "Here's the data-room summary and the brief. Write a two-page memo recommending whether to proceed."
  • Framing. "I'm deciding whether to recommend proceeding. Before drafting anything, list the claims in this packet that would flip the recommendation if they turned out to be wrong, and tell me which of them the packet doesn't settle."

Neither message is better written, and the second one is not more polite or more detailed in any way that matters. It names what would make an answer wrong before an answer exists. Mark the turn and move on.

What counts as framing, in a line each:

  • A criterion arrives before output is requested. A cost ceiling, the audience, what the decision turns on, what is out of scope.
  • The candidate restates the task. Not the brief pasted back, their own shorter version of it, sometimes wrong in a way that tells you something.
  • The unknowns get named. "I don't know the renewal rate and it matters" is worth more than a confident opening.

Two caveats worth holding. Framing can arrive at turn nine, after a first draft came back thin, and a candidate who course-corrects has demonstrated the behavior rather than failed it. And an opening that reads like a rubric entry is not automatically the stronger one: if three criteria get named and then ignored for the rest of the log, you have read a performance, and the later turns will show it.

Which claim did they demand a source for?

One claim, not all of them. Asking for sources in general produces a list nobody opens. What carries signal is a specific load-bearing claim being stopped on, a source demanded for that claim, the source actually opened, and what it did or did not support turning up in the deliverable afterwards. Three turns, usually, and they are easy to find.

The tell is the round trip. The assistant asserts something the recommendation rests on. The candidate names that specific thing and asks where it came from. A later turn reports what the source actually said, often correcting the assistant, and something in the deliverable moves as a result. "Cite your sources," answered with a list of links and no follow-up, is the appearance of sourcing with no evidence that anyone opened anything.

The rarer act is the one worth finding. In the same knowledge-worker survey, 114 of the 319 participants described cross-referencing AI output against external sources, while 23 described checking the sources the output itself had cited 1. Chasing a citation the assistant produced is the less common habit, and a transcript is one of the few places it becomes visible.

Which claim is load-bearing depends on the work, so read for the one the deliverable would collapse without:

  • Financial analysis. The growth rate the recommendation turns on, traced to the filing itself rather than to the assistant's summary of the filing.
  • Healthcare revenue cycle. The payer policy quoted back with its effective date, because a correct rule from last year is a denial.
  • Legal operations. The authority opened and read down to the paragraph, since retrieval is not verification and a real case can be pinned to a proposition it never held.
  • Data and analytics. The column definition checked in the schema, because a fluent explanation of the wrong field reads exactly like the right one.
  • Supply chain. The freight terms behind a landed-cost figure, since three quotes written on three different terms do not compare.
  • Marketing. The market-size number followed past the vendor's blog to the survey underneath it, with the sample size named in the brief.

One asked-and-opened claim beats a bibliography. Write down which claim it was, so a second reader has something specific to disagree with.

What got refused, and what got tested outside the chat?

The two hardest things to fake. A refusal is a turn where the candidate rejects something the assistant produced and says why: an assumption named, a framing dropped, a section cut for doing no work. An outside test is a turn reporting a result the assistant could not have produced, which then changes something in the deliverable. Both leave a mark in the record.

Count refusals as a rate against the answers actually taken up, never as a total. Otherwise thirty prompts with three pushbacks outrank six prompts with three, and you are back to rewarding volume through a side door. Style edits do not qualify. "Make that friendlier" is a preference. "Drop the second recommendation, the packet doesn't support it" is a judgment, and it earns the column. See how Olive measures this.

The outside test has two halves and most logs contain only the first. Someone writes "I checked the numbers and they hold," which is a claim about an act. The evidence is a result the conversation could not have produced, followed by a change: a figure recomputed and different, a page that contradicted the draft, a script that failed. If nothing downstream moves, the check either did not happen or did not matter.

This is where self-report stops being usable, the candidate's and yours alike. In METR's randomized trial, 16 experienced open-source developers working on repositories they had spent an average of five years on completed 246 tasks; afterwards they estimated the AI had cut their completion time by 20%, while the measured effect was a 19% increase 3. People read their own AI-assisted work in the flattering direction. Weigh the turn that shows a result above the turn that reports one, in the log and in the conversation after it.

Absence still proves nothing. A candidate who rewrote three paragraphs by hand in the document made a refusal that never touched the chat. If the artifact carries a version history, ask for it alongside the log and read the two together.

How do you read twenty logs the same way?

Fix the four reads as columns before you open the first log, and fill each with what the log says rather than with an impression. Mark each one demonstrated, partly demonstrated or not demonstrated, and put the turn number beside it, so every mark points at something you could show the candidate. Two readers filling the columns separately should land in the same place.

The wording of the column decides whether it can be filled twice the same way. "Used AI well" collects opinions and nothing else. Write the column as the act instead: named a specific claim and reported what the source said about it. That version is answerable with a yes, a no and a turn number, which is also what makes it reviewable later. Getting two reviewers to score an AI-use rubric the same way is a calibration problem before it is a wording problem, so mark two logs together before anyone marks alone.

Keeping the columns on observable turns has an old and specific reason behind it. The US federal Uniform Guidelines on Employee Selection Procedures, in force since 1978, hold that a selection procedure is supported by content validity "to the extent that it is a representative sample of the content of the job," and that a procedure "based upon inferences about mental processes cannot be supported solely or primarily on the basis of content validity" 4. "Seemed to think carefully" is an inference about a mental process. "Opened the filing and changed the number at turn 7" is a work behavior with a work product attached. Check with counsel on your own program.

Ask for the log in one format, in the brief, before anyone starts: the whole exchange in order with nothing removed, plus a short note of anything done outside the chat and what it changed. Say that the note and the log are graded and the polish is not. Then read every candidate's log against the same four columns, including the one whose prose reads normal to you. Whether the tool is permitted at all sits upstream of every one of these columns, so settle that first: the call on allowing AI on the take-home belongs to the job, not to the assessment.

What a pasted log can't tell you

Whether it is complete. A pasted transcript is a document the candidate assembled, so a missing turn is not evidence of a missing act, and a present turn can be written after the fact. Read it as the strongest claim a candidate is willing to make about their own process, then check the load-bearing parts in twenty minutes of conversation about their own submission.

It is also not a detection instrument, and reading it as one wastes the exercise and the candidate's trust in the same move. The question a log answers is what someone did with the assistant; the question of who typed which sentence is a different one, and whether AI detectors work in hiring is settled enough that no decision should rest there. A transcript read for authorship is a worse instrument than the one you already had.

Three more limits, stated out loud:

  • Missing turns are usually mundane. A second window, a phone, a colleague, a search that never entered the chat. Ask for the note; do not infer.
  • A clean log can be a curated log. Whoever pastes everything, dead ends included, looks messier than whoever pasted the good half. Say in the brief that dead ends are graded as work, or you have built a filter for tidiness.
  • One afternoon is one observation. The log shows what happened on this task with this assistant, not what happens when the model is confidently wrong about something the candidate has no way to check.

Tell candidates what the four reads are before they start. Someone who then aims at them has to frame the problem, demand the source, refuse something on substance and test it outside the chat, which is the behavior you were trying to measure in the first place. A read you would not be willing to describe to the person being read is the wrong read.

See what gets scored

Common questions

How long should a candidate's AI chat log be?

There is no target length, and publishing one turns the log into a word count. Six turns carrying a plan, a sourced correction and a check beat forty turns of rephrasing. If the brief has to say something, say the log should cover the whole task including the dead ends, and that length earns nothing on its own. Candidates who suspect volume is graded will pad, and a padded log takes longer to read and says less.

What if a candidate submits no log at all?

Ask before concluding anything. Some people work in a document with the assistant in a side window and have nothing to paste, some misread the brief, and some used no assistant. Give one chance to supply it, then grade what you have and move the process questions into the follow-up conversation. A missing log is a gap in your evidence rather than a finding about the candidate, and treating it as a finding penalizes careless reading of instructions instead of careless work.

Can you tell from the log who wrote the final text?

No log settles authorship, and authorship is the wrong thing to ask of one. A log shows what was asked, what came back, what was rejected and what was checked. Authorship of finished prose is a detection question, detection is unreliable, and a rejection resting on it is hard to defend to the person on the other end. The answerable version is different: did this person frame the problem, demand evidence for the claim that mattered, keep some of the work, and test something against the world?

Should you tell candidates what you're reading the log for?

Yes, in the brief, before anyone starts. Naming the four reads costs nothing, because none of them can be faked without being done: framing the problem, demanding a source for a named claim, refusing an output on substance and checking something outside the conversation are the acts themselves. Hidden criteria mostly select for candidates who guessed your preferences, and they make the feedback conversation worse, since you end up explaining a standard nobody had a way to meet.

How long does reading one log properly take?

Ten to fifteen minutes once the four columns exist, because you are not reading every turn. Skim for four moments: the opening exchange, any turn where a specific claim gets questioned, any turn where something gets thrown out with a reason, and any turn reporting a result from outside the chat. Note the turn numbers as you go. Most logs give up all four on one pass, and the ones that do not are usually short enough to finish in five.

What if the log shows the assistant did nearly all the work?

Where the line was drawn is the finding, and whether it counts as a problem depends on the role. A candidate who handed over drafting and kept the judgment about what the deliverable had to prove drew a defensible line. One who handed over the judgment too, then shipped the output with the edges tidied, did not, and the log shows it as a run of accepted answers with nothing refused and nothing checked. Read the split, not the proportion.

References

  1. 1. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks and Wilson (Carnegie Mellon University and Microsoft Research), CHI 2025, 2025. microsoft.com Survey of 319 knowledge workers who shared 936 first-hand examples: generative AI 'shifts the nature of critical thinking toward information verification, response integration, and task stewardship'; 'higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking'. Section 4.1.2 reports 114/319 participants cross-referencing external sources and 23/319 assessing the sources referenced in the output.
  2. 2. The Anthropic Economic Index Anthropic, 2025. anthropic.com Across the sampled Claude.ai conversations: 'we saw a slight lean towards augmentation, with 57% of tasks being augmented and 43% of tasks being automated.'
  3. 3. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity Becker, Rush, Barnes and Rein (METR), 2025. arxiv.org 16 developers completed 246 tasks in mature projects on which they had an average of five years of prior experience; after finishing, they estimated AI had reduced completion time by 20%, while the measured result was that allowing AI increased completion time by 19%.
  4. 4. Uniform Guidelines on Employee Selection Procedures, 29 CFR 1607.14 C (Technical standards for validity studies) U.S. Equal Employment Opportunity Commission, 29 CFR Part 1607, via Cornell Legal Information Institute, 1978. law.cornell.edu A selection procedure is supported by content validity 'to the extent that it is a representative sample of the content of the job'; a procedure 'based upon inferences about mental processes cannot be supported solely or primarily on the basis of content validity'.

4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.