Assessment design

AI-Collaboration Rubric: Demonstrated, Partly, Not

Interviewers rating a session where a candidate worked with an assistant. One sheet per candidate, in your own file.

Five behaviors a candidate has to perform themselves, rated on three levels, with every level written as something a reviewer can point at in the record. Fluency and structure are the assistant's contribution and earn nothing. One blank sheet per candidate, no total and no order between candidates. Send the behaviors and the level names with the exercise brief, keep the anchors and any weights with the reviewers, and send the findings after the decision.

See all documents

Five behaviors, three levels, one blank sheet per candidate. Every level is written as an act with a trace, so two reviewers can point at the same moment and agree it is there or concede it is not.

The five behaviors

BehaviorThe question it asksWhat settles it
FramingDid the first move go after the problem, or straight at the deliverable?The opening exchange, before any output exists
EvidenceWas a source demanded for the claim the answer rests on, and opened?What was retrieved, not what was cited
BoundaryWhat was kept, and what was handed over?The acts in the record the candidate did themselves
RefusalWas anything the assistant produced rejected on substance?A section cut, with the reason given at the time
VerificationWas a claim tested against something outside the conversation?A number that moved, or a claim that survived the check

Fluency, structure and length are the assistant's contribution and earn nothing. Neither does volume, in either direction: how much of the work was handed over is not a level of anything.

Refusal and verification sit on one row in the rubric these come from, and are split here because their traces differ. Cut a row a role's record cannot show, rather than rating it blind, and do not add a sixth that restates one already on the sheet.

The grid

BehaviorNot demonstratedPartly demonstratedDemonstrated
FramingThe first request is for the deliverable, and no question about the decision the work serves is asked at any pointOne clarifying question is asked, about scope or format, and the answer changes nothing about what is asked for nextBefore generating, states the decision the work serves and one condition that would make the answer wrong, and the next request changes because of it
EvidenceThe load-bearing figure arrives as the assistant supplied it and travels into the answer unopenedThe source is named or footnoted. Asked where the figure sits in it, the candidate cannot sayOpens the source, locates the figure, compares its definition against the question being asked, and changes the claim where the two do not agree
BoundaryEvery step is handed over, including the one the role exists to doSays afterwards which parts were handed over, with no decision visible at the timeNames a step done by hand and why the assistant was the wrong instrument for it, and the record shows that step done by hand
RefusalEverything generated survives into the answer, in the order it arrivedSomething is cut for length, tone or format, and nothing is cut for being wrongRejects at least one direction on stated grounds, says what was wrong with it, and the answer changes shape as a result
VerificationNo claim is checked against anything outside the session, and correctness is argued from the answer's own confidenceOne input is spot-checked and the result moves nothingRebuilds one load-bearing claim from the material in the room and says what changed when the result came back different

An adjective is what two reviewers disagree about, because each reads it from their own practice. Strong evidence, good rigor and shows judgment are not levels. Every cell above names something a reviewer can point at in the record, or concede is not there.

Anchors in the role's own evidence

The grid gets you a shared shape, not a shared standard, and the standard is the half that does not copy: what a competent colleague in that job asks to see before signing. Write the Demonstrated cell in that evidence first, with people who hold the job, so a reviewer has a concrete behavior to refer to rather than an adjective 1. The other two follow: partly demonstrated is asked for and not opened, not demonstrated is neither.

RoleThe load-bearing claimDemonstrated
Software engineeringHow the code behaves in a case nobody wrote downA test reproducing the reported bug is written first, watched to fail, then made to pass
Financial analysisA figure that feeds the recommendationThe figure is rebuilt from the filing, comes out lower, and the recommendation moves with it
Data and analyticsThe direction or the significance of a resultThe query is re-run against the raw table, with row counts and exclusions stated
Marketing and communicationsAn attributed statisticThe source is opened, the sample turns out to be three years old, and the claim in the brief changes
Strategy consultingA market size or a growth rateThe source is named and opened, the figure located, and its definition compared against the question asked
Legal operationsA case name or a clause positionThe case or the executed contract is opened and the clause quoted

A structured interview has two clauses, not one: the same predetermined questions in the same order, and every response evaluated against the same rating scale and the same standards for acceptable answers 2. Teams ship the scale and skip the standard. The Demonstrated column is the standard.

Almost right is the failure these anchors are pointed at. In a 2025 developer survey the most common reported frustration with these tools was output that is almost right but not quite 5, and purpose-built legal research systems from two major vendors returned hallucinated content between 17% and 33% of the time in a 2024 study, against vendor claims of hallucination-free retrieval 6. A cited-looking claim taken at face value is careless by the standards of the job, not by the standards of the marketing.

The rating sheet

Copy this into your own file, one sheet per candidate, and write beside each behavior the level and the moment it rests on.

BehaviorWhat to record
FramingThe level, and the moment in the session it rests on
EvidenceThe level, and the source that was opened or was not
BoundaryThe level, and the step done by hand
RefusalThe level, and what was cut, with the reason given at the time
VerificationThe level, and what was checked against what

Put the anchor version and the date at the top. A rubric that changed mid-round cannot be compared across the candidates who took it before and after, and a level with no moment cited cannot be argued about, only defended.

The sheet carries no number. Three named levels across five behaviors do not add up to anything, and folding them into one figure hides the disagreement between your reviewers, which is the information you were collecting.

Order of work

StepWhat happens
1Write the Demonstrated cell for each behavior in your own evidence, before the questions exist
2Each reviewer observes, records and rates alone, and discussion comes only after both sheets exist 1
3Rate one behavior at a time across candidates, never one candidate end to end
4Cite the moment behind every level recorded 1
5Version the sheet, log what changed, and carry the version onto every rating made under it

Step two is the one teams skip, and it decides what every agreement figure afterwards is worth: talking first produces one reviewer's opinion held by two people. Step three is the cheapest fix on the list, because reading one candidate end to end lets the first strong level color the next four.

Rate the trace, never the account. In a randomized trial of 246 real issues from experienced developers' own repositories in 2025, the work took 19% longer with the tools allowed, and the same developers still believed afterwards that the tools had sped them up by 20% 7.

Calibration

Both reviewers rate the same three sessions alone, then meet and work only the disagreements. Whoever changes their mind says which sentence of the anchor moved them, and if neither can point at a sentence the anchor is at fault rather than the reviewer. Rewrite it in the room, with the disputed session's actual behavior in it as the worked example. Recalibrate when a reviewer joins, when the task changes, and quarterly regardless.

Measure agreement per behavior, never averaged across the five, because an average will hide one behavior sitting at 0.2 and that is where every disputed decision comes from. Two reviewers who both mark Demonstrated most of the time reach a high raw agreement rate while agreeing about almost nothing, which is what Cohen's kappa corrects for: 0.60 to 0.79 is moderate, 0.80 to 0.90 is strong, and below 0.60 the two are measuring different things 4. Under the same paper's floor of about 30 comparisons, report exact agreement and the list of what split, and call it a calibration check rather than a reliability estimate 4.

Across 19 past-behavior interview studies reviewed in 2014, anchored scales showed higher validity, .35 against .26, and higher agreement between reviewers, .77 against .73 3. A comparison across studies, not a controlled test.

The candidate's copy

Send the five behaviors and the three level names with the exercise brief. Keep the anchors you wrote in your own evidence, and any weight you put on a behavior, with the reviewers.

Sending the criteria does not weaken the instrument, and the reason is in the grid: every Demonstrated cell is an act with a trace, so someone who reads the sheet and then goes and opens the source has done the thing the level is for. What a published rubric usually loses is completeness, and completeness is on no row here.

Sending the anchors and the weights does weaken it, because together they are a specification, and a specification handed to a candidate working with a model comes back satisfied line by line. Where the tool is permitted, correctness of the final answer stops being sufficient evidence and the judging becomes the readable part 8.

The tell that you sent too much is convergence: submissions arriving in the same structure under the same headings, nothing argued. At that point the exercise measures who thought to paste the sheet in.

After the decision, send the rest. There is nothing left to optimize to, and in a 2025 vignette experiment with 921 working-age Austrians an unexplained automated rejection scored lowest on every outcome measured, while an explained one drew the same fairness ratings as a human rejection carrying no explanation 9. Hypothetical rejections on an online panel: an ordering among unhappy outcomes rather than a route out of them.

Duties

Where this sheet decides who advances, the burden below is the employer's.

Once a complaining party shows a particular practice causes disparate impact, demonstrate that it is job related for the position in question and consistent with business necessity. 42 U.S.C. 2000e-2(k), effective November 21, 1991.

So every cell names an act from the job rather than an adjective. An adjective cannot be shown to be job related, because nobody can say what it asked for.

The session is administered material, and the duty below reaches how it is run.

Select and administer any test so the result reflects the skill it measures, not a candidate's impaired sensory, manual or speaking skills, unless those are what it measures. 29 CFR 1630.11.

In practice: offer the same material in a form the candidate can use, and rate the act rather than the pace or the manner of it. Someone who reads the source with assistive software has opened the source.

Limits

An interview answer is an account of work, not the work. This sheet rates what the record contains, and someone who did none of it can describe all of it fluently. Two repairs, both cheap: ask for the artifact rather than the summary, and put the material in front of the candidate live, so the check has to happen in the room.

Agreement is not validity. Two reviewers can be trained into near-perfect agreement about a document showing none of the five acts, and the figure that comes out will look like rigor while describing nothing. Calibration measures whether two people read the same evidence the same way, never whether the evidence was there.

Nothing here predicts performance. What it buys is comparability across candidates asked the same questions against the same standards, a record resting on observable behavior rather than on an inference about mental processes, and a specific next step. Not demonstrated on a behavior means the act is not in this record, under these conditions, with this much time. It does not mean the person cannot. Say that out loud in the debrief, because a sheet of five levels read as a verdict on a person is worse than no sheet at all: it lends arithmetic to a judgment that has not earned it.

Take it

The file and the credit

The publishing entity legal name and postal address are not filled in yet, and both sit inside the disclaimer every packaged format renders. No file is emitted until they are.

Credit line, to paste beside anything you quote from this document.

<!-- Olive template. Licence: https://olive.is/licence/template/1-0 -->
<p><a href="https://olive.is/answers/tools/ai-collaboration-rubric/" rel="nofollow">Olive</a>, AI-Collaboration Rubric: Demonstrated, Partly, Not, version 1.0.0, checked 2026-08-26.</p>

Licence

Important notes

Not a substitute for the advice of an attorney. This is a starting document for your own attorney to edit. It applies no law to your facts. No attorney has reviewed it for your state or your facts, and Olive is not your lawyer.

Published by Olive Independent Study, Inc., [DELAWARE INCORPORATING ADDRESS], United States. Contact hello@olive.is. A person reads every complaint and answers within ten working days. All concerns that Olive has engaged in the unauthorized practice of law are referred to the North Carolina State Bar, wherever the complaint came from.

Checked 2026-08-26 against the sources listed in this file. Version 1.0.0.

This text disclaims no warranty, caps no liability, waives no remedy, and names no court or state for a dispute. Those absences are deliberate.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Read this notice at its own URL · Licence

Packaged files

  • The publishing entity legal name and postal address are not filled in yet, and both sit inside the disclaimer every packaged format renders. No file is emitted until they are.

Checks

  • Sources last re-opened August 26, 2026.
  • Next review due February 26, 2027.

References

  1. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Has subject-matter experts write example behaviors for each proficiency level so raters have concrete behavior to refer to rather than adjectives; requires notes of sufficient quality and quantity to document the reasoning for each rating on each competency; and has panel members individually observe, record and evaluate before discussing and reaching consensus.
  2. Assessment and Selection: Structured Interviews U.S. Office of Personnel Management, 2024. opm.gov In a structured interview all candidates are asked the same predetermined questions in the same order, and all responses are evaluated using the same rating scale and standards for acceptable answers.
  3. The Structured Employment Interview: Narrative and Quantitative Review of the Research Literature Personnel Psychology, 67(1), 241-293 (Levashina, Hartwell, Morgeson and Campion), reporting Taylor and Small (2002), 2014. doi.org Anchored rating scales pay off on the reviewer's side: .35 against .26 for validity and .77 against .73 for interrater reliability across 19 past-behaviour interview studies.
  4. Interrater reliability: the kappa statistic McHugh, Biochemia Medica 22(3), 2012. pmc.ncbi.nlm.nih.gov Kappa corrects raw percent agreement for the agreement two raters would reach by chance; the interpretation table puts 0.60-0.79 at moderate and 0.80-0.90 at strong, states that any kappa below 0.60 indicates inadequate agreement among raters, and gives a heuristic floor of about 30 comparisons.
  5. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co The most common reported frustration with AI tools, at 66%, is output that is almost right but not quite; 46% of developers distrust the accuracy of AI output against 33% who trust it.
  6. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools Magesh, Surani, Dahl, Suzgun, Manning and Ho, Stanford RegLab and Institute for Human-Centered AI (arXiv:2405.20362), 2024. arxiv.org Purpose-built legal research tools from LexisNexis and Thomson Reuters each hallucinate between 17% and 33% of the time, against vendor claims of hallucination-free retrieval.
  7. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 246 real issues from experienced developers' own repositories, randomized to allow or forbid AI tools: the work took 19% longer with the tools allowed, and participants still believed afterwards that the tools had sped them up by 20%.
  8. Reimagining Assessment in the Age of Generative AI: Lessons from Open-Book Exams with ChatGPT arXiv:2605.12363 (Mahmoud, single author), 2026. arxiv.org When the tool is permitted, correctness of the final answer stops being sufficient evidence, and the judging becomes the readable part.
  9. Rejected by an AI? Comparing job applicants' fairness perceptions of artificial intelligence and humans in personnel selection Frontiers in Artificial Intelligence, 2025. frontiersin.org An unexplained automated rejection is the worst case measured, and an explained one rates alongside an unexplained human rejection.

9 sources, numbered by first appearance. How Olive sources claims

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.