Teams
Before You Open the Req, What Can Your Team Already Do With AI?
What can your team do with AI? Read what they shipped. Ask everyone for three deliverables from the last six weeks, one routine, one that took a day, one past their edge, and mark three things on each with the file open: where AI touched it, what got checked and against what, what they kept by hand. Ten minutes an artifact. Say up front it sizes a req and never enters a performance file. It sees six weeks of shipped work, nothing more. The empty column is usually checking.
The takeThe checked column comes back empty for a reason nobody on the team chose. Verification leaves no artifact behind. It reads as slowness, it gets credited to nobody, and at review the person who does it well looks like the person who is behind. Teams stop doing it long before anyone decides to. Hire against that gap and the new person inherits the same incentive, and their checking gets read as a pace problem too. The most useful thing the exercise produces is rarely a req. It is an argument about what review pays for.
Where Olive fits
Open a role and see what the work shows
The three marks on your map are a rough version of what Olive reads formally: what was framed before anything was generated, what evidence was demanded, what the person kept, and what was tested against something outside the conversation. Olive returns six separately-evidenced findings from a real 40-to-60-minute occupational assignment rather than from a self-assessment, and the candidate is granted the same report.
Rank your shortlistWhy won't a survey tell you what your team can do with AI?
Because the answer is a self-assessment with a career attached to it. The person most excited about AI reports capability nobody has watched work; the colleague who quietly rebuilt a reconciliation script ticks "basic". In Microsoft and LinkedIn's 2024 Work Trend Index, a survey of 31,000 knowledge workers across 31 markets, 52% of people who use AI at work are reluctant to admit using it for their most important tasks 1.
The honest respondent is wrong too, which is the part better wording cannot fix. METR ran 16 experienced open-source developers through 246 real issues, in repositories they had contributed to for years that average over 22,000 stars and a million lines of code. With AI tools allowed, the issues took 19% longer. The developers had forecast a 24% speedup, and after living through the slowdown they still believed the tools had made them 20% faster 2. People measured that closely misread their own direction of travel, and a five-point scale on a form will not recover it.
Usage telemetry answers a different question. A seat-license report or an assistant's activity dashboard tells you which accounts were open, never what happened to the output, and it is looking at a fraction of the tools in play: in the same Work Trend Index, 78% of AI users bring their own AI tools to work, rising to 80% at small and medium companies 1. Somebody who has been running an unapproved assistant for six months also has an obvious reason to keep the answer vague on a form carrying their name.
So stop asking what people can do and go read what they shipped. The artifacts are already sitting in a repository, a shared drive or a ticket queue. They are dated, they had a recipient, and nobody had to characterise themselves to produce one. An org-readiness quiz misses for the same reason from the other end: five questions about strategy and data maturity return a verdict about the company, and this req is about nine people.
Which three deliverables should you pull from each person?
Three things they shipped in the last six weeks, each with a real recipient: one routine deliverable, one that took a full day or more, and one where they were working past the edge of what they usually do. Three is the smallest number that separates a habit from a good week. Ask for the file, not for a description of the file.
What counts, by the work your team actually does:
- Software engineering. A merged pull request with its review thread attached, not a snippet.
- Financial analysis. A memo or model that carried a recommendation to someone who then acted on it.
- Marketing. A brief, a launch page or a deck with a market claim inside it.
- Legal operations. A redline or a position memo where the playbook and the counterparty disagreed.
- Revenue cycle. A worked denial or an appeal letter, with the outcome attached.
- Data and analytics. A result somebody made a decision on, plus the query behind it.
- Support and operations. An escalation reply that went to a customer under a person's name.
Ask for the working trail alongside each one where it still exists: the intermediate outline, the branch history, the draft with comments on it, the chat transcript if they kept it. Most people will not have kept the transcript, and that is fine. The marks below can be read off the artifact plus a short conversation, which is the whole reason this is an afternoon and not a project.
One instruction changes the quality of everything you get back: say why you are asking, in one sentence, before anyone sends a file. This is sizing a req, not building a performance file. A team that learns afterward that it was appraised will answer the next round of anything the way people answer surveys. If the map later tells you what AI is actually doing inside the role, that came from a room that was told the truth first.
Mark three things on each artifact: touched, checked, kept
Three columns, one row per deliverable, filled in a ten-minute conversation with the file open: where AI touched this, what was checked before it shipped and against what, and what the person kept for themselves on purpose. There is no rating. A cell is either filled with something specific and checkable, or it is empty, and empty is the finding.
The difference between a filled cell and an empty one is entirely concreteness:
- Touched. Filled: "generated the first draft of the appeal, then rewrote the clinical rationale by hand because the payer policy it cited was the wrong version." Empty: "used it a bit for wording."
- Checked. Filled: "re-added segment revenue against the 10-K and found the model had it a quarter out." Empty: "reviewed it carefully."
- Kept. Filled: "wrote the recommendation paragraph without the assistant, because the client reads that paragraph and nothing else." Empty: "I review everything myself."
Collect examples rather than ratings, for a reason the research spells out. A CHI 2025 study by Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers and gathered 936 first-hand examples of using generative AI in work tasks, and reports that it shifts the nature of critical thinking toward information verification, response integration and task stewardship 3. The same study associates higher confidence in the tool with less critical thinking, and higher self-confidence with more 3. That is the second reason a self-rating misleads: the people surest of the output are checking it least.
Keep the marks per artifact rather than per person. Somebody can run a disciplined check on a client deliverable and none at all on an internal one, and that difference is useful. It is also the difference between an inventory you can act on and a judgment about a colleague. If the exercise pushes you toward a competency model afterward, give AI its own competency only where the deliverable actually changed.
What does the finished map usually show about verification?
Production everywhere, checking thin and uneven. Expect the touched column filled on nearly every row, the kept column on one or two, and the checked column filled only where a past mistake had a name and an owner. Treat that as a pattern to test against your own map rather than a measurement. What the team is short of is people who catch a near miss.
The near miss is the costly one, and the people closest to it say so. In Stack Overflow's 2025 developer survey, the most-cited frustration with AI tools was "AI solutions that are almost right, but not quite," named by 66% of the developers who answered that question, with roughly 45% also naming that debugging AI-generated code is more time-consuming 4. A visibly wrong answer costs a minute. One that is wrong about which quarter a figure came from survives a read-through and fails at the point where somebody spends money on it.
Read the map at task grain, because that is the grain AI arrived at. Anthropic's Economic Index, built from roughly a million Claude.ai conversations matched onto the Labor Department's O*NET task list, reports that around 36% of jobs had some use of AI for at least 25% of their tasks while only about 4% used it for at least 75%, with 57% of tasks augmented against 43% automated 5. One person's "uses AI daily" can mean two tasks out of nine, and the seven it does not cover are where your req is really pointed.
Three patterns are worth naming when you read the finished grid. A person whose checked column is filled every time, undocumented, is the one to ask to write the check down, and nobody has yet. A team where nobody's checked column is filled on the highest-consequence work does not have a hiring problem yet; it has a review problem, and a new hire lands inside it. And a person with one thin artifact may simply not have been given consequential work in the window, which is a staffing fact rather than a skill fact. Hiring against the second and third of those is how a req solves the wrong thing. See how Olive measures this.
Write the req against the gap, not against a tool list
Turn each empty column into a line the job post and the interview can both use. An empty checked column on high-consequence work becomes a requirement written as an act: reconciles a figure to its source before it ships. An empty kept column becomes a judgment requirement. If both columns are full across the team, the gap is capacity, and you are hiring for volume rather than for skill.
Three rewrites, in the order the map produces them:
- Before: "Experience with ChatGPT or Copilot required." After: "You will draft the quarterly variance memo with an assistant and be accountable for every figure in it."
- Before: "Strong attention to detail." After: "You will be the person who reconciles a generated number to the filing before it reaches the partner."
- Before: "AI-first mindset." After: "You will decide which parts of this workflow stay manual, and defend that split at review."
The map also decides which of the two gaps you are looking at, which is the question hiring versus training for AI skills assumes you have already answered. A team with filled touched columns and empty checked ones does not need a new headcount to teach it prompting; it needs one person who models the check, or a review step that makes the check mandatory. A team with empty columns across three functions has a triage problem instead, and which roles need AI skills first is the order to run it in.
Then carry the same three columns into the round you run. If checking is what the map says is missing, the exercise has to make checking observable rather than describable: hand the candidate source material that settles nothing until somebody opens it, and grade what they opened and what changed as a result. Hiring for verification rather than production looks different in every field, and it is the round your map has just justified.
The inventory is a sizing instrument with a short memory. Six weeks of shipped work is the whole window, so it misses skill that had no occasion, judgment exercised in a meeting, and anything a person learned in a job before this one. Use it to write a sharper req, keep it out of anyone's file, and re-run it after the hire lands, because the only real test of the req is whether the next map looks different.
Common questions
How long does this take for a team of eight?
An afternoon, if you prepare first. Ask for the three files a day ahead so nobody spends the session hunting. Then book ten minutes per artifact, which is roughly half an hour per person and about four hours across eight people, plus an hour to read the grid afterward. The conversation is short because the artifact carries the detail. What blows the budget is trying to do it from memory, without the file open, which turns every answer back into the self-report you were avoiding.
What if people did not keep the AI chat transcript?
Most will not have, and the method still works. The three marks come off the artifact and a short conversation: which section reads as generated, which figure has a source behind it, which paragraph the person wrote themselves and why. A transcript is a convenience, not the evidence. If you want transcripts on the next round, say so in advance and give a reason people can accept, because asking retroactively for a record nobody was told to keep reads as an audit and gets answered accordingly.
Won't people just say what you want to hear in the conversation?
Less than on a form, because the artifact is open in front of both of you. "I checked it carefully" invites the obvious follow-up: checked against what, and what did it change. A person who ran the check names the source and the discrepancy straight away. A person who did not answers in adjectives. Nobody is being caught out here. You are asking a specific question that only has a specific answer, which is why it beats a scale.
Should you tell the team the inventory feeds a hiring decision?
Yes, before you ask for the first file. The alternative is that they find out later, and every subsequent request for evidence about their work gets a defended answer. Say what it is for, say the marks are per artifact rather than per person, and say plainly that it is not going into a performance file. Then honor that. An inventory collected under one purpose and used for another is the fastest way to make the second one impossible.
Does this replace assessing AI skills in candidates?
No. They answer different questions. The inventory tells you which gap the req exists to close, at your team's task grain, against work you can already see. A candidate assessment tells you whether one applicant closes it, and it has to create the evidence because there is no shipped work of theirs to read. Running the inventory first makes the candidate round cheaper and sharper, because you know which of the three columns you are hiring against instead of testing everything.
What if the map shows almost nobody uses AI on the work that matters?
Then the req is probably not the first move. A team that has not put AI near its consequential work has no internal standard for what good looks like, so a new hire arrives with nothing to be measured against and no review step to land in. Pick one high-consequence workflow, run it with an assistant for a month with the checking written down, and re-read the map. If the gap is still real after that, you now know what to write, and the new person has somewhere to land.
References
- 1. AI at Work Is Here. Now Comes the Hard Part (2024 Work Trend Index Annual Report) ✓ microsoft.com Survey of 31,000 full-time knowledge workers across 31 markets: 52% of people who use AI at work are reluctant to admit to using it for their most important tasks, and 78% of AI users are bringing their own AI tools to work, rising to 80% at small and medium-sized companies.
- 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ✓ metr.org 16 experienced developers recruited from large open-source repositories (averaging 22k+ stars and 1M+ lines of code) they had contributed to for multiple years, across 246 real issues, took 19% longer with AI tools allowed, having expected a 24% speedup and still believing afterward that AI had sped them up by 20%.
- 3. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers ✓ microsoft.com 319 knowledge workers shared 936 first-hand examples of using GenAI in work tasks; the study reports that GenAI shifts the nature of critical thinking toward information verification, response integration and task stewardship, and associates higher confidence in GenAI with less critical thinking and higher self-confidence with more.
- 4. 2025 Stack Overflow Developer Survey: AI ✓ survey.stackoverflow.co Among the 31,476 developers answering the question, the top-reported frustration with AI tools is "AI solutions that are almost right, but not quite" at 66%, with 45.2% naming "debugging AI-generated code is more time-consuming".
- 5. The Anthropic Economic Index ✓ anthropic.com Built from roughly one million Free and Pro Claude.ai conversations mapped onto O*NET tasks: roughly 36% of jobs had some use of AI for at least 25% of their tasks and only approximately 4% for at least 75%, with 57% of tasks being augmented and 43% automated.
5 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.