Assessment design

How Do You Evaluate a Portfolio When AI Made Half of It?

A portfolio half-made by AI is graded on its decisions, not its output. Assume an assistant made the surface, and ask for the record the candidate's craft kept while the work happened: the cut list, the constraint set, the review trail. Write the rating guide before you open anything, and run the same request past everyone, including anyone whose best work sits behind an NDA and can only be described. Rate specificity. Then treat the result as a filter: it decides who gets a real work sample, and stops there.

The takeThe uncomfortable part of this instrument is who it favors. A decision record survives where somebody kept one: version history, public review threads, specs that got written down. Plenty of strong people worked somewhere that kept none of that, and they will sound vague about choices they genuinely made. Nobody has measured how much of that vagueness is a missing habit rather than missing judgment. If this review ends up sorting a pipeline unevenly, it will not be polish doing the sorting. It will be who happened to work somewhere that kept receipts.

Where Olive fits

Open a role and see what the work shows

A portfolio review reads a decision the candidate made months ago, in whatever form they chose to keep it. Olive is an employer-purchased assignment of 40 to 60 minutes done with an AI assistant on occupational material, returned as six separately evidenced findings (problem framing, evidence sourcing, delegation boundary, working structure, output rejection, verification), each anchored to a timestamped moment, with the candidate granted the same report.

Rank your shortlist

Why Output Quality Stopped Being the Signal

Because a portfolio measures production, and production is the part AI now does. A polished case study, a clean interface, a working demo: all of it can be assembled in an afternoon by someone who could not defend a single decision inside it. What is still expensive is the choosing: what got cut, what got constrained, and what got sent back.

Working out which half was AI is a dead end. Fourteen detection tools tested against generated and human text reached 76% accuracy at best, none above 80%, and fell to roughly 30% on human-edited text and 15% on paraphrased text 1. Asking the candidate to estimate the split does not rescue it either. In a randomized trial on their own repositories, experienced developers took 19% longer with AI tools allowed and still believed afterwards that the tools had sped them up by about 20% 2. People are poor witnesses to their own AI-assisted work, which is a finding about memory rather than about honesty.

So stop trying to partition the artifact. Grade what a portfolio was always supposed to evidence (a sequence of choices made under constraint) and accept that rendering those choices has become cheap. If you also need to settle authorship on one specific project, How to Tell Who Did the Work in an AI-Built Portfolio is the narrower screen for that.

Ask for the Decision Record, Not the Deliverable

Send one request before the review: the record of decisions behind the piece, in whatever form the work already produced it. Not an explanation composed for you afterwards, but the artifact that existed while the work was happening: a cut list, a constraint doc, a review thread, a version history. Spend the session on that, and treat the finished piece as a reference document.

There is a precedent for running this as an instrument rather than a chat. The federal accomplishment record works the same way: candidates submit structured narratives of past work, trained raters score them against fixed dimensions using a written rating guide, and each entry names a verification contact, which is what keeps the descriptions honest. The method carries high content and criterion-related validity and generally shows little or no difference between subgroups 3.

Borrow the mechanics and drop the paperwork. One piece, one decision record, three questions per decision:

  • What were the other options? A real decision has a losing side. "There wasn't one" means it was not a decision, it was a default.
  • What made you drop them? The reason is the artifact. "It was slower" is a memory. "It was slower because every page refetched the whole list, which only showed up at 400 rows" is a decision.
  • What would have made you choose differently? The hardest of the three to borrow and the fastest to rate. Someone who owns the choice names the condition in one sentence.

Rate specificity, not eloquence. An assistant writes the fluent version of any of these answers; it does not know which four hundred rows broke the page. The same separation works on a submitted assignment, which How to Grade Take-Homes When Every Submission Comes Back Polished takes further.

What Counts as the Candidate's Own Work, by Craft

Something different in every field, which is why a single walkthrough script misses. For a product manager it is the cut decisions and the options that died. For a designer it is the constraint set that existed before the first frame. For an engineer it is the review trail: what got sent back, and on what grounds. Ask for the one their craft already keeps.

Product management. The contribution is subtraction: what did not get built, and what would have had to be true to build it. Ask for the spec and read its exclusions before its features. An assistant resolves every ambiguity in a vague request cheerfully and returns a spec with no exclusions at all, so a "not doing" list with reasons attached is the tell. Then ask which requirement they refused to write down until somebody answered a question.

Design. The constraint set, fixed before the first frame. Ask for the file rather than the export: version history, the frames that lost, the component that got rebuilt. A designer who owns the work names the constraint that killed the better-looking option (a contrast ratio, a legal string that could not shrink, a cold-load budget). Generated comps are constraint-free, which is exactly why they look better and work worse.

Engineering. The review trail. Read what got sent back rather than what got merged. 66% of developers name "almost right, but not quite" as their biggest frustration with AI tools, and 45% say debugging AI-generated code takes more time than writing it themselves 4, so anyone genuinely shipping AI-assisted code carries a history of rejections. A branch with no reverted commits, no review comment argued back on and no abandoned approach is not a clean record. It is an absent one.

Writing, marketing and analysis. For writing and marketing, the brief and the kills: the concept that was rejected, and by what argument. For analysis, the reconciliation: the two sources that disagreed, which one they took, and what the answer would have been the other way. Every honest analysis has that moment. A generated one has a confident chart and no memory of the rows that were dropped.

The artifact changes by craft; the question does not. Pick one per role and ask everyone for it, because the comparison is worthless when one candidate hands you a repository and the next hands you a PDF.

Write the Rating Guide Before You Open the Portfolio

Decide what a strong answer contains before you see one, or you will end up rating polish, the one variable AI has flattened. Three or four dimensions, four or five lines each describing a strong answer and a thin one, written down and used unchanged for every candidate on the role. A portfolio review that decides who advances is a selection procedure, with everything that implies 5.

Four dimensions cover most crafts. What constraint was named before anything got generated. What claim was checked against something outside the work. What was refused, and on what grounds. What the candidate kept for themselves instead of handing over. Write the strong and thin versions of each in the candidate's own vocabulary and stop there; a longer rubric is not a better one if nobody reads it twice.

Then apply it identically. A step used to decide who advances has to be job-related and applied consistently across candidates 5. In practice that means the same artifact request, the same questions in the same order, and notes written in the same place for everyone on the role. It also means the review has to survive the failure that actually happens here, which is not a candidate fooling you. It is a reviewer deciding a case study "reads like AI" and quietly marking it down. How to Stop Managers Rejecting Candidates for 'Sounding Like AI' makes the full argument; the short version is that a hunch applied unevenly across a pool is the part that gets challenged.

What a Portfolio Review Still Can't Tell You

Whether they can do it again. A portfolio review captures a candidate describing decisions made months ago, with time to prepare and no way for you to watch a choice being made. It is a strong filter and a weak gate: good enough to decide who gets a real work sample, not good enough to decide who gets an offer.

It also over-rewards one kind of career. Agency work ships under the agency's name, enterprise work sits behind an NDA, and plenty of strong people have nothing public because their best work belonged to a team. Requiring a public portfolio filters for who had permission to build in the open, not for who can do the job. Offer a described-work path up front, rated on the same guide, and say so in the invitation rather than as an exception granted on request.

The deeper limit is that the review is a claim about judgment rather than an observation of it. If the role turns on how someone works with an assistant in the moment (and most now do), the next instrument is a short piece of occupational work done with one available, not a longer conversation about a finished piece. Is a Prompt Engineering Certificate Worth Anything, or Should You Give a Work Sample? and How to Redesign Your Interview So AI Use Is Signal arrive at the same place from opposite directions, and Should You Measure Speed With AI, or Judgment? settles what to rate once you get there.

Read the portfolio as evidence about decisions rather than about craft. The pieces got easier to make. Deciding what to make did not.

See what gets scored

Common questions

Should you ask a candidate how much of the portfolio AI produced?

Ask, but treat the answer as context rather than measurement. People misjudge their own AI-assisted work: in a randomized trial, experienced developers took 19% longer on real issues with AI tools allowed and still believed afterwards that they had been about 20% faster 2. The percentage they give you is a memory. What they cut, what they constrained and what they sent back is checkable, and it is the thing you were trying to learn from the percentage anyway.

What if all the candidate's work is under an NDA?

Run the same review on work they can describe without showing. The decision record still exists in their head (the options killed, the constraint that fixed the design, the thing they refused to ship), and the three questions work on a described project almost as well as on a shown one. Offer that path in the invitation rather than granting it on request, because the candidates who most need it are the least likely to ask.

How long should an AI-assisted portfolio review take?

Thirty to forty minutes on one piece and its decision record. Two pieces at fifteen minutes each is worse than one at forty; the specificity you are rating only shows up past the first layer of explanation. If you still cannot tell after forty minutes, more questions will not fix it. The next step is a work sample, which measures something a walkthrough structurally cannot.

Can you reject a candidate for using AI on a portfolio piece?

Not on that alone, and it is a weak rule to run. How much AI got used is not the signal; what got framed, refused and checked is. A candidate who generated most of a build, caught what was wrong and can tell you how they caught it has shown more than one who hand-wrote everything and tested none of it. Whatever rule you do apply has to be job-related and applied to every candidate identically 5.

Does this work for a junior with only class projects?

Yes, and the questions get easier rather than harder. A class project has a constraint set, a scope that was cut when the deadline moved, and an approach that was abandoned in week three. Ask for those. Federal guidance on rating past accomplishments is explicit that entry-level candidates should be credited for work outside paid employment 3, which includes coursework, volunteer projects and anything they built for themselves.

References

  1. 1. Testing of detection tools for AI-generated text Weber-Wulff et al., International Journal for Educational Integrity, 2023. edintegrity.biomedcentral.com 14 detection tools tested; best 76% accuracy, none above 80%; roughly 30% on human-edited and 15% on paraphrased text, with a bias toward classifying generated text as human-written.
  2. 2. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity METR, 2025. metr.org 16 experienced developers on 246 randomized real issues took 19% longer with AI tools allowed, and still believed afterwards that the tools had sped them up by about 20%.
  3. 3. Assessment and Selection: Accomplishment Records U.S. Office of Personnel Management, 2026. opm.gov Structured narratives of past work rated against fixed dimensions with a written rating guide and a verification contact; high content and criterion-related validity, generally little or no subgroup difference, and entry-level applicants credited for non-paid experience. Undated guidance, verified 2026-08-24.
  4. 4. 2025 Stack Overflow Developer Survey: AI Stack Overflow, 2025. survey.stackoverflow.co 66% name AI solutions that are almost right but not quite as their biggest frustration; 45% report debugging AI-generated code is more time-consuming than writing it themselves.
  5. 5. Employment Tests and Selection Procedures U.S. Equal Employment Opportunity Commission, 2024. eeoc.gov A step used to decide who advances is a selection procedure: it must be job-related and consistent with business necessity, and applied consistently across candidates.

5 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.