Answers
Assessment design
Work samples, take-homes and rubrics, and what an exercise still measures once an assistant can produce the deliverable.
85 articles
Everything on assessment design
- Does Asking for an Accommodation on an AI Test Count Against You? Rarely in a process run well, and the law bans retaliation regardless. Where the request lands, what the panel is told, and the honest residual risk at a small employer.
- Make the Flexible Version the Default So Nobody Has to Disclose An accommodation request rate near zero measures the invitation, not the pool. Give the commonly requested adjustments to everyone and keep a named route for the rest.
- When AI Is Allowed in the Coding Round, Judgment Is the Test When AI is allowed in a coding interview, finishing stops being the bar. Interviewers watch how you frame, what you test, and what you refuse to ship unchecked.
- Outside Engineering, AI Skill Is Judgment About Your Own Work Outside engineering, the AI bar is recognizing when a confident output is wrong about your own field. Show it with an example, not a tool list.
- An AI Work Sample Tests Whether You Catch the Bad Answer An AI-collaboration exercise is usually built to include a wrong answer somewhere in the material. What earns the score is visible checking, not speed or polish.
- Asked to Redo the Exercise Live Without AI: Why, and What Passes A live, AI-free redo is usually a verification stage run on every finalist, not an accusation. Here is why it got added and what actually passes it.
- Send the Session Where You Corrected the Model Asked to share your AI chat log or screen recording? Send the real session, corrections included, and ask what's captured, who sees it, and how long it's kept.
- Before You Let an Assessment Record Your Screen or Your Face Screen capture and a face template are legally different things, and the second carries its own consent rules in a few states. What to ask before you agree.
- Ask for the Instance First, the Hypothetical Only Where There Is None Ask for the instance first, keep the hypothetical for candidates with no track record in it, and give a technical question a written pass bar or cut the slot.
- Same Model, Different Configuration, Different Impact A vendor audit describes an instrument under the vendor's conditions. The cutoff, the stage, the tuning target and the override path are yours, and each moves impact.
- Extra Time Has No Statutory Multiplier. Judge the Work, Not the Rate No US statute or agency document sets a time multiplier. The obligation is individual, and comparability only breaks if the clock was part of the measurement.
- Borrow the Bar, Not the Verdict, When You Cannot Judge the Work Two practitioners, half an hour each, give you the standard. The verdict stays yours if the exercise asks for a decision you can check without the domain.
- Two Hours Is the Ceiling, and It No Longer Bounds Effort Two hours is the working ceiling for a take-home, but the clock stopped bounding effort. Set the length from what you can grade, and bound the deliverable.
- A Question Bank Row Is Four Fields, and One Person Can Delete Keep the bank, shrink it, and give every row four fields: the question, the evidence a strong answer must contain, one follow-up, and the date it was last reviewed.
- Every Number on the Scale Needs an Observable Behavior Each point on an interview rating scale should name a behavior an interviewer could have watched, with a real quote beside it. Three anchored points beat five vague ones.
- Learn AI Is Not Advice Until It Names a Task You Already Do Every course platform and creator repeats learn AI because the vague version sells. Turn it into a two-week task on work you already do, and it becomes advice.
- The Less Discriminatory Alternative Is Usually a Stage, Not a Model Most of the disparity sits in settings you already control. The search order, the four cheap swaps, and the record worth keeping a year later.
- A Portfolio Piece Counts When You Can Defend Its Decisions AI could have built the artifact, so reviewers now grade the decisions behind it. Keep fewer projects and keep the record of why you made each choice.
- Replace the Five-Point Scale With a Call and Its Reason A rating out of five is optional. The call and the evidence behind it are not. Keep a scale only if a stated rule consumes it, and never average it with a tool's value.
- Ask a Reference for One Episode, Not an Overall Impression Rating questions come back as policy language. Ask a reference to walk through one bounded episode with a date on it, ask the same three of everyone, keep the words.
- Do Not Require a Skill Nobody Can Score: Borrow the Evaluator First A requirement nobody on the team can evaluate is one you cannot defend. Get a practitioner to write the answer key first, or post the task instead of the skill.
- Every Requirement Needs a Stage That Tests It Put the requirement list in two columns: the requirement, and the stage that will test it. Any line with an empty right column gets a stage or comes off.
- A Fixed Rule Beats the Same Manager's Holistic Read On the direct comparison the rule predicts better, but only if the weights are fixed before candidates are seen. Allow overrides, require them in writing, and count them.
- Send the Criteria, Withhold the Weights and the Key Candidates are owed the criteria and the intended time. The weights and the worked answer are what a model optimizes to, so those stay with the reviewer.
- The Take-Home Says No AI: What Happens If You Use It Anyway The take-home says no AI. Detection doesn't work reliably, but the rule itself is often part of what gets scored, and breaking it risks more than one bad grade.
- Diagnose a Thin Pipeline Before Widening It, Then Assess Everyone Nine applicants can mean a posting nobody saw or a skill few people have. Check three comparable postings, then spend the unused screening budget on assessing everyone.
- Vendor Validity Evidence Transfers Only If the Jobs Match Borrowing someone else's validation study has conditions attached, and the job analysis is the one that rarely arrives. What to ask a vendor before signing.
- Use a Real Task You Already Solved, and Keep the Outcome Source the exercise from a real problem your team closed six to twelve months ago. It resists an off-the-shelf answer and arrives with a known outcome.
- Pick One Load-Bearing Claim Before the Call, Then Stay on It Choose the claim that makes the rest of the application irrelevant if it is false, name the detail only a real owner would know, and write both down before dialling.
- What Survives Is Knowing Whether the Output Is Right The standard list of durable human skills is unfalsifiable. Here is the one skill employers are actually testing for now, and how to practice it directly.
- Do ADA Accommodations Apply to an AI-Based Assessment, and What Do You Offer? Yes, the ADA covers an AI-based assessment the same as a paper test. The rule that settles every request: change the conditions, never the skill being measured.
- How Do You Write an AI-Use Rubric Two Reviewers Score the Same Way? Anchor every level in the role's own evidence, rate one dimension at a time, calibrate both reviewers on three real sessions, then measure kappa per dimension.
- Is AI Training Enough, or Do You Have to Check the Work? Completion records attendance, not capability. Transfer depends on how much of a role's AI use is judgment. Run the training, then check the work with a work sample.
- The Algorithm Screen Now Measures Prep Time, Not Programming A model clears the standard format in seconds, so the screen now sorts on prep time. What to keep it for, and the multi-file exercise that replaces the ranking round.
- How Do You Assess a Candidate Who Mostly Supervises AI Agents? Skip the build exercise. Hand over a finished-looking agent run with real defects seeded in it, then score what gets checked, what gets rejected, what gets refused.
- How to Run an AI Skills Assessment on the Team You Have Give everyone the same hour of real occupational work with an assistant available, then read four behaviors off the record instead of collecting self-ratings.
- An Assessment Does Not Need to Live in Your ATS Work out what an assessment integration saves at your volume, then ask what it deposits in the candidate record. A link out and a report back is not the lesser option.
- Per-Candidate Pricing Buys Volume When You Need Depth The pricing model is a product argument in disguise. How to read per-candidate, per-seat and credit plans, and what the price buys in human reading time.
- A Timer Measures Speed, and Speed Is Now the Cheap Part A tight clock mostly separates tooling and quiet rooms. Keep a generous cap so an exercise cannot swallow a weekend, and put the clock on the part an assistant cannot do.
- Should You Build Your Own AI Exercise or Buy an Assessment? Building wins at one role and loses by the third. The cost runs per occupation: a realistic task, seeded errors, rubric anchors and reading hours for each one.
- A Multiple-Choice AI Fluency Test Measures Recall, Not Fluency A quiz measures whether someone can name four competencies. Measuring fluency takes an occupational task, an assistant, a deadline, and a planted error.
- What You May Collect and Keep From a Candidate's AI Session Capture the prompt-and-revision trail rather than the screen, say what happens to it before the candidate starts, and file it with the hiring record.
- Whose AI Account the Candidate Uses Changes What You Can Read Supply the tool when submissions must be comparable or the exercise turns on one model's behavior. Otherwise let candidates bring their own, and name the version.
- Should You Disqualify a Candidate for Using AI on a Take-Home? Disqualify only if the brief said no before they started. Otherwise grade the submission again for what was checked, and price the risk your field can't absorb.
- How Do You Run a Coding or Case Interview With AI Open? Hand the candidate AI output with a real flaw in it and score the audit: what they tested, what they refused to ship, and which option they killed.
- How Do You Tell Which Finalist Did the Thinking? Output converges when both finalists use the same model. Ask each the same three questions: how they framed it, which claim they checked, and what the check returned.
- How Do You Set a Defensible AI Bar for a Specific Role? A defensible bar comes from the occupation's own task list, sits at the minimum the job requires, and is documented before anyone is measured against it.
- Detector, Interview, or Work Sample: Which Defends the Decision? The AI detector loses on validity evidence, adverse impact, candidate experience and cost. Which of the other two defends your decision depends on the occupation.
- Two Graders on the First Ten, One After That Two reviewers on the first ten submissions, marked blind. Compare criterion by criterion, repair the wording where they diverged, then grade single with audits.
- How Do You Evaluate an AI-Skills Assessment Before Buying? Ask what the assessment is benchmarked against, what the candidate sees, whether it outputs a single number, and what evidence sits behind a divergent result.
- How Do You Evaluate a Portfolio When AI Made Half of It? Assume AI made the surface. Ask for the decision record each craft keeps (the cut list, the constraint set, the review trail) and rate it against a guide written first.
- What Should an Early-Career Assessment Measure, and Is Yours an Aptitude Test? Measure the job's actual work behaviors, not general aptitude. The vendor questions a repainted aptitude test can't answer, and the federal rule behind each one.
- Delegation, Description, Discernment, Diligence: What Each One Looks Like at Work Delegation, Description, Discernment and Diligence, shown inside a single quarterly attrition analysis. Only one of the four leaves a trace in the finished file.
- How Do You Grade a Take-Home When Every Submission Is Polished? Stop rating the artifact. Grade three decisions instead: what the candidate assumed, what they checked outside the prompt, and what they rejected. Ask for a short log.
- How Do You Hire Someone Who Can Tell When AI Is Wrong? Stop testing production. Hand over a real task with one wrong thing in the source, then grade four acts: which claim was checked, against what, when, and what changed.
- Six Categories of Hiring Assessment and What Each Can Claim Cognitive, personality, situational judgment, job knowledge, work sample, structured interview: what each licenses you to conclude, and how to choose backwards.
- The Validity Table Everyone Quotes Was Corrected in 2022 The 2022 re-analysis puts structured interviews at .42 and job knowledge tests at .40 close behind. The 1998 numbers on most vendor decks run .10 to .20 high.
- Can an Internal Candidate Really Move Into an AI-Heavy Role? Liking the tools is not evidence. Read three artifacts from a task in the destination role (the frame, the refusal, the check) before approving an internal move.
- Live Coding, Take-Home, or AI-Allowed Work Sample: Which Catches Real Skill? Watched coding measures composure, take-homes lose signal wherever the deliverable is a document, and an AI-allowed sample works only if someone authored the constraint.
- Is Your Assessment Broken After the Latest Model Release? The assessment isn't broken. A model release lowered what the task costs to finish, so a fixed cut score passes more people. Re-anchor the bar per occupation.
- One AI-Skills Assessment for Every Department, or Per-Role? One rubric across every department, one case per role: what has to be shared for results to compare, and what has to be local for the assessment to be job-related.
- Should You Run a Paid Trial Instead of an Assessment? A paid trial wins only when a week of the real work is reachable and the finalist can take the week. The full cost: your hours, theirs, and the employment you created.
- Paying for the Take-Home Buys a Shorter Take-Home You are generally not required to pay for a short exercise whose output you never use. Pay anyway: a flat rate, everyone at that stage, and the assignment gets shorter.
- Personality Explains a Few Percent, and Candidates Can Move It Square the validity vendors quote for conscientiousness: four to nine percent of the variation in performance. Use the profile for interview questions, never a cutoff.
- How Do You Pilot an Assessment Vendor Before Making It a Hiring Gate? Run the vendor's assessment beside your live process, act on none of it, and read the results against how your own hires performed. Fix the cohort and criterion first.
- Is a Prompt Engineering Certificate Worth Anything, or Should You Give a Work Sample? A prompt engineering certificate proves a course was finished. Here is the honest ranking of four credentials, and the 45-minute work sample that beats all of them.
- A Rank Order Is a Weaker Claim Than a Bar Write the criterion from the work, judge each candidate against it with the evidence attached, and compare only inside the group that already cleared it.
- What Should You Look for in a Candidate's AI Chat Log? Skip turn count and phrasing. Read where the first move aimed, which claim got a source demanded, what output was refused, and what was checked outside the chat.
- What Replaces an Online Assessment ChatGPT Solves in Thirty Seconds? Stop testing for the answer. Hand campus candidates occupational material with a defect in it, let them use AI, and grade what they framed, opened and refused.
- How Do You Score an Interview Answer the Candidate Produced With AI? Rate four items (framing, demanded evidence, output rejection, verification) against anchors written in your field's own evidence standard. One rubric, filled in.
- How Do You Screen for AI Judgment in Finance or Marketing? The six behaviors that show AI judgment in a diff are just as visible in a memo, a model or a brief. What each looks like in finance, marketing, consulting and product.
- Measure How Fast They Work With AI, or How Well They Judge It? Measure throughput only where a wrong output is cheap to catch and cheap to undo. Everywhere else, measure judgment and use the clock as a cap, not a score.
- Write the Answer Key Before You Send the Assignment A rubric with no reference answer anchors on whichever submission is read first. Do the assignment yourself, write the key from it, then grade one submission twice.
- Do Take-Home Assignments Still Tell You Anything if Candidates Use AI? A take-home an assistant can finish from the brief alone measures nothing. Put a real ambiguity in the material, then grade the questions asked and the checks run.
- Take-Home Assignment or Live Working Session: Which Shows AI Judgment? A live session shows AI judgment when the real work is fast and observable. When it runs across days, run a capped take-home plus a 20-minute walkthrough of the choices.
- What Does 'AI Fluency' Mean on a Job Description, and How Do You Test It? It means six different things across six functions. Write the requirement from the occupation's own tasks, then test one of those tasks with an assistant open.
- What's the Difference Between a Validated and a Bias-Audited Assessment? A bias audit is arithmetic on selection rates across sex and race groups; validation is evidence tied to one specific job. A clean audit proves nothing about your role.
- What Do You Need From an Assessment Vendor for a Bias Audit or EEOC Inquiry? The validity report, the job analysis, the impact figures, the auditor's letter. What each proves, who at the vendor writes it, and the reply that should stop a purchase.
- What Should an AI Skills Assessment Measure Before You Pay? Buy for job-anchored behavior, findings you can open, occupational data dated this year, and a report the candidate also receives. Tool familiarity predicts nothing.
- AI Fluency Is Four Competencies, Not a Tool List AI fluency is working with AI effectively, efficiently, ethically and safely, in four competencies. Nobody certifies it, so the employer writes the bar.
- What Should a Work Sample Test Now That AI Can Produce It? AI compresses output quality, so the artifact stops separating candidates. Keep the task, grade two acts: the framing before generating, and the check outside the chat.
- The Resume Screen Is Regulated AI. A Human-Scored Work Sample Usually Is Not. Decades of test litigation made assessments look like the risky stage. The 2023 and 2025 rules moved the exposure to the automated screen instead. What actually changes.
- Work Sample or AI-Collaboration Exercise: Which Predicts Performance? Neither wins outright. A classic work sample predicts while day-one work still resembles it; the AI-collaboration exercise predicts once an assistant drafts first.
- Work Sample, Interview, or Assessment: Which Catches AI Dependence? Work sample, interview or AI-skills test? One question decides: can a candidate who can't do the work pass with an assistant open? Only the recorded session holds.
- Write the Assignment Brief So Two Submissions Can Be Compared The prompt is the easy half. Name the time budget, the audience, what sits out of scope, and the tool rule, or you grade four assignments against one rubric.
This beat is written by Bill Nguyen & Olive.