Screening
What Should You Ask For on the Application Instead of a Cover Letter?
Replace the cover letter with one scored question about a decision the candidate made in their own work: what they chose, what they gave up, what it cost and how it turned out. Cap the answer at 150 words, ask for one checkable number, date or name, and write the rubric before the question goes live. Score everyone on the same three points, credit school and volunteer work from early-career applicants, and never mark an answer down for reading like AI. It sorts the pile rather than deciding the hire.
The takeWhat the letter really measured was effort, and effort is now free. That is the loss worth naming, because nothing cheap replaces a signal that used to cost a candidate an evening. All you get to choose now is which kind of specificity you ask for, and the one that survives volume is the kind someone outside the room could contradict: a number, a date, a name. The risk, and nobody has measured it, is that reviewers start reading detail as honesty. Detail is not honesty. It is only the part of an account that can be checked, and that is the whole of what it is worth.
Where Olive fits
Open a role and see what the work shows
A good application question shows you a decision the candidate remembers making; it cannot show you one being made. Olive is an employer-purchased assignment of 40 to 60 minutes done with an AI assistant, returned as six findings a human writes against timestamped moments in the session, and the candidate is granted the same report.
Rank your shortlistWhy the Cover Letter Stopped Screening Anything
Because it asks for the one thing a language model produces best: fluent, role-shaped prose with nothing checkable in it. The letter was always a weak instrument: you were reading motivation and writing ability, and usually getting neither. What changed is the floor. Every applicant now clears the bar the letter was set at, so it stopped sorting anyone.
Detection does not rescue it. A test of fourteen detection tools found them neither accurate nor reliable, biased toward calling generated text human-written, and worse again once the text is lightly edited 1. Running letters through a classifier converts a weak signal into a confident wrong one, which is the argument in Do AI Detectors Work on Resumes and Cover Letters?.
The more useful reframe is that the letter was never the evidence. It was a proxy for effort, and effort got cheap. The same thing happened one field over on the resume, and What to Screen On When Every Resume Looks Perfect works through the top-of-funnel version.
Whatever replaces it should not be the same instrument in new clothes. Rating schedules that ask applicants to narrate their background generally relate poorly to job performance, and the weakest items of all are length and recency of education, academic achievement and extracurricular activities 4. "Tell us why you want this role" is that instrument. So is a 300-word essay on your passion for the mission.
Ask for One Decision, Not One Letter
Ask for a decision the candidate made and could have made differently. Name the decision, the option given up, what it cost, and how it turned out, with one number, date or name in the answer. Cap it at 150 words. That shape already has a name in federal hiring practice. It is the accomplishment record, built on the principle that past behavior predicts future behavior 2.
The generic version, before you aim it at an occupation:
> In the last two years, name one decision you made in this work where you had to give something up. What were the options, what did you pick, what did it cost, and how did it turn out? 150 words. Include one number, date or name.
Four details do the work.
- "Give something up" is the load-bearing clause. It rules out the achievement anecdote, which is the format every applicant already has drafted. A decision with no alternative is a task.
- The 150-word cap is for you, not for them. Accomplishment records are slow to score, and scoring is the method's main cost 2. At 150 words a reviewer reads a hundred applications in an afternoon.
- Say the answer may be verified. Applicant inflation is the known failure mode of any self-report, and the documented countermeasures are creating the expectation that responses will be checked and then actually checking some 4. One sentence under the field does most of that work.
- Ask for a name. Accomplishment records normally carry the contact details of someone who can confirm the account 2. On an application form, "who else was in the room" is enough, and you only call it on finalists.
And say out loud that AI help is fine. It is not the thing you are screening on, and a rule you cannot enforce teaches candidates that your process runs on guesses. What a model cannot supply is the memory: the number that came in wrong, the person who pushed back, the week it took to undo. Should You Still Ask for a Cover Letter? covers the case for keeping the letter alongside this; the case against is mostly that a second open text box gets filled by the same tool that filled the first.
Write the Question for the Occupation: Six Versions
Generic prompts get generic answers, and a model fills a generic prompt perfectly. Aim each version at an artifact only someone inside that job has handled: a change that shipped wrong, a client who pushed back, a forecast that missed. Keep the four parts constant across roles so answers stay comparable: the decision, the alternative, the cost, the outcome.
Software engineer. "Describe something you shipped knowing it was wrong in some way. What was wrong, why did it go out anyway, and what did it cost later?" Look for a named constraint (a deadline, a dependency, a migration deferred) and a consequence with a date on it. Vague answers reach for "technical debt" and stop.
Designer. "Describe work a stakeholder rejected. What was the objection, which part of it was right, and what changed in the next version?" Look for the concession. Someone who has shipped design work can say which half of the pushback they agreed with; someone describing the job from outside makes the stakeholder simply wrong.
Marketer. "Name a campaign whose numbers came in under forecast. What did you forecast, what came in, and what did you find when you went looking?" Look for two numbers and a cause the candidate went and checked. "Market conditions" is the tell that nobody opened the report.
Salesperson. "Describe a deal you lost after a verbal yes. When did you know, what changed, and what do you ask earlier now?" Look for the moment and the changed question. Real answers name a person and a stage; borrowed ones name a competitor and a price.
Operations or supply chain. "Name a supplier you moved away from. What number triggered the move, what did the switch cost, and what broke during it?" Look for landed cost, lead time or defect rate, and for something that broke: a first shipment late, a spec drifting, a contract term nobody had read.
Data analyst. "Describe an analysis whose answer changed after you checked it. What was the first answer, what did you check, and what was wrong?" Look for the check itself: a recomputation, a source opened, a join that was duplicating rows. This is the one version that transfers almost unchanged to any analytical role.
For a role not listed, write the seventh yourself by naming the field's own failure artifact, the thing that goes wrong often enough that every practitioner has one and no outsider has any. A paralegal has a document produced late. A support lead has an escalation that should have been refunded on day one. A recruiter has an offer that fell through in week two. How to Screen for AI Judgment in Non-Technical Roles works through the non-engineering cases in more detail.
How to Score It Without Grading Prose
Score three things only, apply the same three to everyone, and fix the rubric before the question is posted. Is the decision specific? Is the tradeoff real, meaning the alternative was genuinely available? Is the outcome checkable, meaning a number, a date or a name appears? Do not grade writing quality, and do not grade whether an answer sounds like AI.
That last rule is not a courtesy. A scored application question is a selection procedure, and the Uniform Guidelines define one broadly enough to cover "informal or casual interviews and unscored application forms" 3. Once a procedure has disparate impact, it has to be job related and consistent with business necessity. The EEOC's phrasing is that it should be "associated with the skills needed to perform the job successfully" and "necessary to the safe and efficient performance of the job" 5. A rubric written after the answers arrive is very hard to describe that way.
The three points, as a scale a reviewer can hold in their head:
- 2: specific. Names the alternative, the cost and the outcome, and at least one of them is checkable. A follow-up question has somewhere to land.
- 1: partial. A real situation with the tradeoff sanded off, or an outcome with no number, date or person attached.
- 0: generic. Could be about any job at any company. This is the score for polished, confident, empty answers, and it is where most of them fall.
Two fairness notes worth writing into the rubric itself. Accomplishment records generally show little or no difference in performance between men and women or between applicants of different races, though that depends on the competencies being assessed 2, which is an argument for keeping the assessed competency narrow and job-related rather than assuming the format protects you. And for early-career applicants, credit accomplishments from school, volunteer work, community service or military service, which is the standard practice for this method and not an accommodation 2.
The failure mode to guard hardest against is a reviewer marking down a candidate whose second language is English for sounding synthetic. How to Stop Managers Rejecting Candidates for 'Sounding Like AI' makes the full argument; the short version is that a hunch applied unevenly across a pool is exactly what turns a screen into a legal problem, and a detector result will not rescue it 1.
What This Still Won't Tell You
Whether the decision happened. A 150-word answer is a self-report, and self-reports inflate; the documented remedy is the expectation of verification rather than better wording 4. So run the question as a sort rather than as a gate. It tells you who is worth fifteen minutes on the phone. It does not tell you who can do the work.
Two cheap follow-ups close most of the gap, and both belong on a scheduled call rather than in the form:
- "What would you have needed to know a month earlier?" Lived decisions have a counterfactual attached and it arrives fast. Invented ones produce a general lesson.
- "Who disagreed with you at the time?" Real ones have a name and an argument. Borrowed ones have consensus.
The honest cost of this format is also worth stating. Some applicants who dislike writing detailed narratives are put off by it and do not apply 2, so a long or vaguely worded prompt shrinks your pool without improving it. One question, one cap, one visible rubric line.
If the decision matters enough that a description is not sufficient, the next step is work rather than more questions: a short occupational task where you watch the choices get made instead of recounted. Should You Drop the Resume Screen for a Work Sample? does the arithmetic on when that is cheaper than screening, and Are Virtual Job Simulations on a Resume Evidence of Anything? covers what the pre-built versions do and do not prove. See how Olive measures this for the same distinction applied to one behavior: a claim tested against something outside the conversation, as an act in the record rather than a described one.
Common questions
Should you drop the cover letter entirely or run both?
Drop it. Two open text boxes on one application get answered by the same tool in the same sitting, and the letter is the one with no rubric behind it. Keeping both also raises the time cost for every applicant while adding nothing you score. If your industry expects a letter, make it optional and unscored, and be explicit that the decision question is the part that gets read.
What if the candidate wrote the answer with AI?
That is fine, and you cannot reliably tell anyway. The question is aimed at material a model does not have: which alternative was on the table, what the number came in at, who objected. A model can tidy that account. It cannot supply it. If an answer is polished and still names nothing specific, that scores zero on the rubric for being generic, not for being AI-assisted, and those two reasons must stay separate in your notes.
Does this work for entry-level applicants with no work history?
Yes, if you say so in the prompt. The standard practice for accomplishment-record questions is to credit examples from school, volunteer work, community service, military service or personal projects. A student who dropped a research approach two weeks in and can say what made them drop it has answered the question completely. Without that sentence in the prompt, you filter for people who already had a job, which is not the competency you meant to assess.
How many words should the answer be?
Around 150, and enforce it. The cap protects your reviewers more than your applicants: narrative answers are slow to score, and scoring is where the cost of this format sits. It also equalizes the field, since length is the one dimension an AI-assisted applicant can inflate without effort. State the cap in the field label and tell candidates that anything past it is not read.
Can you ask for someone who can confirm the story?
Yes, and it changes behavior before anyone is called. Ask for a name and a role ("who else was in the room") rather than full reference details at the application stage. The documented effect of a verification expectation is less inflation in the answers themselves, so most of the value arrives whether or not you make the call. Verify on finalists only, and say that is your practice.
How do you keep this from becoming another unstructured screen?
Settle the scoring before the question goes live, use the same three points for every applicant, and record a score rather than an impression. Have two people score the first twenty answers independently and compare, which surfaces a vague rubric line in about an hour. If reviewers disagree constantly on what counts as a real tradeoff, the fix is a sharper prompt, not a longer discussion.
References
- 1. Testing of Detection Tools for AI-Generated Text arxiv.org Fourteen detection tools judged neither accurate nor reliable, biased toward classifying generated text as human-written, and degraded further by light editing.
- 2. Assessment and Selection: Accomplishment Records opm.gov The method, its behavioral-consistency basis, the verification contact, high content and criterion-related validity, generally little or no subgroup differences, slow scoring, crediting non-paid experience, and applicants deterred by long narratives.
- 3. 29 CFR 1607.16 - Definitions (Uniform Guidelines on Employee Selection Procedures) govinfo.gov A selection procedure is any measure used as a basis for an employment decision, including informal or casual interviews and unscored application forms.
- 4. Assessment and Selection: Training and Experience Evaluations opm.gov Rating schedules relate poorly to job performance, with length and recency of education, academic achievement and extracurricular activities weakest; applicant inflation is countered by expectation of verification and by verifying.
- 5. Employment Tests and Selection Procedures eeoc.gov A selection procedure with disparate impact must be job related and consistent with business necessity, associated with the skills needed to perform the job successfully.
5 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.