Screening
Rebuilding the Application Form Around What a Model Can't Fill In
Keep the fields whose answers depend on the applicant's own week, and drop the ones an assistant can write from the job description. That is the test, and it has nothing to do with length. A motivation paragraph, a self-rated skill slider and a restated work history are all derivable from the posting. What a candidate would want to work on first, one decision they made recently, and a link they can talk through are not.
The takeThe correction most teams reach for is more questions, and it is the wrong one. Extra questions are the cheapest thing in the process to produce, so a longer form buys a bigger pile of uniform answers and one more reason for the person typing by hand to quit halfway through. The form was never the assessment. It is an intake sheet, and the honest version of it is short, specific, and read by someone who will act on the answers.
Where Olive fits
Open a role and see what the work shows
An application form collects claims, and no field on it can show a person working. Olive is the employer-bought version of the other half: a 40-to-60-minute assignment in the candidate's own occupation, done with an AI assistant, returned as six findings that each carry the timestamped moment behind them.
Rank your shortlistWhich Fields Survive the Derivability Test?
The ones whose good answer needs something only that person knows: a decision they made, a project they can talk through, the first thing they would pick up here. Take each field and ask whether a capable assistant could produce a strong answer from the job description alone. If it could, the field now measures the tool. If it could not, the field still returns a person.
Run the current form through it once, field by field. The usual result:
- Name, contact, location, work authorization. Keep. Nothing to derive, and something to act on.
- Employment history, titles, dates, education. Keep, and read them as claims to check later.
- Why do you want to work here. Drop. The posting is the source material and the answer is written from it.
- Self-rated skill sliders. Drop. They ask for a number the applicant has no way to calibrate and you have no way to read.
- The cover letter box. Drop or replace. What to ask for instead of a cover letter is its own decision, and it is the field most teams are actively arguing about.
- One question about the work itself. Add, if nothing like it is there yet.
The test cuts across the usual short-versus-long argument. A three-field form full of derivable questions tells you less than a six-field form where three of the answers could only have come from that person on that day. Derivability is the variable. Length is the thing teams change because it is easy to change.
Why Adding More Questions Backfires
Because generated answers scale and reading time does not. Every question added to the form is free for an applicant running an assistant and expensive for the one typing. Nearly 40% of US adults aged 18 to 64 reported using generative AI in late 2024, and 23% of employed respondents had used it for work at least once in the previous week 1. Assume the tool is in the room.
That is self-reported use in one country at one moment, and it ages in months, so treat it as a floor for the direction rather than a rate to plan against. The mechanism does not depend on the number. One group answers everything, at length, in the register the posting used. The other group answers three questions well and abandons the fourth. The extra field removed exactly the people it was added to find.
Setting a trap produces the same inversion with a worse failure mode. A hidden instruction buried in the posting catches the applicant who pasted the description without reading it, which is carelessness rather than dishonesty, and it silently punishes anyone using a screen reader: hidden instructions in a job posting are worth understanding before anyone suggests one in a meeting.
The real cap is the reading budget. Every field kept has to be read by somebody who will do something about it, and the number of fields is set by that person's week, not by what the applicant tracking system offers to collect.
Keep Three Fields, and Write Them Like This
Three questions resist derivation because they are anchored to the applicant's own week rather than to the posting. What would you want to work on in the first month here, and why that one. Describe a decision you made in the last year and what you would do differently. Send one link to something you can talk through for ten minutes. Cap each answer at 120 words.
Each resists for a specific reason, and the wording is what does the work:
- The first-month question forces a choice among the things in the posting. A generated answer restates the responsibilities. A person's answer names one and gives up the others.
- The decision question needs a real week behind it. The failure mode is the polished non-answer, which is easy to spot because it has no consequence in it.
- The link moves the evidence out of the form. It also sets up the conversation, since the candidate has already chosen what they want to be asked about.
Say plainly that AI is allowed on all three. Catching anyone is not the point: a candidate who used an assistant to tidy an answer still had to supply the decision, the project and the reason, and those are the parts being read. Researchers building an AI literacy assessment for a US Navy robotics programme found that a scenario task simulating real use outperformed the tests they had adopted from prior research or written themselves, and argued that prevailing assessments favour foundational technical knowledge over practical knowledge, such as interpreting model outputs or selecting tools 2. That is one programme with no published effect size, so read it as a design argument rather than a benchmark. The closer a question sits to the actual work, the harder it is to answer from outside the work.
What Do You Do With the Answers?
Route them in one pass. The checkable claims go to verification later. The three anchored answers go to whoever runs the first conversation, as the questions they open with. Nothing in the form decides anything by itself: it is one input, and a process combining two or three different kinds of evidence outperforms one hunting for a single perfect question 3.
A workable pass looks like this. Read the three answers first and the history second, because the history is a set of claims and claims keep. Write one line per applicant naming the question you would ask them. If you cannot write that line, the answers were derivable and the field needs rewording. That is not a reason to reject the applicant.
Then audit the form the way you would audit a report nobody reads. Once a quarter, ask which field last changed a decision. Expect two or three that did, and a long tail nobody has read since the form was built. Deleting the tail pays for the question you wanted to add.
Two parts of the form need a review of their own. The questions that reject people with no human involved are a different mechanism with a different owner: knockout questions are the only auto-reject you configured yourself. And the document attached to the form has quietly changed meaning, which is worth settling before rewriting anything around it, because a resume now works as a claims index rather than a writing sample.
Common questions
Should the application still ask for a cover letter?
Not in its open-ended form. A box asking why the candidate wants the role produces the most uniform writing in the entire application, and it did before generated text existed. If you want what the cover letter was supposed to give you, replace it with one specific question, a word limit, and a commitment that someone reads every answer. If nobody will read them, delete the field instead of collecting text nobody opens.
How many fields is too many?
The cap is a reading budget, not a number. Count how many answers one person can genuinely read and act on in a week, divide by expected applicants, and you have your field count. Someone with an hour a week and a hundred applicants is not reading five paragraphs each, and the arithmetic says so before anyone argues about the number. A field nobody reads is worse than a missing field, because it costs the applicant time and creates a record you never looked at.
Do these questions have to be required, or can they be optional?
Make them required and short rather than optional and long. An optional 500-word box is answered by whoever has an assistant open, which reintroduces the bias you were trying to remove. Three required 120-word answers cost every applicant the same and give you a comparable set. Publish the word limits on the form so nobody guesses, and offer an alternative route for anyone who needs one.
Does asking for a link disadvantage people with no portfolio?
It does when the link has to be a portfolio. Ask for something the candidate can talk through for ten minutes and accept anything that qualifies: a repository, a document, a spreadsheet, a process they built at a previous employer that has no public artifact at all. Say so in the field label. The purpose is a shared object for the conversation, not proof that someone had time to build a personal website.
Is it worth asking which AI tools the candidate uses?
A tool list ages faster than the form does and tells you almost nothing about judgment. Two people naming the same three tools can work in completely different ways, and the answer is derivable from any job posting that names its stack. If the question matters for the role, ask what the candidate used AI for on a specific piece of work and what they changed afterwards. That answer is about them.
References
- 1. The Rapid Adoption of Generative AI (NBER Working Paper 32966) nber.org Supports the claim that an application form should assume the applicant has an assistant: nearly 40% of US adults aged 18 to 64 reported using generative AI in late 2024, and 23% of employed respondents had used it for work at least once in the prior week.
- 2. AI Literacy Assessment Revisited: A Task-Oriented Approach Aligned with Real-world Occupations arxiv.org Supports the design argument that a question anchored to real work reads better than an abstract self-assessment, stated as the authors' own comparison rather than as a measured margin.
- 3. Revisiting the design of selection systems in light of new findings regarding the validity of widely used predictors cambridge.org Supports the claim that the form is one input among several: combining two or three kinds of evidence beats searching for one perfect question.
3 sources, numbered by first appearance. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.