Pipeline
Drafting a Job Description With AI Without Inventing Requirements
A model can draft the job description, with one section fenced off. Give it the task list and the constraints instead of the job title, because a title makes it average public postings and the qualifications section comes back full of requirements with no origin. Then delete every requirement you cannot trace to a specific task, and for each survivor name what in your process would show whether a candidate has it.
The takeA team that drafts its posting with a model and then screens applicants for sounding like one is holding a position it cannot say out loud and will be asked about. The honest resolution is to allow AI on both sides. Put it in the posting: AI is allowed, and the round checks judgment rather than polish. That sentence costs nothing, it is the only version of the rule anybody can enforce, and writing it forces you to decide what the loop actually checks, which was coming anyway.
Where Olive fits
Open a role and see what the work shows
Olive's twelve item banks are each grounded in one occupation and its SOC code, so an assignment is tied to an occupation rather than to a job title. A person writes all six findings, and every released report exports with its rubric, scorer and bank versions attached.
Rank your shortlistCan you draft a job description with AI?
Yes. Drafting a posting is one of the better uses available. Short, self-contained professional writing is where the measured gains are clearest, and a job description is exactly that. What the model cannot do is know your job, so the drafting help is real everywhere except the one section that decides who gets screened out.
The evidence on this kind of writing is unusually direct. In a pre-registered online experiment, 444 college-educated professionals did occupation-specific writing tasks, and the half given ChatGPT finished 10 minutes faster, 37%, against a control group averaging 27 minutes, with graders scoring their output 0.45 standard deviations higher and the largest gains going to the weakest writers 1. Those were one-sitting tasks with no revision cycle and no colleagues, which describes a first draft of a posting well and describes almost nothing else in hiring.
So split the document before you prompt anything. Three parts of a job description are drafting problems:
- The summary of the role. The model tightens what you feed it.
- The context paragraph. Team, stack, who the person works with, what the day looks like.
- The benefits, location and process boilerplate. Genuinely repetitive, genuinely fine to generate.
One part is not a drafting problem at all: qualifications and requirements. That section is a set of claims about who can do this job, and every line in it is a bar somebody may be rejected against. Generating it is the difference between using a model to write faster and using one to decide who gets considered.
Give it the task list, never the job title
A title is an instruction to average. Asked for requirements for a "senior marketing manager," the model returns the central tendency of every public posting under that title, which is a real document made of other companies' compromises and none of your work. A task list gives it something to compress instead of something to imitate.
What to put in the prompt, in this order:
1. Fifteen to twenty task statements for the role as it is performed now, marked for which ones an assistant already drafts. 2. The constraints: level, range, location rules, who this person reports to, what they own end to end. 3. Three things that go wrong when the role is done badly. This is the input that produces a specific description rather than a competent one. 4. An instruction to omit qualifications entirely, and to list assumptions it would otherwise have made.
That last instruction is the one worth keeping. The assumptions list is genuinely useful output: it shows you which requirements the averaging would have inserted, which is a decent map of what your competitors ask for and a terrible map of what your job needs.
The prompt cannot supply a task list you never made. Finding out what AI actually does in this role takes about half a day and it is the input every later step reads. Without it, a model-drafted description is a plausible average, and a plausible average is what a thousand applicants feed straight back into their own model, since the job description becomes the candidate's prompt the moment it is published.
Which requirements survive the trace-back?
The ones that point at a task in your list and at evidence in your loop. Take each requirement the draft produced, name the task it comes from, then name what in the process would show whether a candidate has it. Two answers and the line stays. One answer and it moves to preferences. No answers and it goes.
The reason to be strict is not tidiness. A hiring test has to measure the person for the job rather than the person in the abstract, and the touchstone the US Supreme Court set for a practice that screens people out is business necessity, with nothing forbidding testing as such, only giving a device controlling force unless it is demonstrably a reasonable measure of job performance 2. That is 1971 language about aptitude tests at a power plant, quoted here for the principle rather than as advice, and the principle is the one a generated requirement fails: nobody can say where it came from.
The trace-back also catches the subtler failure, which is a requirement that is true, traceable, and still not checkable. "Strong analytical judgment" traces to real tasks. Nothing in a four-stage loop distinguishes a candidate who has it from one who describes it well, so it functions as a line that admits everybody and justifies rejecting anybody.
A draft that reads well earns no extra trust. A field experiment with 758 Boston Consulting Group consultants included one task deliberately chosen to sit outside AI capability, and on that task the consultants using GPT-4 were 19 percentage points less likely to reach the correct answer, 84.5% of the control group against 60% and 70% in the two AI conditions 3. They could not tell which side of the line their task sat on, so the failure arrived as a confident wrong answer on work that looked like work the model handles well. A qualifications section nobody traced has that shape. For the wording itself, an AI-skills requirement that is not legally vague is shorter than most drafts produce.
Say what applicants may use, since you used it
Put one line in the posting stating that applicants may use AI and that the process checks judgment rather than polish. A team drafting its posting with a model while penalizing candidates for sounding like one is holding a position it will be asked about, and the version that survives the question is the permissive one, because the alternative depends on telling generated writing from careful writing, which nobody can do reliably.
The practical shape of the line is short: AI is allowed at every stage, the work is checked for judgment, and anything a candidate submits they should be ready to talk through. That gives you a defensible round and it gives strong applicants a reason to believe the process is not a guessing game. It also heads off the thing that actually damages a pipeline, which is managers rejecting candidates for sounding like AI on no evidence at all.
Your sense of how much the model helped is not a measurement. In a randomized trial, 16 experienced open-source developers forecast that AI tools would cut their completion time by 24% and estimated afterwards that it had cut it by 20%, while the measurement showed the tools made them 19% slower 4. Sixteen developers on codebases they knew intimately is a narrow setting and the magnitude does not transfer. The gap between belief and measurement does.
So measure the description the way you would measure anything else: applications per week, the share that clear the first real check, and whether the hiring manager recognizes their own job in the document. A generated draft that shortens the writing by an hour and adds two unowned requirements has cost more than it saved, and the cost lands on people who never see the document that rejected them.
Common questions
Is a model-drafted job description legally risky by itself?
The drafting is not the exposure; the unowned requirement is. A description written by a person can carry a requirement nobody can justify, and a description drafted by a model can be perfectly traceable. What changes with generation is volume and speed: it is easy to publish twenty descriptions carrying inherited qualifications nobody examined. Keep a short record of where each requirement came from, dated, attached to the file. That record is cheap to make while writing and expensive to reconstruct later, and what counts as defensible in your jurisdiction is a question for your counsel.
What should the prompt actually contain?
Task statements, constraints, three failure modes, and an instruction to omit qualifications and list its assumptions instead. Feeding it a job title invites averaging, and averaging is where inherited requirements come from. Length helps: a prompt carrying twenty specific task statements produces a specific description, while a two-line prompt produces something you have read before. Keep the prompt itself with the description, since it documents what the draft was built from and makes the next refresh a twenty-minute job.
Should the posting disclose that AI helped write it?
Not necessary, and not the interesting disclosure. Candidates care far more about what happens to their application than about how the posting was produced. The line worth writing states whether AI is allowed on their side and what the process examines. If a team is uncomfortable saying the posting was model-assisted, that discomfort is usually a signal about the rule they were planning to apply to applicants, and it is worth resolving before publishing either sentence.
Can a model check a description for bias?
It can flag obvious wording patterns, which is useful and limited. What it cannot do is tell you whether a requirement screens people out unnecessarily, because that depends on your applicant pool and on what the job needs, neither of which is in the prompt. Use it as a proofreading pass over gendered or exclusionary phrasing, then run the substantive check by hand: for each requirement, what task it comes from and what evidence the process gathers about it.
How do I tell whether the generated description is any good?
Send it to two people who do the job and one who does not. Practitioners catch the invented requirement and the duty that has not existed for a year. The outsider catches whichever sentence only makes sense to somebody already on the team. Then watch the first fifty applications: if most of them look plausible against the requirements and few clear the first real check, the description is describing something the loop does not examine, which is a document problem rather than a sourcing problem.
References
- 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) ✓ economics.mit.edu Supports the claim that short professional writing is where AI drafting help is best evidenced: 444 professionals, 10 minutes and 37% faster against a 27-minute control, grades up 0.45 standard deviations, largest gains to the weakest writers.
- 2. Griggs v. Duke Power Co., 401 U.S. 424 (1971) ✓ law.cornell.edu Supports the claim that a screening device must be tied to the job: the touchstone is business necessity, and a test must measure the person for the job rather than the person in the abstract.
- 3. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality (Working Paper 24-013) ✓ mitsloan.mit.edu Supports the claim that a wrong answer arrives looking like a right one on work the model appears suited to: in a field experiment with 758 consultants, on one task outside AI capability those using GPT-4 were 19 percentage points less likely to be correct, 84.5% against 60% and 70%.
- 4. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (arXiv:2507.09089) ✓ arxiv.org Supports the claim that a felt speedup is not a measurement: developers forecast a 24% time reduction and estimated 20% afterwards, while the trial measured them 19% slower.
4 sources, numbered by first appearance. Every one was opened and checked against the claim it carries. How Olive sources claims
General guidance for hiring teams. What works at one company and one volume may not transfer to yours.
Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.