Interviewing

AI-Written Interview Questions Are the Ones Candidates Already Practiced

Use a model to draft interview questions, not to pick the final set, and never to score the answers. Before you keep any generated question, paste it back and ask the model to answer it cold. Whatever it answers well is a question a prepared candidate will also answer well, so it carries almost no signal. Keep the survivors, then write by hand the two questions that turn on your own artifacts, because those are the ones no model can produce from a posting.

The takeDrafting with a model and judging with one arrive in the same product and get treated as one decision, which is the mistake. A bad drafted question costs twenty minutes and is fully reversible. A generated judgment about a person is a selection procedure that nobody in the room can explain afterwards, and it is the part a candidate, a regulator or your own counsel will eventually ask about. Draft freely, sort ruthlessly, and keep the verdict in a human hand where it can be defended.

Where Olive fits

Open a role and see what the work shows

An interview captures a candidate describing how they would check a confident claim; it cannot capture them checking one. Olive puts a role-grounded assignment in front of the person with an assistant that will do all of it if nobody stops it, and a human reviewer writes six findings, each anchored to a moment in the session.

Rank your shortlist

Should you use a model to write the questions at all?

Use it for the draft. Producing thirty candidate questions in two minutes is exactly the shape of task where the measured gains are real, and a recruiter who starts from thirty and cuts to four does better work than one starting from a blank page. What the model cannot do is pick, because picking requires knowing which answers you have already heard forty times and which artifact sits on your own drive.

The evidence for drafting is narrow but solid. In a pre-registered experiment, 444 college-educated professionals did occupation-specific writing tasks; the half given ChatGPT finished 10 minutes faster, 37% quicker than a control group averaging 27 minutes, and graders scored their output 0.45 standard deviations higher, with the largest gains going to the weakest writers 1. Mid-level professional writing, done once, online, for pay. That is a fair description of drafting an interview question and a poor description of deciding whether to hire someone.

The convergence problem is what the guidance leaves out. Your model and the candidate's model, handed the same posting, are drawing on the same public corpus of interview advice, so the questions land in the same place. A generated set is therefore the most anticipated set it is possible to bring into a room. Generating is still the right move here. The output is raw material, and it has to be tested before anyone asks it out loud.

Ask the model to answer your own question list

Run every generated question back through the model with a one-line description of the role and nothing else, and read the answers as a candidate would deliver them. Any question the model handles fluently is a question your best-prepared applicant will also handle fluently, which makes it a test of preparation rather than of the person. Delete those. What is left is a much shorter and much more useful list.

This takes about fifteen minutes for thirty questions and the sorting is unambiguous. Questions like "how do you check AI output for accuracy" and "tell me about a time you used AI to solve a problem" come back with polished, complete, entirely reasonable answers, because the model has read every article that ever answered them. Questions that reference a specific artifact, a specific constraint or a specific failure in your own workflow come back vague, because there is nothing to draw on.

Two refinements worth the extra ten minutes:

  • Ask twice, in different sessions. If two independent runs return substantially the same answer, that answer is the consensus one your candidates have also read.
  • Ask it to answer badly. A question where the model can produce a convincing wrong answer is one where you need to know in advance what separates the two, which is a rubric problem rather than a question problem.

The same sorting is what stops the rewrite cycle that follows a leaked question set. Rotating interview questions turns into a treadmill precisely when the questions being rotated were the anticipated ones to begin with.

Write the two questions a model cannot produce

Write both questions out of your own work, because a posting is all a model has. It has never seen the campaign brief that went out with the wrong pricing, the ticket that got closed twice, or the client email your team still quotes. Those are the raw material for the two questions worth writing by hand, because their answers cannot be rehearsed from public advice and their standard for a good answer already exists inside your team.

Both questions have the same construction. Take a real artifact from the work, remove anything confidential, put one thing in it that the role should catch, and ask what the candidate would do next. The answer is checkable against the artifact rather than against your impression of the person, and it stays checkable when three interviewers read the transcript later. Deriving those questions from the role's own tasks is a short afternoon of work and it is the part of this that does not need a tool.

Structure matters more here than novelty does. In the 2022 re-analysis of the selection literature, structured interviews estimate at .42 against .19 for unstructured ones, from a sample-size-weighted combination of two earlier meta-analyses 2. Structured in that coding means the same questions rated on a common scale, and it is not a product or a script. The gap between two interviews run differently is larger than the gap between two methods, which means how consistently you ask beats how clever the question was. Running a structured interview about AI use so candidates are comparable is the mechanical half of that.

Keep the model out of the scoring

Drafting and judging are different acts with different consequences, and they should not both fall to the tool because they arrived in the same product. A generated question that misses costs one round of interview time. A generated judgment about a candidate is an employment decision with no author, and the first person to ask how it was reached will get an answer nobody in the room can give.

Under US federal selection law the definition is deliberately broad. The Uniform Guidelines, written in 1978, define a selection procedure as any measure, combination of measures, or procedure used as a basis for any employment decision, covering the full range from paper-and-pencil tests through informal or casual interviews and unscored application forms 3. Swapping a scored assessment for a conversation does not move a step outside that definition. Those Guidelines name no software, and they impose a validation burden where adverse impact appears rather than auditing everything, so the reach of a 1978 rule over a generated judgment is an argument from that definition rather than a closed question. It is still the definition your counsel will start from. Where a tool substantially assists or replaces the decision, some jurisdictions attach duties on top of it: New York City's Local Law 144, in effect since January 1, 2023, requires a bias audit within the prior year and notice to the candidate 10 business days before the tool is used 5.

Illinois attaches its duties to a different trigger. Since 2020 its AI Video Interview Act has required notice, an explanation and the applicant's consent from any employer using AI to analyze recorded video interviews for a position in the state, whatever weight that analysis carries in the decision 6. The Act covers that one technology and names no penalty, so read it as evidence about what the trigger class is rather than as a rule with teeth.

Leave the statutes aside and the practical objection still stands. Interviewers make sense of almost anything put in front of them. In a controlled test, 76 undergraduates predicted classmates' semester GPAs; predictions made after an unstructured interview correlated .31 with the actual result, against .65 for prior cumulative GPA alone, and in a later study 96 of 169 participants chose to conduct an interview in which the interviewee answered questions at random over conducting no interview at all 4. Undergraduates predicting a classmate's grades is a lab analogue rather than a hiring result, and the mechanism is what travels: low-diagnostic material dilutes good evidence, and a plausible-sounding summary is exactly that kind of material.

So the working rule is short. Generate, sort, hand-write two, then close the tab before the interview starts. If you want an assist during scoring, a rubric for an answer the candidate produced with AI is a better place to spend the effort than a generated verdict.

See how it works

Common questions

Is it a problem if the candidate also used AI to prepare?

It is expected, and it is only a problem for questions that reward preparation over judgment. Someone who researched your company, anticipated the questions and rehearsed answers has done what candidates have always done, with a faster tool. The questions built on your own artifacts survive that preparation, because knowing the format does not tell anyone what is wrong inside the specific draft you hand over. If your whole list collapses under preparation, the list was the weakness rather than the candidate.

What about vendor tools that generate a question set from the job description?

They have the same convergence problem plus a version you cannot inspect. The question set is derived from the posting text, and the posting text is what the candidate fed their own model, so both sides land in the same neighborhood. Treat the output the way you would treat anything you generated yourself: run each question back through a model, discard what it answers well, and add the artifact-based question the vendor could not have written because it has never seen your work.

Can a model write the follow-up probes as well as the questions?

It can draft them, and follow-ups are where the drafts are weakest. A good probe depends on what the person just said, which is the one input a pre-written list does not have. Generate three probes per question as a safety net for a nervous interviewer, then train the panel on the single move that matters: ask how they know, and then ask how they checked. Those two probes cover most of what a scripted follow-up was going to try to do.

Does using AI to draft questions need to be disclosed to candidates?

Drafting help almost never triggers a disclosure duty. The notice rules in force attach to tools pointed at the candidate, not at the question list you drafted. New York City's duty keys to a tool that substantially assists or replaces discretionary decision-making, and a question you wrote and then chose yourself is not that 5. Illinois keys on the other variable: its AI Video Interview Act, in force since 2020, attaches notice and consent to AI analysis of an applicant's recorded video interview, whatever its role in the decision 6. Other jurisdictions draw the line differently and it moves by year, so check with counsel on the specific tool and the candidate's location rather than on the category.

How many questions should survive the sorting?

From thirty generated, expect four to six to survive, and expect two of the survivors to still need rewriting against a real artifact. That is a normal yield and it is the reason generating is worth doing: producing thirty by hand would take an afternoon, and cutting thirty to five takes twenty minutes. If almost everything survives, the sorting was too gentle. If nothing does, the role description you fed it was too thin to work from.

Should the questions go to candidates in advance?

Sending them in advance is a defensible choice and it changes what you are measuring. Preparation time flattens nervousness and favors people who cannot rehearse on demand, which is often a fairness gain, and it kills any question whose value came from surprise. Artifact-based questions survive being sent early, since the candidate cannot see the draft until the room. Decide once, apply it to every candidate for that role, and write down which way you went.

References

  1. 1. Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence (working paper, not peer reviewed) MIT Department of Economics, 2023. economics.mit.edu Supports the claim that short self-contained writing tasks, which is what drafting a question set is, are where the measured AI gains actually sit.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the claim that consistency of administration matters more than question novelty: .42 for structured interviews against .19 for unstructured.
  3. 3. 29 CFR Part 1607 - Uniform Guidelines on Employee Selection Procedures (1978), sections 1607.16(Q) and 1607.3(A) U.S. Government Publishing Office, Code of Federal Regulations (Title 29, Vol. 4, 2023 edition), 1978. govinfo.gov Supports the claim that a generated judgment about a candidate is a selection procedure under the Uniform Guidelines, as an informal interview also is.
  4. 4. Belief in the unstructured interview: The persistence of an illusion Judgment and Decision Making, 8(5), 512-520 (Society for Judgment and Decision Making), 2013. sjdm.org Supports the claim that low-diagnostic material dilutes good evidence and that interviewers construct meaning from answers even when the answers are random.
  5. 5. Automated Employment Decision Tools: Frequently Asked Questions NYC Department of Consumer and Worker Protection (DCWP), 2023. nyc.gov Supports the claim that a disclosure duty attaches to tools that substantially assist or replace a hiring decision, with NYC Local Law 144 as the named jurisdiction, date and notice period.
  6. 6. Artificial Intelligence Video Interview Act, 820 ILCS 42 Illinois General Assembly, Illinois Compiled Statutes, 2020. ilga.gov Supports the claim that a notice and consent duty can attach to AI analysis of a candidate's recorded video interview whatever weight that analysis carries in the decision.

6 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.