Interviewing

Structure Is Four Things, and Same Questions Is Only One

A structured interview is four things together: content drawn from an analysis of the job, the same questions asked of every candidate in the same order, an anchored rating scale applied to each answer, and independent scoring recorded before anyone discusses the candidate. Semi-structured is not a middle setting on a dial. It names which of the four got dropped, and the one teams drop most is the last, which is also the cheapest to keep.

The takeThe validity figures everyone quotes were measured on a coding of two things: the same questions, and every answer rated on a common scale. A loop that asks fixed questions and then settles it in a discussion has only the first of those, so it is quoting a number measured on something else. A claim to run structured interviews has become a statement about paperwork rather than about how the decision gets made, and the paperwork carries none of the evidence.

Where Olive fits

Open a role and see what the work shows

Structure is what makes two candidates comparable, and the same logic applies to whatever an interview cannot reach. Olive runs one role-grounded assignment for every candidate on a role, with an AI assistant that will do the whole task if nobody stops it, and a human reviewer writes the six findings.

Rank your shortlist

What the four parts are

Job-analysed content, fixed questions in a fixed order, anchored scales, and independent scoring before discussion. Two of those four have a public definition behind them. A US federal practical guide defines a structured interview by three properties that land on the middle two: the same questions in the same order, a common rating scale, and interviewers agreeing in advance on what an acceptable answer looks like 1. The first and last are what a hiring team adds.

Taken one at a time:

  • Job-analysed content. Questions come from what the role actually does, not from a list of good interview questions. This is the part that makes the rest defensible, because it is what ties the thing being measured to the position being filled.
  • Fixed questions, fixed order. Every candidate faces the same set, with follow-ups allowed where they were planned in advance. Invent a new topic mid-interview and that candidate has answered a question nobody else was asked, which produces nothing comparable.
  • Anchored scales. Each answer is rated against written descriptions of what a weak, an adequate and a strong answer contains. Anchors are what stop two interviewers using the same number to mean different things.
  • Independent scoring first. Ratings are recorded before the panel talks. Cheapest of the four, and first to go.

That guide dates from 2008 and was written for US federal merit-system hiring. It is guidance rather than statute, and it binds nobody in the private sector 1. It is useful anyway, because it is a definition written by a hiring authority with nothing to sell, and because it predates every current argument about interviewing.

Does semi-structured count?

Not as a level of rigour. Semi-structured describes which parts were dropped, and the honest version of the term names them: fixed questions without anchored scales, or anchored scales without independent scoring. Each omission has a specific cost, and stating it turns a vague label into a decision somebody actually made.

The corrected meta-analytic estimates put structured interviews at .42 and unstructured at .19, from a sample-size-weighted combination of two earlier meta-analyses; before any correction the raw observed figures were .32 and .13 2. That gap is why the method gets recommended, and the label behind it is narrower than the four parts: one of the two meta-analyses graded interviews on four levels of structure and the .42 collapses its top two, so the coding turns on standardised questions and answers rated on a common scale 2. Job analysis and independent scoring sit outside what that number measured, which makes them the two parts a team has to argue for on its own.

The same finding carries an implication teams tend to skip past: the difference between two interviews is larger than the difference between two methods. Adding structure to the interviews you already run does more than swapping interviews for a different instrument. A team with no budget for a new instrument can still capture most of what is available.

Where the interview is about how a candidate works with AI, the rule holds and gets harder to keep, because the topic invites exactly the improvised follow-ups that leave two candidates with two different interviews (run a structured interview about AI use).

Why independent scoring is the part to fix first

Of the four, only one protects the readings after they exist. Job analysis, fixed questions and anchored scales all shape what gets collected. Independent scoring is what keeps four separate observations from collapsing into one. Drop it and the loop still produces four scorecards, filled in during or after a discussion that has already reached its conclusion.

Self-reported adoption gives a rough sense of how common the full method is. In SHRM's benchmarking of its member organisations, with data collected through 2021, 34% of organisations used structured interviews to assess executive candidates, 37% for middle management and 36% for individual contributors, against in-person interviews at 79%, 78% and 76% 3. Those are organisations describing themselves against SHRM's own definition, and organisations routinely claim structure they do not run, so read them as an upper bound rather than a measurement.

If all four are missing, work backwards through the list. Start with independent scoring, because it costs a deadline and an email. Then anchors, because they are what make two ratings comparable at all. Then fixed questions. Then the job analysis, which is the most work and the thing that makes the whole set defensible, so give it the time it needs.

Independent scoring only holds if the meeting afterwards respects it, and the meeting is where most of the gain is handed back (collect the scores before the debrief).

Check your loop against the four

Take the last role you filled and answer four questions about it. Where did the questions come from. Did every candidate get the same ones. Was there a written description of what a strong answer contains. Were ratings recorded before the debrief. Anything you cannot answer from a document is a part you do not have, whatever the process is called internally.

Two things this check usually surfaces. Interviewers improvising follow-ups that then become the basis for the rating, so the comparison between two candidates is a comparison between two different interviews. And a rating scale with numbers but no anchors, which looks like a common scale until two raters compare notes, because a 4 means whatever each rater's private standard says it means.

Structured interviews come out top-ranked in that 2022 re-analysis at .42, paired there with a Black-White subgroup difference of .23, against .79 for cognitive ability tests 2. A subgroup difference is not adverse impact, which depends on how ratings are used, the selection ratio and the applicant pool, and a low difference does not make a method good on its own. Those figures describe the method at large. The four questions above are what tell you whether you are running it.

The last check is whether the loop needs the number of stages it has, because structure applied to a stage that establishes nothing new is careful work spent on a duplicate (count the claims your loop tests).

See how it works

Common questions

Is a structured interview the same as a scripted one?

No. What is fixed is the set of questions every candidate faces and the scale their answers are rated against; what stays free is how far the interviewer probes an answer, as long as that probe is available to every candidate who gives a similar answer. A structured interview usually goes deeper per question than an unstructured one, because the interviewer is not spending attention deciding what to ask next.

Does structure make the interview worse for candidates?

It usually makes it better, and it certainly makes candidates comparable. Job-related, consistent questions set clearer expectations, and a candidate who gets a follow-up nobody else got has been assessed on a different interview without being told. What does frustrate people is structure applied badly: reading questions off a sheet without listening, refusing to answer clarifying questions, or asking about a job the questions do not resemble.

How do you write an anchored rating scale?

Take the question and write what a weak, an adequate and a strong answer actually contains, in the words a real answer would use. “Names the risk and says how they would check it” is an anchor. “Demonstrates good judgment” is not, because two raters will fill it in differently. Three levels are enough for most competencies, and the middle one is the hardest and most useful to write. Test the anchors against a few remembered answers before the first real interview.

Is a job description enough, or does this need a job analysis?

A description is a starting point and rarely enough on its own, because it lists duties rather than what separates someone who does them well. The lightweight version is a conversation with two or three people doing the job now: which tasks take the most judgment, where new hires struggle in the first quarter, what a mistake looks like. Four or five competencies out of that is enough to build questions from, and the notes are the record that ties the interview to the role.

Are panel interviews structured by default?

No, and a panel can make things worse. Several people hear the same answer and then discuss it, which produces one shared reading rather than several independent ones. A panel is structured when the questions are fixed, the scale is anchored, and every panelist records ratings before anyone speaks. Without that last step, a panel interview is an unstructured interview with more witnesses.

Do candidates get the questions in advance?

Often worth doing, and it costs less than teams fear. Sending the competencies or the questions ahead reduces the advantage held by candidates who have been coached on interview format, and it moves the assessment toward what someone can actually do rather than how quickly they can compose under pressure. Keep back anything where the surprise is the point, such as a case with a deliberate error in it.

References

  1. 1. Structured Interviews: A Practical Guide U.S. Office of Personnel Management, 2008. opm.gov Supports the three defining properties quoted here (same questions in the same order, a common rating scale, agreement in advance on what an acceptable answer looks like) and the scope statement that this is 2008 US federal guidance rather than statute, binding no private employer.
  2. 2. Revisiting Meta-Analytic Estimates of Validity in Personnel Selection: Addressing Systematic Overcorrection for Restriction of Range Journal of Applied Psychology (American Psychological Association), 107(11), 2040-2068, 2022. gwern.net Supports the .42 and .19 estimates for structured and unstructured interviews, the raw observed .32 and .13, the research coding of structure behind those labels, and the Table 3 pairing of the top-ranked .42 with a Black-White subgroup difference of .23 against .79 for cognitive ability tests.
  3. 3. SHRM Benchmarking: Talent Access (Selection Criteria, Overall) Society for Human Resource Management, 2022. shrm.org Supports the self-reported adoption percentages, which count organisations using the technique rather than candidates assessed, presented as an upper bound because organisations report themselves against SHRM's definition.

3 sources, numbered by first appearance. How Olive sources claims

General guidance for hiring teams. What works at one company and one volume may not transfer to yours.

Olive assesses how a person works with AI. It does not detect AI-written documents, and it never produces a score, a ranking, or a match percentage for a person. Candidates read the same report the employer reads.

Back to answers

Open your first role Ten attempts a month against a live item bank, with a human-written report on every one.