assessment
Work sample tests vs interviews: what the evidence says
Structured interviews predict performance better than work samples on corrected 2022 estimates. What changed, and how to run both properly.
Structured interviews predict job performance better than work sample tests do. That is the opposite of what the recruitment industry has repeated for twenty-five years, and it is what the best current evidence says. The number that matters is not which format you choose. It is whether whatever you choose is structured, because the gap between a structured and an unstructured interview is larger than the gap between an interview and a work sample.
The short answers
- Structured interviews are the single best-validated selection method, at an operational validity of .42 against job performance. Work samples come fourth at .33, behind job knowledge tests and empirically keyed biodata.
- The famous figures are wrong. For twenty-five years the field quoted work samples at .54 and cognitive ability at .51. Those numbers systematically overcorrected for range restriction. Corrected, they are .33 and .31.
- Unstructured interviews are the weak method, at .19. That is what most companies actually run, and it is the thing worth replacing.
- Work samples show larger group differences than structured interviews in the US research, so they are not the low-adverse-impact option they are often sold as. The UK legal test is different and this is not legal advice.
- The gain is in structure, not format. The structured and unstructured estimates are .42 and .19. Those come from separate bodies of research rather than a head-to-head trial, so read the gap as a large difference between two literatures rather than as a measured effect of adding structure.
- The two combine well because they fail differently. A structured interview tests reasoning the candidate can articulate. A work sample tests reasoning under a constraint they cannot rehearse.
- A recorded, narrated work sample scored against a published standard is a work sample with the property that makes structured interviews work. That is the design this site argues for, and the evidence for it is the structure, not the format.
Two clarifications before the numbers, because this post is easy to misread.
This is not an argument for dropping work samples. It is an argument that the property doing the predicting is structure: a fixed task, a scoring key written in advance, independent scores, and a standard rather than a comparison. That structure is available in both formats. An unstructured work sample is not a better instrument than an unstructured interview.
It is also not an argument that validity is the only thing worth optimising. Later sections cover what a work sample does that no interview can, and why that matters more for some roles than the average coefficient suggests.
What the evidence actually says
In 2022, Paul Sackett, Charlene Zhang, Christopher Berry and Filip Lievens published a systematic re-analysis of the meta-analytic estimates the selection field had been using since 1998. Their finding was that the standard corrections for range restriction had been applied in ways that substantially overstated how well most selection methods predict performance.
The corrected table is the one to work from.
| Selection method | Corrected validity | Previously quoted | Change |
|---|---|---|---|
| Structured interviews | .42 | .51 | −.09 |
| Job knowledge tests | .40 | .48 | −.08 |
| Empirically keyed biodata | .38 | .35 | +.03 |
| Work sample tests | .33 | .54 | −.21 |
| Cognitive ability tests | .31 | .51 | −.20 |
| Integrity tests | .31 | .41 | −.10 |
| Assessment centres | .28 | .37 | −.09 |
| Situational judgement tests | .26 | not reported | not reported |
| Contextualised conscientiousness | .25 | not reported | not reported |
| Unstructured interviews | .19 | .38 | −.19 |
Two things in that table matter more than the rest.
Work samples fell furthest. The .54 figure came from a 1974 narrative review that predated meta-analysis as a technique. Replacing it with a proper meta-analysis of 54 studies dropped it to .33. Anyone still quoting .54 at you is quoting a number the field abandoned.
Structured interviews came top. Not by a wide margin, and the ranking is less interesting than the level, but the method the industry spent a decade calling unreliable is the best-validated thing available.
The paper’s own summary of its finding is blunt: the same procedures that ranked high before still rank high, but with validity estimates reduced by .10 to .20, and “selection predictor–criterion relationships are considerably lower than previously thought.”
Why the old numbers were wrong
This is worth understanding rather than taking on trust, because it determines which other statistics in hiring you should distrust.
A validity study correlates a selection score against later job performance. The sample is people you hired, which is a narrower range of ability than the people who applied, because you rejected the bottom. Correlations computed on a restricted range understate the true relationship, so the field corrects for it.
The correction requires knowing how restricted the range was. Most meta-analyses did not know, so they assumed a distribution and applied it to every study. Sackett and colleagues show that these assumed distributions were routinely too aggressive, and that applying them to concurrent studies, where you test people already employed and range restriction largely does not apply, inflates the result substantially.
For structured interviews, the range restriction correction alone increased the estimate by 38%. For unstructured, 41%. Those increases were mostly artefact.
The recommendation the authors give is the one to carry into any statistic you are quoted: absent a sound basis for estimating the degree of range restriction, it is better not to attempt a correction at all.
Structure is the variable, not format
Look at the two interview rows again. Structured .42, unstructured .19. The same method, the same candidates, the same interviewers, and a difference of .23. That is larger than the entire gap between structured interviews and work samples.
That is the finding to design around. The recruitment industry frames this as a format question, because formats are products and structure is discipline. But no work sample rescues a process that scores candidates against each other from memory, and a genuinely structured process gets most of the available benefit whichever format it runs in.
Structure, concretely, means four things.
The same questions, in the same order, for every candidate. Not a topic list. The actual wording, because a question rephrased is a different question and scores from differently-worded questions are not comparable.
A scoring key written before you meet anyone. What a 1 looks like, what a 3 looks like, what a 5 looks like, in behavioural terms, with an example of each. Written after you have met candidates, it will describe the ones you liked.
Scores recorded independently before discussion. A panel that debriefs before scoring produces one opinion held by three people. The most senior voice in the room sets the anchor within about ninety seconds.
Scores against the standard, not against the field. Ranking three candidates tells you who was best of three. A published standard tells you whether any of them clear the bar, which is the question you actually have. A weak field otherwise produces a hire and a strong field produces a rejection.
Most processes that describe themselves as structured do the first of these and none of the other three.
What a work sample is genuinely better at
The validity coefficient is one number summarising an average relationship. It is not the only property worth having, and there are things a work sample does that a structured interview cannot.
It is not rehearsable. Interview answers improve with practice at answering interviews, which is a skill with a weak relationship to the job. A candidate can prepare thoroughly for a work sample and still has to do the work in front of you.
It surfaces process rather than outcome. An interview answer is a reconstruction assembled after the fact with the false starts removed. The false starts are the diagnostic part. Watching someone work shows the order they attack a problem in, which is most of what operations competence is.
It is evidence a third party can re-examine. A score plus a recording can be reviewed by someone who was not in the room. An interview impression cannot. That matters for consistency across interviewers and it matters if a rejection is ever challenged.
It transfers. A candidate who has done a real piece of work has something to show the next employer. An interview produces nothing the candidate keeps.
Those are real advantages and none of them is a validity coefficient. Be honest about which claim you are making.
The cost nobody selling work samples mentions
Validity is not the only property that matters. A selection method also differs in how much its scores vary between groups of candidates, and a method with larger differences produces more adverse impact when you select on it.
The 2022 paper reports these alongside each validity estimate. The short version is that structured interviews are unusual in being both the most valid method and among the least differentiating, while work sample tests are less valid and show substantially larger group differences, roughly three times the gap on the same standardised measure. Cognitive ability tests show the largest differences of any common method.
Two caveats belong with that, and they matter more than the ranking.
The underlying data is American. These estimates come from US validation research, and the group comparison they report is a US one. Whether the same differences appear in a UK applicant pool is not established by this research, and it should not be assumed. Treat the direction as informative and the magnitudes as belonging to the population they were measured in.
In the UK the legal frame is different. A selection method that disadvantages a protected group raises indirect discrimination under the Equality Act 2010, where the question is whether the practice is a proportionate means of achieving a legitimate aim. That is not the same test as the US one, and nothing here is legal advice. If your process has adverse impact you need an employment lawyer, not a validity table.
What survives both caveats is narrow and still useful: adopting a work sample because a vendor described it as the fairest option available is adopting it on a claim the research does not support. If reducing adverse impact is a stated goal, that is a question to take advice on rather than to settle with a format choice.
The follow-up paper makes the constructive version of the point. Because cognitive ability is no longer the dominant predictor, dropping it from a combination of predictors now costs about .05 of validity rather than the .20 it used to. Equally valid combinations with smaller group differences are now buildable, which was not true under the old numbers.
Where work samples go wrong
They are not automatically better, and four failure modes account for most of it.
Too long. A multi-day take-home filters for candidates with spare evenings. The strongest operations people are employed, busy, and decline. You have built a filter for availability and called it a filter for capability. Sixty to ninety minutes is the ceiling for anything unpaid.
Testing the wrong task. A work sample only measures the job to the extent it resembles the job. A task chosen because it is easy to score, a tool configuration exercise or a written case with one right answer, measures something adjacent and scores it confidently.
Unpaid work with commercial value. Asking a candidate to produce a strategy for your actual accounts is not an assessment, and candidates who have seen it before will withdraw. Use a realistic scenario built from a situation you have already solved.
No scoring key. An unstructured work sample is an unstructured interview with a spreadsheet attached. Whatever validity the method has comes from consistent scoring, and a rubric written after the fact describes the candidates you liked.
The last one is the common one. Companies adopt the format and skip the discipline, then conclude the format did not work.
Where interviews reliably mislead
Everything below applies to unstructured interviewing, the .19 method, and is largely fixed by structure.
Fluency reads as competence. A confident, well-organised answer sounds like expertise. It is evidence of having answered the question before. The correlation between how well someone describes work and how well they do it is real but much weaker than interviewers believe.
Reconstruction beats observation. Ask how someone would debug a discrepancy and you get a tidy narrative with the dead ends removed. You are assessing their editorial judgement about their own past work.
The bar moves. Without a written standard, panels score against the field.
First impressions propagate. An unstructured interviewer forms a view early and spends the remainder gathering support for it. Structure interrupts this by fixing the questions before the impression forms.
None of this makes interviews useless. It makes unstructured interviews weak, and almost every company running “four rounds” is running four unstructured ones and counting the repetition as rigour. Four unstructured interviews is not more valid than one. It is the same weak measurement taken four times, and the agreement between them feels like corroboration.
How to combine them
The two methods fail differently, which is the condition under which combining predictors adds something rather than duplicating.
A defensible process for a senior operations role, in order:
- A structured screening conversation, 30 minutes. Same questions, scored against a written key. This is your highest-validity instrument. Treat it as an assessment, not as a chat.
- A work sample, 60 to 90 minutes, narrated and recorded. A realistic problem, worked through out loud. Scored against a published standard rather than against the other candidates.
- A structured panel on the work sample, 60 minutes. Questions derived from what they actually did, which is unrehearsable by construction and where the two methods compound.
- References after the offer conversation, asked about specific behaviours you observed rather than for a general impression.
What that sequence buys you is two independent readings of reasoning, one they can prepare for and one they cannot, with a written standard behind both.
What it does not buy you is certainty. A validity of .42 is the best available and it is not close to deterministic. Any process claiming to eliminate hiring risk is selling something. The honest claim is narrower: the research on structured processes reports substantially higher validity than the research on unstructured ones, and that difference is large enough to be worth the discipline.
Why the recording matters
A scored work sample gives you a number. A recorded one gives you the reasoning that produced it, and for operations roles the reasoning is the thing being bought.
Four things are visible in a narrated recording and invisible in a score.
Decision sequencing. Whether they check the obvious causes before the interesting ones. Ops work rewards boring order, and the person who checks whether a provider changed its schema before theorising about market conditions will fix problems faster for years.
Elimination logic. Whether possibilities get ruled out systematically or a conclusion gets picked and then defended.
Pause points. Hesitation is information. The question is whether it reads as care or as unfamiliarity, and that is audible.
Acknowledgement of gaps. In operations the person who guesses silently is the liability. Listen for whether they say when they do not know.
The recording also does something a score cannot: it makes the assessment auditable. A hiring manager who disagrees with a score can watch the session. That is the mechanism by which a published standard stays honest, and it is why the standard has to be published rather than described.
How settled is any of this
Less than the decimal points suggest, and the authors say so themselves. Their closing position is that they do not put a stake in the ground on these estimates and expect further refinement as better data becomes available.
Read the figures as the best current estimates rather than as constants. Two practical consequences follow. A gap of .02 between two methods is not a reason to choose one over the other. And the ranking is more durable than the levels: the methods that came top before still come top, at lower absolute values.
What is not in doubt is the direction of the 2022 correction. Every widely quoted validity figure from the 1998 compilation is an overestimate, and the three that circulate most in recruitment marketing are the ones that moved furthest.
| Claim still in circulation | What the corrected research reports |
|---|---|
| Work samples are the most predictive method, at .54 | .33, and fourth in rank |
| Cognitive ability predicts performance at .51 | .31 |
| Interviews barely predict performance | True of unstructured interviews. Structured interviews rank first at .42 |
If a supplier quotes the older figures, the useful question is which meta-analysis they come from and what year, because the answer is usually 1998.
Related reading
- What a recording shows that an interview cannot, for the mechanism behind the narrated assessment.
- RevOps interview questions that predict performance, for the structured questions this post argues you should be scoring.
- What AI-written applications did to screening, for why the written layer stopped filtering.
- How to hire a RevOps manager, for the process these stages sit inside.
- Why your RevOps job description attracts the wrong people, for the filter that runs before any of this.
- How long a RevOps search takes, if you are weighing an extra stage against elapsed time.
Questions
What people ask about this.
- Are work sample tests better than interviews?
- Not on validity. The best current meta-analytic estimates put structured interviews at .42 against job performance and work sample tests at .33. The widely quoted claim that work samples lead the field at .54 comes from a 1974 narrative review that a 2022 re-analysis corrected. What is true is that unstructured interviews are much weaker than both, at .19, and that is what most companies actually run.
- What is the most valid selection method?
- Structured interviews, at an operational validity of .42, followed by job knowledge tests at .40 and empirically keyed biodata at .38. Work samples come fourth at .33 and cognitive ability tests fifth at .31. These are the corrected 2022 figures. The values quoted before that correction were between .10 and .20 higher across almost every method, because the standard range restriction corrections were applied too aggressively.
- Why did the validity numbers change in 2022?
- Because the corrections applied to them were wrong. Validity studies are run on people you hired, which is a narrower range than the people who applied, so the field corrects for range restriction. Most meta-analyses assumed a restriction distribution rather than measuring one, and applied it to concurrent studies where restriction largely does not apply. Sackett and colleagues showed this systematically overstated validity, by 38% for structured interviews from the range restriction correction alone.
- Is cognitive ability the best predictor of job performance?
- No. This was the largest single correction in the 2022 re-analysis, from .51 down to .31, and structured interviews now rank above it. A 2023 follow-up found that removing cognitive ability from a combination of predictors now costs about .05 of validity against the .20 it used to, so a composite that leans on it less is no longer the sacrifice it once was.
- Are work samples the fairest selection method?
- Not on the available evidence, though they are often sold that way. In the US research, work sample tests show substantially larger differences between candidate groups than structured interviews do. That data is American and the UK legal test is different, so treat it as a reason to take advice rather than as a settled answer. Nothing here is legal advice.
- How long should a work sample be?
- Sixty to ninety minutes for anything unpaid, and pay for anything longer. A multi-day take-home selects for candidates with spare evenings rather than for capability, and the strongest operations candidates are employed and will decline. Length is the most common way a work sample fails, and it fails silently because the people it excluded never tell you why.
- What makes an interview structured?
- Four things, and most processes do only the first. The same questions in the same wording for every candidate. A scoring key written before you meet anyone, with behavioural anchors for what each score looks like. Scores recorded independently before any discussion. And scoring against a published standard rather than against the other candidates. Skipping the scoring key is what turns a structured process back into an unstructured one.
- Are four interview rounds better than one?
- Not if they are unstructured. Four unstructured interviews is the same weak measurement taken four times, and the agreement between them feels like corroboration when it is mostly shared bias and a propagating first impression. One structured stage with a written scoring key predicts better than four unstructured conversations and costs the candidate a great deal less time.
- Should you combine a work sample with an interview?
- Yes, because they fail differently, which is the condition under which adding a second method adds information rather than duplicating the first. A structured interview tests reasoning the candidate can articulate and prepare. A work sample tests reasoning under a constraint they cannot rehearse. A structured panel built on what the candidate actually did in the work sample compounds both.
- Why record the assessment rather than just score it?
- Because for operations roles the reasoning is the thing being bought, and a score compresses it away. A recording makes decision sequencing, elimination logic, hesitation and admissions of uncertainty visible. It also makes the assessment auditable: a hiring manager who disagrees with a score can watch the session, which is the mechanism that keeps a published standard honest.
- Does a structured process help if a hiring decision is challenged?
- Consistency is easier to evidence than an impression, which is the same property that makes a structured process more valid: the same task, the same scoring key, applied identically and written down. That is a general observation rather than legal advice. If you are worried about how a specific process would stand up, that is a question for an employment lawyer, and in the UK the relevant framework is the Equality Act 2010.
- What is the single highest-return change to a hiring process?
- Writing the scoring key. Whatever stages you run, define what a good answer looks like before you meet anyone, have each assessor score independently, and compare only after everyone has written their scores down. Most of the difference between an unstructured process at .19 and a structured one at .42 is that document, it costs half a day, and almost nobody does it.
Tell us the role. We will tell you honestly whether we can fill it.
Nothing owed until someone starts.
Book a call