assessment

What a recording shows that an interview cannot

An interview measures how someone describes their work. A recorded, narrated task measures the work. The gap between those two is where hiring mistakes live.

Rareix · · updated · for employers

What a recording shows that an interview cannot

An interview measures how well someone describes their work. A recorded, narrated task measures the work.

Those are different skills, and the gap between them is where expensive hiring mistakes live. It has widened sharply now that every candidate arrives with a polished written account of themselves.

This post is about the mechanism rather than the format. The argument for work samples over interviews has already been made on the evidence, and the honest version of it is narrower than the industry version. What follows is what a recording adds on top of a scored exercise, why that matters more for operations roles than for most, and what it does not do.

The short answers

  • The claim is not that recordings predict better than interviews. On the corrected 2022 estimates, structured interviews reach .42 against job performance and work samples .33. Anyone selling you a recording on validity grounds is quoting numbers the field abandoned.
  • The claim is that structure predicts and a recording makes structure checkable. A fixed task, a rubric written in advance, independent scores against a published standard. The recording is the evidence layer under all of that.
  • Auditability is the actual product. A score compresses reasoning into a number. A recording lets a third party re-examine the reasoning that produced it, which is the mechanism that keeps a published standard honest.
  • Four things are visible in narration and invisible in a score: decision sequencing, elimination logic, where they slow down, and whether they say what they do not know.
  • This got urgent rather than merely nice. Applications per hire have tripled since 2021 to more than 300, and candidates are roughly half as likely to reach an interview. The written layer stopped carrying information.
  • It is not a lie detector and not proctoring. It raises the cost of faking. It does not make faking impossible, and we say so.
  • Recording a person working is processing their data. Consent that is genuinely free to refuse, minimal retention, and the recording belonging to the candidate are the conditions that make it defensible. Nothing here is legal advice.

Start with what the evidence does and does not support

It is worth being precise, because this is the part the assessment industry consistently overstates and we sell an assessment.

The best current meta-analytic estimates put structured interviews first among common selection methods at an operational validity of .42, job knowledge tests at .40, empirically keyed biodata at .38, work samples fourth at .33 and unstructured interviews at .19. Work samples fell furthest in the 2022 correction, from a .54 figure that came from a 1974 narrative review. If a supplier is still quoting .54, they are quoting a number the field replaced.

So the defensible claim is not that a recorded work sample beats an interview. It is this:

Structure is what predicts. A fixed task, a scoring key written before you meet anyone, independent scores, and a standard rather than a comparison. Those four properties account for most of the distance between the .42 literature and the .19 one, and they are available in both formats. An unstructured work sample is an unstructured interview with a spreadsheet attached.

The follow-up research favours job-specific measurement. The 2023 design paper finds job-specific measures outranking general construct measures, and that removing cognitive ability from a combination of predictors now costs about .05 of validity rather than the .20 it once did. In plain terms: testing the actual work is a better use of a stage than testing general aptitude, and that conclusion is newly available because the old numbers said otherwise.

The recording adds a property that validity coefficients do not describe. Whether the evidence behind a score can be examined by someone who was not present. No selection method’s validity estimate captures that, and for operations roles it is the thing worth buying.

What you learn from watching, that you cannot learn from asking

The order they do things in. Ask a candidate how they would debug a reporting discrepancy and you get a tidy narrative, usually reconstructed after the fact. Watch them do it and you see whether they check the boring causes before the interesting ones, meaning currency, date ranges, record types and deleted records, before theorising about market conditions. That ordering is most of what separates operators, it compounds over years of small decisions, and nobody describes it accurately about themselves because the false starts are edited out of the retelling.

What they rule out, and why. Strong operators eliminate possibilities out loud. Weak ones pick a plausible answer and defend it. This is visible in thirty seconds of narration and almost invisible in an interview, where a confident conclusion sounds like competence. The diagnostic question is not whether they reached the right answer. It is whether they could have reached a wrong one and noticed.

Where they slow down. The pause before someone opens a tool they claim daily familiarity with is information. So is the moment the data contradicts their hypothesis, and what they do in the four seconds after. Hesitation reads as care or as unfamiliarity, and the difference is audible in a way it is not writable.

Whether they say what they do not know. In operations this is not modesty, it is the core professional skill. The person who guesses silently is the liability, because the guess enters a report and the report enters a board pack. An interview rewards the opposite: fluency, confidence, and an answer to every question.

Whether they ask before they act. A candidate who clarifies the question before touching anything is showing you what they will do with an ambiguous request from a VP. That behaviour has no representation in a CV and only a self-reported one in an interview.

What a score compresses away

A scored exercise gives you a number. For most roles that is enough, because the number stands in for a performance you are content to take on trust.

Operations is the awkward case, for a specific reason: the output of the job is usually a decision about what the numbers mean, and the quality of that decision lives in the reasoning rather than in the result. Two candidates can reach the same conclusion about a broken forecast, one by systematic elimination and one by recognising a pattern they saw at their last company. Both score the same. Only one of them will find the next problem, which will not resemble the last one.

A number cannot carry that distinction. A recording can.

It is also the case that operations candidates are routinely assessed by people who have never done the job. The research on GTM engineering found only 45% of practitioners saying their own company understands the role, from a survey of 228. A recording lets a hiring manager who cannot design the assessment still evaluate the candidate, because the evidence is in a form they can watch rather than a claim they have to accept.

Why this got urgent rather than merely nice

A remote operations role now draws hundreds of applications inside days, and the written layer has become uninformative: the CV, the cover letter and increasingly the take-home are all things a model produces competently.

The volume data is unambiguous. Applications per hire have tripled since 2021 to more than 300, and candidates are roughly half as likely to reach an interview, across more than 100 million applications and 200,000 jobs. By mid-2025, job seekers were submitting applications to LinkedIn at roughly 11,000 a minute. In the UK specifically, an analysis of 8.8 million applications found a median of 72 applications per posting, a 4.3% interview rate and a 1.1% offer rate.

Screening still assumes the paperwork correlates with the person. It no longer does. That leaves the interview carrying the entire weight of the decision, and interviews were never good at carrying it alone. The evidence on conventional hiring practice is not flattering, and Peter Cappelli’s survey of the field remains the standard reference. Unstructured interviewing was survivable when the CV was a weak signal. It is not survivable now that it is no signal.

There is a second-order effect worth naming, because it falls on candidates. When volume rises and screening capacity does not, processes silently stop replying: one analysis of more than 200,000 conversations found 72% of candidates in an active conversation on an open role going 30 days or more with no logged follow-up, and a survey of 1,024 job seekers found 53% reporting having been ghosted. A cheap, consistent measurement early in the funnel is partly an argument about efficiency and partly an argument about not doing that.

Auditability is the actual argument

Here is the mechanism, stated plainly.

A published standard is only honest if someone can check a score against it. If the rubric is public but the evidence is private, the standard is a marketing document. You are asked to trust that a score means what the rubric says it means, and nothing available to you could contradict it.

A recording closes that loop. Four things become possible that are not possible with a score alone:

A hiring manager who disagrees can check. Not argue about it. Watch the four minutes in question and form their own view against the same rubric.

Scores stay comparable across assessors. Drift is the normal condition of any rubric applied by humans over months. It is only correctable if the underlying sessions can be re-marked, which requires them to exist.

The candidate can see the basis of their own result. Which is the condition under which it is reasonable to ask them to do the work at all.

A rejection has evidence behind it rather than an impression. Consistency is easier to evidence than a recollection: the same task, the same rubric, applied identically, with the working visible. That is a general observation about process rather than legal advice, and a recorded process brings obligations of its own, covered below.

That is why the rubric is published rather than described. A standard nobody can audit is a claim, and this business is in the business of not making those.

What a recorded assessment is not

Four things it is worth being explicit about, because the category is oversold.

It is not a validity claim. See the top of this post. The published estimates do not put work samples above structured interviews, and a recording does not change a method’s validity coefficient.

It is not a lie detector. It raises the cost of misrepresentation substantially, because narration under time pressure with follow-up questions is hard to fake. It does not make it impossible. A determined candidate with a second machine could get through a short exercise, which is exactly why the assessment that carries weight is a full sitting, recorded, narrated and followed up.

It is not proctoring, and should not become it. Eye-tracking, keystroke analysis and browser lockdown are a different product with a different relationship to the candidate, and they measure compliance rather than capability. The design here shows the work; it does not police the candidate.

It is not a culture-fit instrument. A single sitting measures diagnostic reasoning and communication under constraint. It does not measure whether someone will still be motivated in year two, and no short exercise does.

The design that makes it work

The properties doing the work are structural, and any of them can be removed to produce something that looks the same and measures nothing.

A realistic task built from a solved problem. Realistic, not real: asking a candidate to produce a strategy for your actual accounts is unpaid commercial work, and the candidates who recognise it withdraw.

A rubric written before anyone takes it. With behavioural anchors for what each score looks like. Written afterwards, it describes the candidates you liked.

Narration as a requirement of the task, not a bonus. The instruction is to think out loud, and it is scored, because unnarrated work returns you to marking an output.

Independent scoring against the standard, not the field. A published bar answers whether this candidate clears it. A ranking answers who was best of the people who happened to apply this month, which is not the question you have.

A fixed length. Sixty to ninety minutes for anything unpaid. Length is the most common way a work sample fails and it fails silently, because the strong candidates it excluded never tell you why.

Remove the rubric and you have an unstructured exercise. Remove the narration and you have a take-home. Remove the published standard and you have a ranking. The format is the least important part.

If you are building one yourself

The design above is not proprietary and the process in this post works without us. Four decisions do most of the work, and each has a cheap wrong answer.

Choose a task you have already solved. Take a real problem from six months ago, strip the identifying detail, and reconstruct the broken state. You need a known answer to mark against, and you need the task to be one where a wrong turn is recoverable inside the time limit. A problem nobody at your company solved is not an assessment, it is a consultation.

Write the rubric before you pilot it, then pilot it on people you already know. Run two current employees through the exercise, ideally one strong and one you would not hire again, and check whether the rubric separates them. If it does not, the rubric is measuring something other than what you believe. This step takes an afternoon and it is the one everybody skips.

Decide in advance what you will do with a disagreement. Two assessors scoring the same session differently is not a problem to be averaged away. Write down who re-marks, against what, and how the rubric gets amended when the disagreement turns out to be legitimate. A rubric that never changes is not being used.

Fix the length and stop selling it as flexible. Sixty to ninety minutes, the same for everyone, communicated before the candidate agrees. “It usually takes about an hour but take as long as you like” is a multi-day take-home with a friendlier sentence in front of it, and it produces exactly the same filter for spare evenings.

The thing not to do is run it unstructured. A recorded conversation with no fixed task and no rubric is a video interview, which is the .19 method with a camera pointed at it. The recording adds auditability to structure; it does not substitute for it.

The objections worth taking seriously

It is more intrusive than a conversation. True. Recording someone working is more exposing than answering questions, and the defence is not that it is harmless but that the exchange is fair: the recording is theirs, it is shared only with employers they name individually, a weak score is never disclosed, and they can withdraw at any point. If the arrangement were one-sided, the objection would stand.

It takes longer than a phone screen. Also true, and it replaces the phone screen rather than adding to it. The comparison that matters is against the whole process, where one recorded session usually removes a screening call and at least one interview round.

It disadvantages nervous candidates. It disadvantages candidates who cannot explain their reasoning, which is a real requirement of the job. The rubric marks what is said rather than how smoothly, which is a design claim about our rubric rather than a property of the format, and it is checkable, because the rubric is published and the sessions exist.

Some candidates need adjustments. Offer them in the invitation rather than waiting to be asked. Extra time, a written brief in advance, or narrating by typing rather than speaking all preserve what the exercise measures. In the UK the relevant framework is the Equality Act 2010, and how it applies to a specific process is a question for an employment lawyer rather than for us.

What happens to the recording

Recording someone working is processing their personal data, and the arrangement only works if that is handled properly.

The design we run: the recording belongs to the candidate, it is shared only with employers they name individually, a weak score is never disclosed, and they can withdraw at any point. Consent in a hiring context has an obvious problem, because a candidate is not in a strong position to refuse. That is why refusing has to cost them nothing, and why an alternative route through the process has to exist and be offered.

The ICO’s employment guidance and its guidance on consent as a lawful basis are the reference points, and the second is the one most often got wrong: consent that a person cannot freely decline is not consent. Beyond that, the ordinary obligations apply, including being clear about retention and collecting no more than the exercise needs.

None of this is legal advice, and it is a description of how we run it rather than a template. If you are building a recorded assessment inside your own process, take advice on it before you record anyone.

What it produces on your side

Three things, before the first interview:

  • The full session, so you can watch the parts you care about
  • A score against a published rubric, part by part
  • A one-page guide of what to ask this specific person, drawn from specific moments

The last of those is the one that changes the interview. A hiring manager who watches none of the recording still runs a materially better conversation with a guide that says at 41:15 they modelled it correctly then over-claimed on the result, probe whether that is a communication habit or a judgement problem.

That is not a question anyone thinks to ask from a CV, and it is the reason the two methods compound rather than duplicate. The structured interview tests reasoning the candidate can prepare. The recorded task tests reasoning under a constraint they cannot rehearse. A panel built from their own session is unrehearsable by construction.

What we cannot claim

Three limits, stated because a standard that only publishes its strengths is not a standard.

We have no published validity study of our own instrument. What we can point to is the design literature on structure, which is not the same thing as evidence about this assessment. A vendor claiming a validity coefficient for their own product should be asked what it was correlated against, in what sample, and over what period.

A single session does not measure long-run performance. It measures diagnostic reasoning and communication under a constraint. Those are load-bearing for an operations role and they are not everything.

The people a score is checked against are a self-selected group. Everyone who has taken this assessment chose to. That is true of every benchmark of this kind and it is worth knowing when you read a score.

The claim that survives all three is narrow and still worth something: you get to see the work before you interview, the standard it was marked against is published, and the evidence behind the score is available for you to check.

Questions

What people ask about this.

Is a recorded assessment more predictive than an interview?
Not on the published evidence, and we will not claim it is. On the corrected 2022 meta-analytic estimates, structured interviews reach an operational validity of .42 against job performance and work sample tests .33. What a recording adds is not validity, it is auditability: the evidence behind a score can be re-examined by someone who was not in the room. That is a different property and it is worth having for different reasons.
Is this not just a take-home test with extra steps?
A take-home measures the output, which is the part a model can now produce. Narration measures the reasoning, which is much harder to outsource in real time, and which is the thing you are actually buying for an operations role. The second difference is that a take-home produces an artefact nobody can check the provenance of, while a recorded session produces evidence with a timestamp on it.
Will good candidates agree to be recorded?
In our experience yes, when three things are true: it is a single sitting rather than a multi-day project, the recording belongs to them, and it is shared only with employers they explicitly name. Strong candidates generally prefer being measured to being guessed at, particularly in a market where a median UK posting draws 72 applications and a 4.3% interview rate. The ones who decline usually decline the time cost, not the camera.
Does narration disadvantage people who are nervous?
It disadvantages people who cannot explain their reasoning, which is a genuine requirement of an operations job rather than an artefact of the format. It should not reward polish, and the rubric is written to mark what is said rather than how smoothly. That is a design claim about our rubric, not a general property of recorded assessments, and it is the reason the rubric is published rather than described.
What about candidates using AI during the assessment?
The screen is recorded, and follow-up questions drawn from specific moments tend to expose someone who did not derive their own answer. We are honest that a determined cheat with a second machine could get through a short exercise, which is why the assessment that carries weight is a full sitting, recorded, narrated and followed up. The realistic claim is that it raises the cost of faking rather than that it makes faking impossible.
What happens to the recording afterwards?
It belongs to the candidate. It is shared only with employers they name individually, a weak score is never disclosed to anyone, and they can withdraw at any point. Recording someone working is processing their personal data, so the arrangement has to be transparent, minimal and consented in a way that is genuinely free to refuse. The ICO's employment guidance is the reference point. Nothing here is legal advice, and if you are designing your own recorded assessment you should take some.
How long should a recorded assessment be?
Sixty to ninety minutes, and pay for anything longer. Length is the most common way a work sample fails, and it fails silently: a multi-day take-home selects for candidates with spare evenings rather than for capability, and the people it excluded never tell you why. Our own assessment is one sitting in four parts.
Can a recording replace the interview?
No, and it should not try. The two fail differently, which is the condition under which adding a second method adds information rather than duplicating the first. The sequence that works is a structured conversation, then the recorded task, then a panel built from what the candidate actually did, with questions that are unrehearsable by construction because they come from their own session.
What does a hiring manager actually watch?
Usually not the whole thing, which is why the output includes a guide keyed to timestamps. A manager who watches none of the recording still runs a materially better conversation with a note that says at 41:15 they modelled it correctly then over-claimed on the result, probe whether that is a communication habit or a judgement problem. The recording is there so the claim can be checked, not so that everyone has to watch an hour.
Does a recorded assessment help if a hiring decision is challenged?
Consistency is easier to evidence than an impression: the same task, the same rubric, applied identically, with the working visible. That is a general observation rather than legal advice, and a recorded process brings its own obligations around consent, adjustments and data. In the UK the relevant frameworks are the Equality Act 2010 and UK data protection law, and a specific process is a question for an employment lawyer rather than for a validity table.
What about candidates who need adjustments?
Ask before the assessment, not after it, and make the offer of adjustments part of the invitation rather than something a candidate has to request. Extra time, a written brief in advance, or narration by typing rather than speaking all preserve what the exercise measures. A format that cannot accommodate any of that is measuring the format.
What does this not test?
Stakeholder management over months, judgement about when to stop building, and whether someone is right for your specific team. A single session measures diagnostic reasoning and communication under a constraint. Anyone claiming a short assessment measures culture fit or long-run performance is selling something, and the honest version of our own claim is that it replaces the CV screen, not the hiring decision.

Tell us the role. We will tell you honestly whether we can fill it.

Nothing owed until someone starts.

Book a call