assessment

AI-written applications broke technical screening

Every candidate now arrives with a polished written account of themselves. What that does to CV screening, and what still carries signal once it stops.

Rareix · · updated · for employers

AI-written applications broke technical screening

A remote operations role now draws several hundred applications inside a few days, and nearly all of them are written or rewritten by a model.

This is not a complaint about candidates. Using the available tools to produce a better-written application is rational, and anyone advising otherwise is asking candidates to compete with one hand tied. The problem is what it does to the employer’s side of the process, which was built on an assumption that no longer holds.

The short version, which several other posts on this site borrow: the written layer stopped carrying information, and the weight that layer was holding has moved onto stages that were never built to hold it.

The short answers

  • Volume tripled and selectivity halved. Applications per hire have gone past 300 since 2021, and candidates are roughly half as likely to reach an interview.
  • CV screening was never accurate. It was cheap, and it worked because producing a polished application took effort, and effort correlated loosely with interest.
  • Effort collapsed toward zero and took the correlation with it. The filter now costs the same to run and selects for close to nothing.
  • The weight moved onto the interview, which was never good at carrying a decision alone: unstructured interviews sit at .19 operational validity against .42 for structured ones.
  • Banning it, detecting it, adding stages and adding requirements all fail, each for a different and instructive reason.
  • Specificity still carries signal. A named system at a named scale, a number with a direction and a baseline, a decision that would have to be defended.
  • The cost falls on candidates too. Processes stop replying when volume outruns capacity, and the ghosting data is unambiguous.
  • The fix is to invert the funnel, not to add to it: make the real measurement cheap enough to run first.

What screening actually assumed

CV screening was never accurate. It was a cheap filter that worked for a reason nobody wrote down: producing a polished, well-targeted application took an hour of a person’s evening, and an hour of a person’s evening correlated loosely with interest, diligence and, weakly, competence.

That is the whole mechanism. Not the content of the CV. The cost of producing it.

Both halves broke at once. Effort collapsed toward zero, and the correlation with competence went with it. What remains is a filter that costs the same to run and selects for nothing in particular, operated by people who have not yet been told the mechanism it depended on has gone.

What actually changed, in numbers

ThenNow
Applications per hireRoughly a third of today’s, in 2021More than 300
Chance of reaching an interviewnot statedRoughly half what it was
LinkedIn application ratenot stated~11,000 a minute by mid-2025
UK applications per postingnot statedMedian 72
UK interview ratenot stated4.3%
UK offer ratenot stated1.1%
Time to fill43.64 days in 202259.67 days in 2025

The volume figures come from an analysis of more than 100 million applications across 200,000 jobs; the UK figures from 8.8 million UK applications; the time-to-fill series from more than 6,000 companies and 640 million applications. Ashby’s benchmark data puts median time to first fill at 71 days for senior roles and 75 for technical ones.

Two honest caveats about that table.

These are different datasets measuring different populations, and the rows are not strictly comparable to each other. They point the same way, which is the most that can be said.

The market is soft, and that also produces volume. UK vacancies stood at 707,000 in May to July 2026, down 2.7% on the year and 10.3% below early 2020, with about 2.5 unemployed people per vacancy. KPMG and the REC reported permanent placements stabilising in July 2026 after a 45-month downturn. A slack market raises applications per role on its own, and none of these datasets separates that from the effect of cheaper application production. What a slack market does not explain is the specific thing this post is about: applications got better written while carrying less information.

Where the weight went

This is the part worth dwelling on, because it explains why processes feel worse rather than merely busier.

A hiring process is a sequence of filters, and each one is implicitly assigned an amount of the decision. When the first filter stops working, its share does not disappear. It moves down the sequence, onto stages that were sized for a different job.

The stage it lands on is usually the interview, and unstructured interviewing is the weakest common method in the field: an operational validity of .19 against job performance, on the corrected 2022 estimates, against .42 for structured interviews. That was survivable when the CV was a weak signal and the interview was carrying part of the decision. It is not survivable now that the CV is no signal and the interview is carrying all of it.

The second place the weight lands is the screening call, which was designed to confirm availability and salary expectations and is now being asked to assess capability in twenty minutes by someone who often has not done the job. Employ Inc’s benchmarks put 7.2 days between application and that first screening conversation, across 6,640 companies, which tells you the stage is a bottleneck as well as an overloaded filter.

The third place, increasingly, is a take-home, which is the worst of the available options: it measures output, output is precisely the thing that got cheap, and it costs the candidate several hours for a signal that was already degraded before it arrived.

The responses that do not work

Banning it. Unenforceable, and it selects for rule-followers over strong candidates. The people who ignore the instruction are not a worse cohort, and you have no way to identify them anyway.

Detectors. Not reliable enough to reject someone on. The specific problem is not that they are useless in principle. It is that no detector we are aware of publishes a false-positive rate for the populations you would be applying it to, and a screening tool whose error rate you cannot state is one you cannot explain to a candidate or defend if a rejection is questioned. In the UK the relevant framework is the Equality Act 2010, and how it bears on a particular screening tool is a question for an employment lawyer rather than for us. Nothing here is legal advice.

More stages. Adding interviews increases the cost of the process without adding a filter that measures anything different. Four unstructured conversations is the same weak measurement taken four times, and the agreement between them feels like corroboration when much of it is a shared first impression propagating. You also lose the candidates with options first, which is a filter running in exactly the wrong direction.

More requirements in the advert. Everyone now meets every stated requirement on paper. A longer list does not raise the bar; it lowers the response rate, and it does so unevenly: the people who read a requirements list as a specification rather than a wish list are disproportionately the ones you wanted.

Keyword filtering harder. Keyword matching was another proxy for effort and familiarity, and both got cheap at the same time. It still removes the applications that did not try. That is a smaller filter than it used to be, and it is not one you can build a decision on.

There is a structural parallel worth noticing. Google spent the last two years rewriting its spam policies around scaled content abuse, because cheap generation broke a ranking system that had implicitly priced effort. The response was not to detect generated text. It was to change what the system rewards. Hiring has the same problem and has not yet had the same reckoning.

What still carries signal

Specificity that would be costly to fabricate. “Migrated 40,000 accounts from Salesforce Classic with a three-day freeze” is checkable in conversation. “Experienced in CRM migrations” is not. A model will produce the second happily, and the first only if the candidate supplies the detail, at which point the detail rather than the sentence is the signal.

Numbers with a direction and a baseline. Not “improved conversion” but “improved stage-3 conversion from 22 to 31 per cent over two quarters, mostly by killing a qualification step”. The baseline is the part that is hard to invent, because it invites the follow-up question.

Decisions the candidate would have to defend. Something they chose, an alternative they rejected, a cost they accepted. Fabricated decisions collapse on the second question, and real ones get more interesting.

Anything observed rather than reported. Which is the whole argument below.

What does not carry signal, and largely never did: writing quality, length, enthusiasm, tailoring, and the presence of the right nouns. Those were proxies for effort. They are now proxies for nothing.

Invert the funnel

The conventional funnel puts the cheap filter first and the expensive measurement last, on a shortlist. That ordering made sense when the cheap filter worked.

The response that adds information rather than adding stages is to move a real measurement to the front. Marking a structured exercise against a fixed rubric is now inexpensive enough to run on everyone rather than on a shortlist, which inverts the usual order: the measurement comes first and the CV becomes context.

Three properties make that work, and removing any one of them returns you to the problem.

It has to be short. Sixty to ninety minutes, or you have built a filter for spare evenings.

It has to be marked against a standard written in advance, not against the other candidates, or you have built a ranking of whoever applied this month.

It has to measure something a model cannot produce on the candidate’s behalf. Which in practice means watching the work happen rather than reading its output.

A model can write a convincing account of how someone debugs a reporting discrepancy. It cannot, in real time and under a follow-up question drawn from a specific moment, produce the hesitation that reveals whether the person has done it before. The order someone works in, what they rule out and why, the moment they notice the data contradicts them: none of that is available to be written in advance, because it is not a document. It is a behaviour.

That is the argument for a recorded, narrated work sample, made at length in what a recording shows and published in full as the assessment. The evidence for it is the structure rather than the format. Structured interviews still lead the validity table, and a work sample without a rubric is not an improvement on anything.

What the process looks like after the inversion

Concretely, for a specialist operations role. The conventional funnel is on the left, the inverted one on the right.

StageConventionalInverted
1CV screen, ~300 applicationsAdvert that states the problem, band published
2Recruiter screening call, ~30 candidatesShort structured exercise, marked against a rubric, open to everyone who applies
3Hiring manager interview, ~10CV read as context, for the 15 who scored above the bar
4Take-home, ~5Structured interview on the scorecard, ~6
5Panel, ~3Panel built from what they did in the exercise, ~3
6OfferOffer

What moved: the measurement went from stage 4 to stage 2, and the CV went from a filter to a document you read after you already have an opinion formed from evidence.

Three consequences, two good and one that costs money.

The stage that carries the decision is now the one designed to carry it. The interview stops being asked to do a job it is bad at, and becomes a conversation about something you have both seen.

Candidates who would have been filtered out on paper get measured. Which is the part that changes who you hire rather than merely how you feel about the process. The people the written layer was misranking are, by construction, the ones you were missing.

It costs more at stage 2 than a CV screen does. This is the honest catch. A cheap filter has been replaced by one that is merely inexpensive, and it runs on a larger population. Whether that is affordable depends entirely on how much of the marking is automated against a fixed rubric, which is a tooling question rather than a philosophical one, and it is why this response was impractical five years ago and is not now.

The thing not to do is run the inverted funnel with an unstructured exercise. A recorded conversation with no fixed task and no rubric is a video interview, which is the .19 method with a camera pointed at it, applied to more people.

If you are dealing with the volume right now

Three things that help before any of the above is in place.

Close the advert early and reopen it if needed. A posting left open for six weeks accumulates applications at a rate unrelated to how many you can assess. Most of the value arrives in the first week.

Screen on one specific, checkable thing rather than on overall impression. Pick the single requirement that genuinely cannot be worked around, usually a named system at a named scale, and sort on that alone. It is a weak filter, but it is a weak filter you can describe, which is more than overall impression offers.

Reply to everyone you rejected, with a template. It costs one afternoon of setup. Given the ghosting figures below, it is also the cheapest reputational advantage available in a market where most processes have stopped doing it.

Fix the advert while you are here

If the top of the funnel is the problem, the advert is the cheapest thing to change.

Appcast’s analysis of its own advertising network, across more than 400 companies, found apply rate peaking at 8 to 8.5% for postings of 201 to 400 words, against about 4.5% below 200 words and under 5% above 701. Job titles of one to three words averaged 3.41 clicks against 2.75 for titles of 13 or more, with the best-performing range at four to six words. The publisher does not break those figures out by country, so treat them as directional rather than as a UK benchmark.

The practical version: lead with the problem the role exists to fix, keep the whole thing under 400 words, put the tool list at the bottom marked as context rather than requirement, publish the band, and include one paragraph on what the first six months actually contain. That last one is the closest thing to a working filter left in the written layer, because it is the paragraph a strong candidate reads to decide whether the job is real. The longer treatment is in why your RevOps job description attracts the wrong people.

The cost that falls on candidates

Worth stating plainly, because it is the part of this that is not an operational inconvenience.

When volume outruns screening capacity, processes stop replying. One analysis of more than 200,000 conversations found 72% of candidates in an active conversation on an open role going 30 days or more with no logged follow-up. A survey of 1,024 job seekers found 53% reporting having been ghosted, 28% of those after submitting an application and 20% after a first interview.

The comparison that makes it sting: when a recruiter approaches a candidate, the median reply comes in 3.9 days and 70.8% reply within seven, across more than 230,000 outreach threads. Candidates answer quickly. The asymmetry is not about attention.

A cheap, consistent measurement early in the funnel is partly an efficiency argument and partly this: a process that measures everyone against the same standard can also answer everyone, and one that cannot keep up with its own inbox will not.

What this post does not claim

Three limits.

None of the datasets here isolates AI as the cause. They measure volume, selectivity and elapsed time, in a period when the labour market also loosened. The direction is consistent across independent sources; the attribution is an inference and is labelled as one.

“AI-written” is not measurable at the scale being discussed. Nobody knows what share of applications are model-produced, because there is no reliable way to count them, which is the same fact that makes detectors undependable.

Watching the work is not a solved problem either. It raises the cost of faking without eliminating it, and the evidence supporting it is evidence about structure rather than about recording. The honest claim is that it replaces a filter that has stopped working with one that has not, not that it is certain.

The evidence on conventional hiring practice was unflattering before any of this. Peter Cappelli’s survey of the field remains the standard reference. What changed is that the cheapest stage stopped working, and most companies have responded by keeping it and pretending.

Questions

What people ask about this.

How much has application volume actually changed?
Applications per hire have tripled since 2021 to more than 300, and candidates are roughly half as likely to reach an interview, measured across more than 100 million applications and 200,000 jobs. By mid-2025, job seekers were submitting applications to LinkedIn at roughly 11,000 a minute. In the UK specifically, an analysis of 8.8 million applications found a median of 72 applications per posting, a 4.3% interview rate and a 1.1% offer rate.
Is this just a soft labour market rather than an AI effect?
Both are happening and they are hard to separate, which is worth saying rather than glossing. UK vacancies stood at 707,000 in May to July 2026, down 2.7% on the year and 10.3% below early 2020, with about 2.5 unemployed people per vacancy. That is a market that produces more applications per role regardless of how they were written. What a slack market does not explain is why the written quality of applications rose while their information content fell.
Should we ban AI use in applications?
Unenforceable, and it selects for candidates who follow instructions rather than candidates who are good. Using the available tools to produce a better-written application is rational, and the people who ignore the instruction are not a worse cohort. The productive response is to screen on something a model cannot produce on the candidate's behalf.
Do AI detectors work?
Not well enough to reject someone on. The specific problem is that no detector we are aware of publishes a false-positive rate for the populations you would be applying it to, and a screening tool whose error rate you cannot state is one you cannot defend if a rejection is ever questioned. In the UK the relevant framework is the Equality Act 2010; this is a general observation rather than legal advice.
Is a take-home test the answer?
Not on its own. A take-home measures the output, which is the part that is now cheap to produce, and it filters for candidates with spare evenings rather than for capability. What survives is watching the work happen, because the reasoning is much harder to outsource in real time and a follow-up question drawn from a specific moment is hard to answer if you did not do the work.
What still carries signal in an application?
Specificity that would be costly to fabricate: a named system at a named scale, a number with a direction and a baseline, a decision the candidate would have to defend. Generic competence claims carry none, and never did. What changed is that generic claims are now well written, so the writing quality that used to stand in for effort no longer separates anything.
Should we add more interview rounds to compensate?
No, and this is the most common expensive mistake. Adding a fourth unstructured conversation is the same weak measurement taken a fourth time. Unstructured interviews sit at .19 operational validity against job performance, against .42 for structured ones. More rounds also cost you the candidates who have options first, which is a filter running in the wrong direction.
Does an ATS keyword filter still help?
Less than it did, and in the same direction as everything else here. Keyword matching was a proxy for effort and domain familiarity, and both are now cheap to produce. What it still does reliably is remove the applications that did not try, which is a smaller and less useful filter than it was when trying was expensive.
How should the advert change?
Shorter, and specific about the problem. Appcast's analysis of its own advertising network found apply rate peaking at 8 to 8.5% for postings of 201 to 400 words against under 5% above 701, and one-to-three-word titles averaging 3.41 clicks against 2.75 for titles of 13 or more. The publisher does not break those out by country, so treat them as directional. A long requirements list no longer raises the bar, because everyone meets every stated requirement on paper now.
What does the volume do to candidates?
It makes processes stop replying. One analysis of more than 200,000 conversations found 72% of candidates in an active conversation on an open role going 30 days or more with no logged follow-up, and a survey of 1,024 job seekers found 53% reporting having been ghosted. That is a cost of the same breakdown, it falls on people rather than on companies, and a cheap consistent early measurement is partly an argument about not doing it.
Is hiring getting slower because of this?
Time to fill has been rising, though the causes are not cleanly attributable. Greenhouse's benchmarks report time to fill going from 43.64 days in 2022 to 59.67 in 2025 across more than 6,000 companies. Ashby puts median time to first fill at 71 days for senior roles and 75 for technical ones. More applications with less information in them is consistent with that direction, but neither dataset isolates it as a cause.
What is the actual fix?
Invert the funnel. Make the measurement cheap enough to run early rather than reserving it for a shortlist, and let the CV become context rather than the filter. That is a tooling change more than a philosophical one: marking a structured exercise against a fixed rubric is now inexpensive enough to run on everyone, which is the only response that adds information rather than adding stages.

Tell us the role. We will tell you honestly whether we can fill it.

Nothing owed until someone starts.

Book a call