A few days ago I came across a piece on how AI screens résumés: https://www.quora.com/How-does-AI-actually-read-a-resume-What-is-it-even-looking-for
It is well written and clearly argued. One passage goes roughly like this:
Early systems relied on keyword matching. If you wrote "managed projects" and the job ad said "project management", a crude matcher would miss it. Today it's different: current systems use semantic matching. They know the two mean the same thing, and that Python and Django are related. They turn both the job description and the résumé into vectors and compute a similarity, which gives a score rather than a simple keyword hit.
My first reaction was: that sounds reasonable — but it describes how it should work.
I had recently measured how it actually works on real résumés, and found two sizeable gaps.
The data and the method
Twelve real résumés from nine people. One person gave me three different write-ups of the same career history, and that will matter later. Nine job advertisements, also real.
I broke each advertisement into individual requirements and judged them one at a time: is there anything in this résumé that supports this requirement? The rules were fixed before reading any résumé. A job title on its own does not count as evidence — you have to be able to point at a specific line. Synonyms, broader terms and reasonable inferences all count; it is not a string comparison. I also ran a control group, deliberately pairing résumés with advertisements from unrelated fields, to see whether the judgements would separate.
Three things up front, or the numbers below can't be judged:
The sample is small. Twelve résumés is not a study; it is a probe.
The judging was done by a frontier language model, not by hiring managers.
Only one of the 13 pairings was judged independently by two judges (agreement about 87%); the rest were judged by a single judge. So I cannot say "two independent judges".
Now, the substance.
First: drop in the whole document, and semantic matching all but fails
I happened to have a natural control: the three résumés mentioned above.
Same person, same employers, same years, same education, same experience. The only difference is how it is written.
Judged requirement by requirement, the three versions could show evidence for 44%, 24% and 27% of requirements. The best is close to twice the worst. That gap is not small; a person can feel it at a glance.
Then I encoded each full résumé as a vector and computed its cosine similarity with the same job description: 0.6584, 0.6551 and 0.6660.
The human judgement saw a gap of 0.21. The vectors saw a gap of 0.011.
Almost twenty times smaller.
But what really stopped me was that the order was reversed. The human judgement ranked B best; the vectors ranked B last, and put the version the human judged weakest in first place.
This is no longer a matter of "not recognising synonyms well enough". It cannot even tell good from bad.
On reflection that is not surprising. A résumé of a thousand-odd words compressed into one vector of a few hundred dimensions: the small differences that come from how it is written get washed out by the overall meaning. Three résumés about the same person's same history ought to look alike in vector space — and they do, so alike that the differences disappear.
Interestingly, with the same data, when I split the résumés into individual entries and encoded each one separately, the separation jumped from 0.011 back to 0.25, close to the human 0.29.
So the conclusion is not "AI can't read résumés". It is that AI can't read a résumé dropped in whole — cut into pieces, it reads them fine.
That distinction matters to job seekers: whether how well you write can be seen at all depends on which kind of system is on the other side. And you have no way of knowing.
Second: what is missing isn't synonyms — it was never written
The piece implies you should switch your wording to the job description's terms so the machine recognises it.
But in my data, for pairings in the right field, more than half of the requirements had no supporting evidence anywhere in the résumé.
I won't give you a precise figure, because no such figure exists. It varies enormously between pairings, from 25% to 82%, averaging about 55% over two rounds. The variation itself is the finding: for the same set of requirements, people differ by a factor of three.
As for what is missing: the tools used day to day are missing about 30% of the time; the scope of responsibility — how many people, what budget, which markets, how far end to end — is also missing about 30% of the time. Soft skills such as communication and collaboration are missing only 7% of the time, the only thing not missing. Because everyone writes them.
I also ran a check that uses no model at all: the share of content words that appear in the job description but nowhere in the résumé, correlated with the share of requirements that have evidence. The correlation was strongly negative, ρ around −0.7 to −0.8.
In other words, the main cause of the gap is not "a different wording the machine didn't recognise". The thing was simply never written down.
(A caveat here: after de-duplicating the three versions from the same person, only 8 samples remain, and the correlation keeps its direction but is not statistically significant. So I consider the direction credible and the strength uncertain.)
Three specific people
Numbers go numb after a while, so here are three specific cases.
First, a licensed financial adviser whose résumé lists no licence at all.
She works as a licensed adviser at a large insurer, and under local regulation she can hardly not hold the relevant qualifications. But the résumé doesn't say a word. Judgement: no evidence.
The same person was matched against four financial sales roles. Under tools she listed Salesforce, Apollo, Notion and Canva — no Excel.
It isn't that she can't use Excel. It is so everyday that it didn't seem worth writing.
Second, seven more years of history, one more requirement answered.
Version J has seven more years of full work history than version C, and more brand-related detail. Out of 34 requirements, it turned over just one more.
The advice to "make your résumé more complete" had almost no effect in this case.
Third, the same experience written as a summary line, and half the evidence gone.
Version B describes specific actions and numbers: which channels, what the campaign data was, the order of magnitude of the spend, how it was tracked and evaluated. In the other two versions the same person compressed all of that into a single line: Led overall brand communication strategy.
Evidence dropped from 15 items to 8. Same experience, same person, just written differently.
Where that piece is right
Parsing matters, layout can break parsing, and structured data is easier to read than free text — I agree with all of that completely.
What it doesn't address is that it assumes the information in the résumé is complete, and the only problem is whether the machine recognises it.
In my data, the biggest loss happens one step earlier. The thing was never written down. However clean the layout and however well chosen the synonyms, what isn't written isn't there.
Incidentally, one sentence in that piece strikes me as the key one: the "intelligence" of different vendors varies enormously.
That sentence effectively admits that any general advice about "how the machine will handle your résumé" has no basis.
Including this piece. So I have written down the method, the data and every point I think is unreliable, for you to judge.
What it means for us as job seekers
These five points are read straight off the gaps above:
Write the names of the tools you actually use every day. The more it feels like "surely that goes without saying", the more you should write it.
Write the scope you are responsible for. How many people, what budget, which markets, how far end to end.
Verb plus object plus magnitude. "What I did" carries far more information than "what I was responsible for".
Delete the words that describe yourself. Self-motivated, Team player, Mature — in my data these earned zero evidence.
One thing per line; don't merge them into summary sentences. The difference between versions B and C above is mostly this.
The method, the original judging rules and the aggregated data are here: \[https://drive.google.com/file/d/1vMhM6RknIic\_udqNzv2Tak-4L5J6-UA1/view?usp=sharing\]