Frustrated job applicants who suspect that no human being has ever looked at their resumé might have a point.
As both employers and job applicants increase their use of AI and automation, researchers and regulators are taking a closer look at these tools. Two key questions: does employment decision technology make the hiring process fairer, by reducing the impact of hiring managers’ personal biases? Or does this technology make the hiring process less fair, by automatically weeding out the same applicants over and over, based upon the tools’ default programming? Or both?
Job seekers want answers. Employers don’t have them.
Fortunately for New Yorkers, the city passed legislation, Local Law 144 (LL144), that requires employers and employment agencies to notify employees and job seekers that they are using “automated employment decision tools.” They are also required to say how they are using them and publish the results of that tool’s “bias audit.” The law went into effect in 2023, but attorneys say, it hasn’t made an impact yet.
The current situation is a difficult one both for the unemployed and the employers. In some cases, technology is not only being used to screen resumés anymore; applicants are being filtered by personality quizzes, online skills tests, and AI video screening interviews with an attractive AI agent… all before the applicant speaks to a human. Recent research indicates that while some of these tools perform very well at AI bias audits, at least some of these tools return results that adversely impact Asian and Black applicants, and might potentially “blackball” individual applicants.
So, what can job seekers do? Many applicants are trying to fight AI with more AI, — but that’s actually part of the problem.
What Hiring Looks Like Today
“We’re not getting 50 applications anymore. We’re getting 500, and a significant chunk of them are AI-generated cover letters attached to resumés that have been keyword-stuffed to beat the ATS.”
—Jacob Wickett, headhunter, founder of Live Digital - Go to Market Recruitment Agency
If you are a job seeker yourself, you might be familiar with applicant tracking systems (ATS). You might even have saved your passwords for Greenhouse, Workday, Workable, SuccessFactors, and BambooHR or others because so many of the jobs you apply to use their software platforms to manage applicants.
An ATS supports human resources teams in many ways, but one of its major functions is serving as the bouncer of the talent acquisition team.
An ATS resumé scanner is the first one to review the candidates and suggest who should make it to the next stage. It doesn’t own the business, so its recommendations of who gets through the door and who waits in the cold can be overruled by the person in charge. But the person in charge is very busy, so if the ATS says you’re not getting in, you’re probably not getting in.
And like any bouncer, the ATS does not respond to human language. That’s why career services experts typically advise job seekers to “ATS optimize” their resume, to make it a machine-readable document. (Professional ATS optimization can easily cost hundreds of dollars, but there are free services through the library.)
Career services experts recommend that job seekers further optimize the resumé for each individual job. So applicants run it through an “ATS resume scanner” like Rezi or Jobscan to get an estimate of how well it scores against the job description.
Things that potentially dock points off a resume are: too few years experience, too many years of experience, having the dates formatted in the less desirable way, using the word “creative” when they were looking for the word “creativity,” or failing to have the precise target job title (Associate Vice-President and Undersecretary of Whatnots, APAC Region) listed among your prior experience.
Tidying up each resume can be a time-consuming process, and frustrating for applicants who put forth the effort, but fail to get interviews.
Many job-seekers have therefore brought more automation into their own process. Some resume scanners now have AI tools built in that will “auto-optimize” the resume.
And then there are extreme automated tools. LazyApply, FastApply, LoopCV, AutoApply, and others that will automate the entire process—searching for jobs, optimizing the resumés/cover letters, and bulk-applying for them without the applicant even knowing what the roles are.
Sometimes 50 jobs a day.
Sometimes a hundred a day.
Sometimes a thousand.
HR’s Dilemma
“It’s a chase. Cat and mouse.”
Tatiana Teppoeva, data scientist and AI hiring strategist
“I’d push back on the idea that ATS tools are the villain here,” says recruitment agent Wickett. “When ‘apply to everything’ became the default job search strategy, recruiters had to start filtering faster. The tools only exist because the problem started existing first.”
Wickett says that the application volume problem has significantly escalated just in the past year, as the auto-apply tools hit the market.
Recessions, tight job markets, and desperation might have contributed to the original “apply to everything” advice in the first place. Whatever chicken or egg started the problem, the challenge both HR teams and job applicants are following—volume—is real.
Tatiana Teppoeva, AI hiring strategist for ONE Nonverbal Ecosystem, explains that the volume of applications isn’t the only problem. Because everyone has optimized their resumés for ATS systems, HR teams cannot distinguish one candidate from the next. Everyone sounds the same. Including some cyber attackers posing as job candidates.
This is why companies are requiring more from applicants—projects, references, cover letters, portfolios, games, tests, video interviews, etc. Teppoeva was a data scientist and researcher for Microsoft and Boeing for 17 years (including eight years working on AI) before founding her own business. She now specializes in helping companies understand the strengths and limitations of their AI video screening tools—what they find and what they miss about candidates.
“Companies are not using these tools just because they’re so cool,” she says. It’s out of necessity.
So now it’s a battle between employers’ AI and applicants’ AI. She likens it to the arms war between cybersecurity professionals and cyber criminals. Applicants try to break through employers’ defenses, and employers respond by adding more layers of defense.
“It’s a chase,” she says. “Cat and mouse. It’s a real problem.”
It Must Be Working…Right?
“Am I happy with the results? Honestly, no.”
Jacob Wickett, recruitment agent
While the technology cannot be blamed for creating all problems, it isn’t solving all of them either.
In December, 36% of respondents to an Express Employment Professionals-Harris Survey said they have open positions they cannot fill, mostly because applicants lack necessary experience. A report last month from the National Federation of Independent Businesses found that 84% of the small businesses who were hiring or trying to hire found few or no qualified applicants.
For some roles, experience is hard to come by, but for others, it might just be hard for the technology to recognize.
“Am I happy with the results? Honestly, no,” says Wickett. “I’ve seen people who would have been great hires screened out because their resume doesn’t tick the right boxes in the right order.”
“The tool isn’t exactly wrong, it’s just not asking the right questions,” says Wickett. “It’s looking for a shape, not a person and some very good people don’t have the right shape on paper.”
What are the right boxes? What’s the right shape? That’s not that easy to say.
Explainability and Transparency
“We don’t know why it did it, and we can’t explain it, and this is the biggest problem.”
—Tatiana Teppoeva
Teppoeva describes how the mystery doesn’t unfold with AI video screening interviews. These interviews might be “asynchronous,” which is a video the interviewee pre-records and uploads, or “synchronous,” which is a live interview between the candidate and an attractive AI agent.
What are these tools measuring? Like a usual screening interview, they are trying to measure “professionalism,” communication skills, enthusiasm for the position, knowledge of the subject matter.
But how?
As they record or analyze a video, they may measure and rate the candidate by the content of their responses, the length of their responses, the background they are using, their attire, their facial expression, how close they are to the camera, etc.?
But…how?
What makes a candidate score poorly? What attire or setting is considered professional? If a person is naturally a bit shy, or has an accent, Teppoeva says, that might impact their score. Or it might not. It depends upon how the AI model was trained.
“From a statistical standpoint,” she says, “all of these models – whether it’s video or it’s resumé screening—they look for patterns. This is how models are trained.
(To “train” an AI model, you first feed it enormous sets of information—called “training data.” The more thorough, and more accurate the training data set, the better the AI will be at “inferring” accurate conclusions later. There’s a concept in software programming called “garbage in, garbage out.” If you fed it a bunch of pictures of dogs and said “that’s a dog” but only showed it pictures of golden retrievers, it might have a harder time later inferring that a Pomeranian is a dog.)
As Teppoeva explains, the tools measure candidates — or whatever they’re measuring — against a “standard deviation curve.”
“If a person fits that curve well, then they score well or whatever the measurement system is incorporated. If person doesn’t fit the curve, it will be an outlier and most likely filter it out. So if model is trained on American population, any accent will be more of an outlier, because it will not be fitting the pattern of the main speaking American population. However if we take a population like me for example, with Russian accents and we train model on this accent, then I will be a majority, I will fit the pattern and an American-speaking person will not fit and they will be the outlier. So it really depends on what they are trying to do.
Tech companies don’t typically make those kind of details clear. They show the answers; but they don’t show their work.
Wickett says, “Transparency depends on the software and some tools are better than others at showing their thought process. But I haven’t used one yet where I felt like I fully understood why it ranked someone the way it did.”
“When companies get results,” says Teppoeva, “they say ‘WOW, this person Tatiana, you know her from previous years, and now she’s trying to do promotion at our company, you know she’s good worker, but the tool screens her and gives her a bad rating.’
“We don’t know why it did it, and we can’t explain it, and this is the biggest problem,” she says. “The results and the output cannot be easily explained.”
Is There Discrimination?
“We find clear racial disparities in applicant outcomes.”
—from ‘Algorithmic Monocultures in Hiring,’ Bommasani, Bana, Creel, Jurafski, and Liang
So, if the technology cannot be explained, how do you know that it’s not discriminatory?
That is the role of AI bias audits.
An example is this one, from AI recruitment company Braintrust AIR, which published the thorough results of an independent AI bias audit of its technology. (This is the type of document that employers should be posting on their website if they use this technology to hire in New York City, as required by local law.)
Braintrust AIR’s audit shows that the tool doesn’t create adverse outcomes by gender, ethnicity, intersectional ethnicity, age, veteran status, or disability. However, Teppoeva points out that bias audits test for anti-discrimination on certain groups protected by law, but they miss subtler things—like shyness. They also don’t provide clarity on how the tool treats other common issues — like employment gaps, job switching, immigration status, a degree from a different country, or whether the applicant used too much boldface text in their resumé document.
A study released in May by researchers from Stanford, Chapman, and Northeastern Universities reported evidence of bias in a gamified recruitment tool called pymetrics, which was developed specifically to be bias-free. Pymetrics scores applicants based upon their performance on a variety of online games.
They had two main findings. First: there was evidence of clear racial disparities that negatively impacted applicants who self-reported as Asian or Black—enough so that it would violate US civil rights legislation.
We find clear racial disparities in applicant outcomes. Of all applications submitted by Asian and Black applicants, 14.74% and 25.87% are submitted to positions that adversely impact Asian and Black applicants, respectively, according to U.S. employment discrimination standards.
Plus, they saw reason for concern about an “algorithmic monoculture”—meaning that while there are thousands of employers, there are only a small number of recruitment tech companies, so whatever a tool’s algorithm decides about a person could impact their application at multiple jobs.
Individuals also receive homogeneous outcomes: 4% of all applicants who apply to 10 positions are recommended for rejection from all positions, a rate higher than expected by chance.
…
Our findings establish empirical evidence to support the claim that shared dependence on a single hiring algorithm vendor yields homogenous outcomes.
In other words, an applicant who is “algorithmically blackballed” as they call it, could be rejected from every job they apply to, and to the best of their knowledge “this pattern of systemic rejection is unique to algorithmic hiring.”
The researchers are clear that they only tested one tool, and they cannot confirm that the same results would apply to other gamified recruitment tools, resume screeners, or AI video interview assessment tools.
However, if these results could be representative of other tools—if these tools can create discriminatory adverse impacts against entire groups of people, and if widespread use of the same tools across thousands of companies can cause certain individuals to be blackballed from the workforce—who’s responsible?
Well, we’re still working that out.
Where the Law Stands Now
Tomorrow I’ll give a run down on a few things:
The status of a discrimination case a job seeker brought against software company Workday,
What NYC Local Law 144 is, and how it can (or should) help workers understand how companies are using automated employment decision laws like these
How workers can pursue anti-discrimination charges outside of LL144
What else workers can do to help navigate the job search, in light of the amount of automation being used today.
Check back for that tomorrow. It’s all exhausting.


