Skip to main content

The hiring-AI problem: bias audits, notices, and the law already watching hiring AI

The short answer

Hiring AI is the most-watched AI you run

A single resume screener can owe United States federal statutes, New York City's Local Law 144, the Illinois Human Rights Act, and the EU AI Act's high-risk regime at the same time. No other system in your estate is under this many spotlights, which is why it earns its own topic and its own record.

What you will be able to do

  • Identify every hiring-related AI system in your organization by pulling from your AI systems inventory (built in (see Topic 0.2)) and naming, for each, exactly which stage of the hiring funnel it touches: sourcing, screening, ranking, assessment, or interview analysis.

The lesson

Job applications fly into the digital ether, and instantly, at 2 in the morning, automated digital rejections fire back. The system clears dozens of candidates before a human recruiter even wakes up. Over 100 of those instantaneous rejections hit a single candidate, Derek Mobley, a Black applicant over 40 with a diagnosed disability.

A shared piece of applicant screening software from Workday connected the companies that turned him away, sitting deep inside thousands of corporate hiring pipelines. In 2023, Mobley sued the vendor directly, and in 2025, a federal judge let that case proceed as a nationwide collective action, drawing roughly 14,000 opt-ins and changing the risk math for every organization using automated screening. The software does not get to hide behind being software.

Because hiring eight-eyed is now subject to scrutiny from city ordinances and state statutes, federal enforcement agencies, and international regulations, it requires concrete mathematical proof of fairness. The court's logic in the Mobley case established a specific theory of agent liability. Delegating a hiring decision to a software algorithm is legally equivalent to delegating it to a human agent.

This legal reality voids the common corporate defense that liability rests solely with the vendor who built the model. When you delegate a hiring function, you inherit the liability of the agent carrying it out. These algorithms are already bound by civil rights statutes, drafted decades before the first neural network was ever trained.

The Age Discrimination and Employment Act of 1967 was recently held to reach modern software screening. The judge noted that reading legislative silence as a lack of protection for algorithm-screened applicants is a dangerous way to interpret existing, historic legislation. Organizations using hiring AI today are managing strict, non-delegable civil rights liabilities.

The law treats the algorithm as an active participant in the hiring decision, not a passive tool. Most professional mapping of hiring AI is too narrow, focusing only on a single box at the end of the process. Regulation reaches the entire hiring funnel.

It covers the tools placing targeted job advertisements, sourcing candidates, screening applications, scoring video interviews, and ranking final recommendations. A single tool can be bound by four different bodies of law simultaneously, depending on where the candidate is located. If you hire in New York City, local law 144 requires an annual independent bias audit of the tool and published metrics on your website.

In Illinois, the Human Rights Act makes discriminatory AI outcomes a civil rights violation, regardless of whether the discrimination was intentional. And if you hire in the European Union, the AI Act classifies recruitment systems as high-risk, triggering mandatory risk management and data governance duties. Effective governance must account for the fact that these systems operate under local, state, federal, and international jurisdictions all at once.

To test for disparate impact, United States law relies on a three-step burden-shifting framework established by the Supreme Court in Griggs v. Duke Power. Step one requires the applicant to prove the algorithm produces a measurably worse pass rate for a protected group. Regulators test this using the four-fifths rule.

Group A has a 40% selection rate. Group B has 24. Divide the lower rate by the higher to find the impact ratio, 0.60. Falling below the 80% threshold proves adverse impact.

This arithmetic is the only thing that matters in discovery. A glossy vendor PDF claiming a model was designed to be unbiased offers zero protection once these numbers are on the table. Once a gap is proven, the burden shifts to the employer.

You must prove the tool is job-related and consistent with business necessity. Algorithmic accuracy is not the same as legal validity. HireVue dropped its facial analysis feature because an audit found it added almost nothing to predicting actual job performance.

It carried adverse impact risk with no validity to defend it. Surviving the legal burden shift requires definitive mathematical proof and validation studies, not marketing claims or good intentions. Many teams assume that scrubbing protected fields like race or age from the model's inputs automatically prevents bias.

In reality, algorithms mathematically reconstruct those traits through correlated variables. This is proxy discrimination. Using data points like a zip code to stand in for race.

Because of these proxies, you must audit the selection outcomes and provide candidates with clear notice that the automated tool is being used before it ever runs. The law also requires meaningful human oversight. A recruiter clicking approve on an AI-generated shortlist without reviewing the auto-declined applications is performing decorative oversight.

True oversight requires a human who can see exactly why the tool rejected a candidate and possesses the authority to reverse that algorithmic decision. Relying on input scrubbing and decorative oversight manufactures strict legal liability, leaving the algorithm as the sole undefended decision maker. There is a massive liability gap between organizations that sound compliant and those that are legally ready for discovery.

Myth 1. We don't use AI. Math. The law covers the whole funnel.

Myth 2. The audit passed. Math. Passes hide intersectional discrimination.

Myth 3. It is accurate. Math. Accuracy is not legal validity.

Myth 4. The vendor set the threshold. Math. You own that decision.

Relying on these myths is exactly how organizations find themselves ambushed by regulators and class action discovery. The objective of AI governance is to preload your legal defense documents before any challenge arrives. Answer three mandatory questions before an algorithm processes an application.

1. Where does output land? Map every screening location. 2. What can it prove about fairness? Demand independent audits. 3. Who signs the decision and can they override? Ensure meaningful human oversight.

At procurement, extract exactly what you need to defend the tool. Force your vendor contracts to guarantee audit rights, supply full validity studies, and require notifications for model updates. If a system is out of compliance today, execute an operational pivot.

Switch the tool from making automated auto declines to creating human-reviewed flags. Governing hiring AI requires a continuous cadence of re-audits and monitoring. Build your evidence annex today.

When the regulator's letter arrives, you should be able to produce mathematical documents, not empty assurances.

The ideas, one by one

The category is the whole funnel, not the resume box

Targeted job ads, candidate sourcing, screening, assessments, and video-interview analysis are all named as employment AI. Map your systems stage by stage or you will govern one box and miss three.

"The vendor built it" is not a defense

Mobley v. Workday established that a hiring-AI vendor can be liable as the employer's agent, and that a company delegating hiring to software is delegating to an agent the same as to a human. Liability is shared; treat vendor hiring AI as your own system with your own audit rights.

A fairness statement is not a bias audit

The document that failed in court says "designed to be unbiased." A real audit reports per-group selection rates, impact ratios, the four-fifths result, the auditor, and the population. Read the numbers, demand the intersectional and full-funnel cuts, and treat a bare "pass" as unaudited.

Proxy discrimination is why inputs are not enough

Illinois banned ZIP code as a stand-in for a protected class because a model can discriminate through correlated variables it never labels as protected. Audit outcomes, not just inputs.

Notice is the cheapest obligation and the most skipped

Name the tool, say what it evaluates, offer a human alternative, deliver it before the tool runs. A few lines of plain text separate a documented good-faith process from an ambush.

Human oversight must be able to override

A recruiter who cannot see why the tool rejected someone and cannot reverse it is decorative, and the tool is really deciding, which is the exact configuration the agent theory targets.

Where you set the bar is a decision you must own

A hiring model outputs a score; a human picks the cutoff, and that choice moves every selection rate and impact ratio. Document who chose the threshold, what a fairer bar would have cost, and why you set it where you did, because the threshold is often the least discriminatory alternative a plaintiff will point to.

A clean audit expires

Applicant pools shift and vendors retrain models, so a passing audit can quietly fail; that is why the law wants an audit within the prior year. Run governance as a cadence with re-audit triggers and continuous monitoring, and keep the record current in your dossier so "we audited it" is always answerable with a recent date.

Disability is the class a standard audit can miss

The ADA adds accommodation duties and screen-out limits, and disability is often invisible to a group-by-group audit, so a real, specific accommodation route in the candidate notice is frequently the primary protection, not an afterthought.

Passing the four-fifths rule is not the finish line

The four-fifths rule is a rule of thumb. Courts also weigh statistical significance, so a small gap in a large applicant pool can still be actionable while a scary ratio in a tiny pool may be noise. Report both the four-fifths result and whether the disparity is significant for your sample; never let one stand in for the other.

Validity is the defense, and "accurate" is not validity

Once adverse impact is shown, the burden shifts to you to prove the tool is job-related and consistent with business necessity, which means a validation study showing it predicts success at this job. The HireVue facial-analysis feature was dropped because it added almost nothing to predicting performance: adverse-impact risk with no validity is indefensible. If your vendor cannot supply a validity study, treat that gap as a live liability now, not paperwork for later.

Know the shape of the case before it is filed

Disparate impact runs in three ordered steps: they show the gap, you show the tool works, they show a fairer tool would have worked too. Governing hiring AI is pre-loading the document that answers each step. Ask, on an ordinary day, "which document answers step two, and do we have it today," and the whole job becomes concrete.

This record persists

The bias-audit-and-notice record you build here is not a one-off; it feeds your conformity file (see Topic 5.6) and your evidence annex (see Topic 10.6), so when the regulator's letter arrives (see Topic 5.7) you produce documents, not assurances.

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 37 of the podcast.

Read the full conversation

You know, usually when we talk about a medical diagnosis, there's this, this expectation of precision. Right. Yeah.

It's very clinical. Exactly. It's like engineering.

You break your arm, you go to the hospital and the x-ray shows that in the jagged white line on the screen. It's right there in front of you. And the doctor just points and says, there it is.

Right. It's a binary state. Totally.

Broken or not broken. It's clean. It's undeniable.

And well, it gives you a clear path forward. And honestly, it's comforting. I mean, human beings, especially those of us who operate in management or, you know, legal compliance, we desperately like things to be visible.

Oh, absolutely. We want to be able to categorize a risk, put it in a neat little box and just apply a known framework to it. But then you step into the world of algorithmic decision making inside a corporation and suddenly that x-ray machine is just, it's fundamentally broken.

Yeah, it really is. You can't just shine a light through a neural network and see where the fracture is. We're looking at a diagnostic landscape that is, well, honestly, it's incredibly murky.

It is the absolute definition of diagnostic muddy waters. And when you are a sharp, busy professional, you know, someone responsible for human resources or technology procurement, or maybe you're the general counsel, that lack of an x-ray isn't just frustrating. It's terrifying.

It is terrifying. Because the stakes aren't just technical glitches. They are massive legal liabilities.

Welcome to the Deep Dive. Okay, let's unpack this because today we are opening the black box of your HR department. We're looking at the single AI use case with the most legal crosshairs pointed at it right now from multiple jurisdictions.

And a reality you just have to face is that hiring AI is the most watched AI you run. And what's fascinating here is that this often surprises leadership teams. Really? Why is that? Well, because if you read the headlines, you'd think the biggest danger is like a generative AI model hallucinating.

Oh, right. Like a chatbot promising a customer a refund it shouldn't have. Exactly.

But we aren't waiting for some futuristic, you know, sci-fi AI legislation to catch up with recruitment technology. The law has been here for years. It's already in the building.

It is actively hunting noncompliant systems today. I mean, a city ordinance in New York has required published quantitative bias audits since 2023. Wow.

Illinois made discriminatory hiring AI illegal on the first day of 2026. And a federal age discrimination statute written all the way back in 1967 was just held by a federal judge to reach a modern hiring algorithm. Without even updating the law.

Without needing a single word of amendment. So the law is already in the room. And to really understand the shockwave this is sending through the corporate world.

We have to look at how this plays out in reality. We have to shatter this very persistent illusion in procurement that if you just buy the software as a service, the vendor holds all the risk. Which brings us to a specific, highly explosive federal case that proves exactly how this technology is being watched.

And frankly, who is actually left holding the bag when things go wrong? Let's talk about Derek Mobley. Because his story really changes the entire risk calculus for anyone buying HR software. So Derek is a black man.

He is over 40. And he has a diagnosed disability. Now he was highly educated, holding degrees in finance and IT.

He applied for over 100 jobs. Over 100? 100 jobs. And he was rejected from every single one of them.

Which is statistically staggering on its own. Right. But it wasn't just the volume of rejections that raised a red flag.

It was the velocity. Sometimes these rejections arrived within an hour of him hitting submit. Sometimes the rejection email arrived in the middle of the night.

You know, vastly faster than any human recruiter could have possibly opened his file, read his resume, evaluated his qualifications, and actually made a determination. Right. Because when you are getting auto-rejected at 3 a.m. for jobs you are fundamentally qualified for, you start to realize you aren't being evaluated by a human being.

Not at all. Now, what connected all these different, completely unrelated companies that turned him away? That's the key. It wasn't a shared bias hiring manager who happened to work for 80 different corporations.

It was a shared piece of software. It was an applicant screening system sold by Workday, Inc. And that specific vendor's platform sits inside the hiring pipelines of thousands of employers globally.

Thousands. So in 2023, Mobley did something that made every governance team in the country freeze in their tracks. He sued.

He sued. He sued. But he didn't just sue the dozens of employers who rejected him.

He aimed higher up the chain. Right. He sued the vendor.

He sued Workday directly. And this case, Mobley v. Workday, stretching from 2023 into 2026, has permanently rewritten the rules of engagement. It truly has.

Judge Rita Lynn, who is sitting in the U.S. District Court for the Northern District of California, looked at the mechanics of how the software operated. And she allowed the case to proceed as a nationwide collective action. Wow.

Under the 1967 Age Discrimination in Employment Act, you know, the ADEA, we are talking about roughly 14,000 plaintiffs who opted into this collective action. 14,000 people. Yes.

And in March 2026, Judge Lynn made a definitive ruling. She completely rejected Workday's motion to dismiss. That is a massive precedent.

It establishes a rule that every executive needs to memorize right now. The vendor built it is not a defense. It absolutely isn't.

And to understand why, we need to dive into a critical legal concept called agent liability. Okay, let's break that down. So under traditional common law, if you delegate a task to an agent, you know, a third party acting on your behalf, and they violate the law while executing that task, you, the principal, retain liability.

Right. The court in the Workday case ruled that delegating a hiring function to an automated screening tool is legally identical to handing it to a human recruiting firm. Okay, let me play the role of a very stressed out procurement officer here, because I can hear them yelling at their steering wheels right now.

Go for it. Wait a second. Signing a software subscription contract is legally the same as hiring a prejudiced human recruiter.

But I have the master service agreement right here. Our vendor contract specifically says they are not liable for final hiring decisions. It has a massive limitation of liability disclaimer in bold text.

Oh, I hear that exact argument in boardrooms all the time, like constantly. There is this massive misconception that you can just contract away civil rights obligations. Right, just bury it in the terms of service.

But a private contract, no matter how strongly worded by your outside counsel, cannot erase an employer's statutory duties to its applicants. It just doesn't work that way. It doesn't.

You cannot sign a piece of paper with a software company that nullifies the Civil Rights Act of 1964. The software in these applicant tracking systems isn't passively recording decisions. Right.

It's not like Microsoft Excel just holding data in cells. Precisely. It is actively participating in the decision-making process.

It parses the resume, extracts the features, weighs them against a model, and then recommends some applicants while silently burying others. It is an active participant. So the algorithm is literally acting as an agent of the employer.

Exactly. The anti-discrimination statutes do not distinguish between delegating a hiring function to a human being and delegating it to lines of code. That is wild.

I mean, think about it. If you hire a third-party contracting firm in another state to source candidates for you, and that firm decides to just throw all the resumes from female candidates in the trash, you, the employer, are going to face the EEOC. You can't just blame the contractor.

Exactly. The exact same logic applies to software. The employer cannot point at the dashboard and say, well, the vendor built the model.

We didn't know how it worked. Right. And the vendor cannot point at the employer and say, we only sold neutral software.

They chose how to use it. The liability creates a blast radius that hits both of them. So if a private contract disclaimer is effectively a useless shield here, what concrete actions does a procurement leader actually take when they're sitting at the desk staring at a renewal contract for an HR platform? You have to shift your leverage.

You move away from contractual disclaimers to evidentiary demands. You have to extract actual mathematical evidence from the vendor before you sign. Okay.

What kind of evidence? Four things. First, you must extract audit rights. You need the contractual ability to independently test the tool with your own third-party auditors.

You can't just accept their internal report. Okay. Independent testing.

Got it. Second, you need validation data. This means the vendor must provide statistical proof that the tool actually predicts job performance, not just that it successfully matches keywords.

Which are two very different things. Completely different. Third, you need proxy testing data to see exactly what hidden features the model might be relying on to make its decisions.

Okay. And finally, you need model update notifications. Because these models aren't static, right? They learn, they update.

Right. They push updates constantly. You need a contractual guarantee that they cannot silently change the algorithm's weights or training data without notifying your team.

Because an update could just introduce bias overnight. Overnight. Absolutely.

If a vendor pushes back and offers you, say, a stronger legal indemnification clause instead of giving you audit rights and validation data, you need to recognize that as a massive red flag. Like, walk away. Yes.

It means that when the federal subpoena inevitably arrives, you will be standing in court defending their black box entirely blind. Wow. Okay, so we have firmly established why the employer is liable.

The agent liability doctrine really locks you in. But to really build a defensible perimeter, we have to expand our understanding of where this liability actually lives inside the organization. That's a crucial next step.

Because I guarantee you, if you walk into a Fortune 500 company and ask the C-suite, do you use AI to hire? A good portion of them will confidently say no. Which is one of the most dangerous blind spots in corporate governance right now. They say, we don't use AI to hire because their mental model of AI is heavily influenced by pop culture.

Oh, like Herminator or something. Right. They picture a generative AI chatbot conducting an interview or some sci-fi robot literally reading printed resumes and making a final offer.

Since they don't have a robot making final offers, they think they're in the clear. But the law doesn't care about the robot making the final offer. The law looks at the mechanics of the entire pipeline.

And this brings us to a fundamental reality of this space. The category is the whole funnel, not the resume box. Yes.

If we break this down to the operational level, we have to map the hiring funnel stage by stage. The legal frameworks recognize five distinct stages where employment AI operates. And each stage carries a totally different flavor of risk.

Let's walk through them. Stage one. First, there is source.

This is the very top of the funnel. It involves algorithmic targeted advertisements and candidate mining from massive online databases like, you know, LinkedIn or specialized job boards. This is the AI deciding who even gets to see the job posting in their feed in the first place.

Right. Based on what the algorithm thinks a quote unquote good candidate looks like. Precisely.

The algorithm is optimizing for engagement who is most likely to click. But if the historical data tells the algorithm that men in their thirties usually click on engineering jobs, it will stop showing the ad to women in their fifties. Oh, wow.

We'll get into why this is so legally insidious in a moment. OK, stage two. The second stage is screen.

This is resume parsing and application filtering is by far the highest volume stage. Right. You have 10,000 applications for 50 warehouse jobs.

The software automatically screens out 9,000 of them based on keywords, tenure or maybe employment gaps. So it's just a massive filter. Exactly.

This is where agent liability and disparate impact exposure really concentrate simply because it's where the sheer highest number of human beings are automatically rejected. The biggest bottleneck. OK, then we hit the third stage.

Assess. Yes. Assess is where things get more behavioral.

This includes cognitive skills tests, personality quizzes and those infamous culture fit evaluations. And culture fit is a phrase that should instantly make any compliance officer's blood run cold. It really should.

Because how does an A.I. define culture fit? Right. How do you mathematically quantify that? You basically look at your existing top performers and train a model to find people who match their treats, their backgrounds, their communication styles. An A.I. trained on culture fit often just mathematically encodes similarity to your current workforce.

Oh, I see. So if your current engineering team is 90 percent white and male, the algorithm will quietly penalize candidates who don't fit that demographic profile. It's just reproducing and locking in whatever demographic imbalances you already have.

Exactly. It's just automating the status quo. OK, see here.

The fourth stage is interview. This isn't just like calendar scheduling. This involves asynchronous video interviews where the A.I. conducts biometric and voice analysis.

Wait, really? This is the software turning on your webcam, tracking your eye movements, measuring how often you smile or scoring the prosody and tone of your voice? Yes. And as you can imagine, this stage carries the absolute sharpest risks regarding disability accommodations and biometric privacy laws. I can only imagine.

Finally, we reach the fifth stage. Decide. This is the ranking and recommendation engine.

It presents a dashboard to the human recruiter with a list of candidates scored from one to one hundred. And the big trap I see leaders fall into here is thinking, well, the software just gave us a ranked list. A human being, our VP of hiring, makes the final choice from the top five candidates.

Therefore, it's not A.I. making the decision. The human is in the loop. But the human is only in the loop at the very end of a deeply automated funnel.

The A.I. decided who made the top five out of 5,000 applicants. The human is just picking from the survivors. The legal harm, the discriminatory screening already happened to the 4,995 people the human never even saw.

Let's go back to the source stage for a second. Those targeted ads. Yeah.

Because there's a really terrifying data illusion that happens here. If a biased ad delivery algorithm decides not to show your logistics management job to older women, those women never apply. They never enter your applicant tracking system.

They don't exist in your pipeline. So at the end of the year, when you run a bias audit on your stage two resume screener, the data will look perfectly clean. It will show that men and women are passing the resume screen at equal rates.

But the audit is a mirage because the denominator, the total pool of applicants you are auditing against was already poisoned upstream by the ad algorithm. You hit the nail on the head. The people who were suppressed are just ghosts in the data.

You can't audit what you can't see. OK. So this entire five stage funnel is an absolute minefield.

And depending on where you are stepping and where your candidates live, different regulatory rule books apply. We aren't just dealing with one federal agency here. No, definitely not.

As of 2026, there are four overlapping legal frameworks currently acting on this funnel simultaneously. Let's break these down, because understanding the differences and how they operate is, well, it's crucial for compliance. The first rule book is the foundation, United States federal statutes.

This includes Title VII of the Civil Rights Act of 1964, the ADEA from 1967, and the Americans with Disabilities Act, the ADA, from 1990. And what's vital to remember here is that these laws are old. They don't have a section titled artificial intelligence.

They don't mention AI algorithms or machine learning at all. And that is exactly what makes them so dangerous to ignore. They are incredibly durable because they are technology agnostic.

Well, that makes sense. They reach any selection procedure based on its effect on a protected class. Right.

As Judge Lynn showed in the Workday case, a plaintiff doesn't need the statute to say the word algorithm to apply it to one. A selection procedure is a selection procedure, whether it's a written test from 1975 or a neural network from 2025. Exactly.

Now, moving from the federal level to the municipal level, we have the second rule book, New York City Local Law 144. This was passed in 2021, and the grace period ended with active enforcement beginning on July 5th, 2023. How exactly does NYC 144 operate? What is the demand of a company? It governs any automated employment decision tool used for a role located in New York City, and it has actual teeth.

It imposes civil penalties of $500 to $1,500 per violation per day. Wait, per violation? And a single violation can be counted as a single candidate processed improperly. So if you screen a thousand candidates over a weekend without compliance, you could be looking at a million dollar fine by Monday.

Potentially, yes. The law requires employers to conduct an annual independent bias audit. Okay.

Crucially, you cannot just keep the results internal. You must publish a summary of the quantitative results on your career's website. Furthermore, you must provide advance notice to the candidates that an automated tool is being used, giving them a chance to request an alternative.

Okay, so that's the city level. But the third rule book, which is arguably moving the fastest and creating the most complex patchwork, is the state level. Specifically, let's look at Illinois HB 3773, which became effective January 1st, 2026.

Right. Illinois represents a significant escalation. Under this law, discriminatory AI in hiring is illegal regardless of the employer's intent.

Intent doesn't matter. It is a strictly outcome-based liability model. If the tool discriminates, you are liable full stop.

You can't just say, well, the vendor told us it was fair. No. And Illinois went a step further by specifically targeting the inputs.

The law explicitly bans using a candidate's ZIP code as a proxy for a protected class. Right. We will delve into proxy discrimination shortly, but Illinois recognized that geographic data is deeply tied to racial demographics in the U.S. and they outlawed its use in algorithmic screening.

And it's not just Illinois, right? The state level is highly volatile. It is incredibly volatile. For instance, Colorado passed a broad AI act, but then repealed and replaced it in 2026 with a much narrower disclosure statute.

California is constantly debating new regulations. If you are a national employer, you cannot just set a policy and forget it. You have to re-verify your operating states every single hiring cycle because the ground is shifting beneath your feet.

And if you operate globally, the complexity multiplies. Yeah. Which brings us to the fourth rule book reaching across the Atlantic.

Yeah. The European Union AI Act, specifically Annex III.4 of the Act. The EU AI Act is a completely different beast.

In this case, Annex III explicitly classifies AI systems used for recruitment, selection, and filtering of candidates as high-risk systems. But high-risk status doesn't mean you can't use it, right? It's not a ban. It is not a ban.

But it is an intense, heavy set of affirmative duties. If you operate a high-risk system, you are subject to stringent requirements regarding risk management systems, data governance, continuous human oversight, detailed technical logging, and you must pass a formal conformity assessment before the tool can even be put on the market. But the timeline on the EU AI Act has been a bit of a moving target.

It has. In 2026, the EU passed what's known as the Digital Omnibus Simplification Package. This package deferred the heavy standalone system obligations, the massive technical conformity assessments to December 2, 2027, largely to give the vendor ecosystem time to catch up.

So a company might think, great, I have until late 2027 to worry about Europe. And that would be a catastrophic misread. Because while the heavy technical requirements were deferred, the Article 4 AI literacy obligation is live right now, as of February 2025.

What does AI literacy actually mean in a legal context? Do all my recruiters need computer science degrees? No, no. But it means that any staff deploying or interacting with these high-risk systems must be legally competent to understand their outputs, their limitations, and their risks. You cannot just hand a complex ranking algorithm to a junior recruiter and say, trust the machine.

You have an affirmative duty to train them on how the algorithm can fail. You know, if we step back and look at these four rule books, the federal statutes, NYC 144, Illinois 3773, and the EU AI Act, there's a massive philosophical difference in how the U.S. and Europe approach governance. It really comes down to the sequence of events.

That is a brilliant way to frame it. The U.S. model is essentially outcome first, litigate after. In the U.S., you are generally free to deploy the technology.

There is no central government body that has to stamp your software before you use it. But if that technology produces a discriminatory outcome, you get hit with plaintiff lawsuits, collective actions, and EEOC investigations after the fact. It's a system based on deterrence through litigation.

Exactly. The EU model, conversely, is duty first. It is a preventative regime.

You have to prove conformity, do the risk assessments, and establish the human oversight documentation before you are legally allowed to deploy the system on a single European citizen. Which puts multinational companies in a very difficult position. Because if your tool is evaluating candidates globally, say, for a remote engineering role where applicants are in Chicago, New York, and Dublin, you owe both answers at once.

You have to prepare the exhaustive documentation and conformity assessments before deployment to satisfy the EU, and you have to be ready to defend the statistical outcomes in a federal court to survive the U.S. legal system. Exactly. You are fighting a war on two fronts with two completely different rulebooks.

And since the U.S. rulebook is litigate after, we need to know exactly how that litigation works. We need to understand the mechanics of a lawsuit so you can prepare the documents you need today, not frantically try to assemble them when a subpoena arrives. Which brings us to the most important legal doctrine in U.S. employment AI, disparate impact.

If you take nothing else away from this deep dive, you must understand how disparate impact works. Let's define this clearly, because it is the engine of almost every major algorithmic hiring lawsuit. Disparate impact occurs when an employer uses a facially neutral selection procedure, meaning a test or an algorithm that doesn't explicitly mention race, gender, or age, but that procedure produces a disproportionately negative outcome for a legally protected group.

And the absolute most critical part of this definition, it does not require intent. That is the cornerstone. You do not have to mean to discriminate to be found liable.

Your programmers could be the most egalitarian, well-intentioned people on earth. But if the mathematical outcome of their model is discriminatory, the tool is illegal, period. The law only looks at the landing point, not the launch pad.

Yes. And the way this is actually litigated in court follows a highly formalized three-step framework. This framework was established by the Supreme Court all the way back in 1971 in a case called Grigs v. Duke Power.

And it was later codified by Congress in the Civil Rights Act of 1991. It is known as the burden shift. The burden shift.

Right. Understanding the shape of this burden shift is the single most useful mental model you can carry out of this discussion. Let's walk through it step by step.

I want you to imagine you were sitting in a federal courtroom. Step one. Step one is the plaintiff's burden.

The plaintiff, let's say it's Derek Mobley's lawyers, must show adverse impact. They don't have to prove how the AI works. They just have to identify the specific selection procedure, which is the AI model, and prove mathematically that it produces a worse pass rate for a protected group compared to another group.

This is typically demonstrated using statistical models, most commonly the four-fifths rule, which we will calculate together shortly. So in step one, the plaintiff just points to the spreadsheet and says, look, Your Honor, this software is advancing 50% of white candidates, but only 20% of black candidates. Here's the math.

We have adverse impact. Once the judge accepts that math, what happens? Then the burden shifts. That's why it's called the burden shift.

We move to step two, which is the employer's burden. Once the plaintiff mathematically demonstrates adverse impact, you, the employer, are on the defensive. You must now prove that the AI tool is job-related and consistent with business necessity.

How do you prove that? You can't just say, well, it helps us hire faster. No, efficiency is not a defense for discrimination. To prove job-relatedness, you must produce a formal validation study.

Validation is a very specific concept in industrial organizational psychology. What does a proper validation study actually look like? If I ask my vendor for one, what am I looking for? Under the EEOC's uniform guidelines, a validation study must take one of three forms. First, there is criterion-related validity.

This requires statistical evidence showing that higher scores on the AI tool directly correlate to higher actual job performance. So if the AI gives someone a 95, you have data showing people who score 95 actually sell more products or write better code than people who score a 70. It actually predicts success.

Yes. Second is content validity. This means the tool samples the actual physical content of the job.

The classic example is a typing test for a typist. If the AI is evaluating a coding simulation for a software engineer, that has high content validity. Makes sense.

Third is construct validity. This means the tool accurately measures a specific psychological trait like reliability or spatial reasoning that is demonstrably necessary to perform the job safely and effectively. Okay, so let's say the employer manages to produce a brilliant criterion-related validation study.

They prove the tool actually predicts who will be a great employee. Do they win the lawsuit right there? Not automatically, because if you carry your burden in step two, the burden shifts back to the plaintiff for the final phase. Step three is the plaintiff's reply.

Okay. The plaintiff can still win the case, even if your tool is valid, if they can show that a less discriminatory alternative existed that served your business needs just as well and you refused to adopt it. Ah, that is a massive trapdoor.

So even if the tool successfully predicts job performance, if there was a fairer way to configure it and you just didn't bother to look, you lose. Exactly. The plaintiff's expert witness will come in and say, yes, the employer's model predicts job success.

But if they had simply adjusted the threshold score by two points, or if they had dropped this one specific proxy feature from the dataset, the algorithm would have been just as predictive, but the adverse impact against minority candidates would have dropped by half. They didn't test for that alternative. That failure to explore alternatives becomes the hammer that wins the case against you.

Let's ground this very abstract legal framework with a highly publicized real-world example. Let's look at the HireVue case, stretching from 2019 to 2021. HireVue was, and is, a widely used vendor for video interview software.

And for a time, they sold a product that scored candidates partly on automated facial analysis. Yes. The software would track microexpressions, how your eyes moved, how your brow furrowed, the tone of your voice.

And they claimed that this biometric analysis could measure complex traits like emotional intelligence or cognitive ability. Which sounds incredibly futuristic, but also deeply unsettling. And in 2019, a privacy rights group called the Electronic Privacy Information Center, or EPIC, filed a formal complaint with the Federal Trade Commission.

They argued this facial analysis scoring was an unfair and deceptive trade practice. HireVue came under immense public and regulatory pressure. So they hired an independent algorithmic auditing firm, O'Neill Risk Consulting or CAA, to conduct a massive audit of the tool.

And in 2021, HireVue did something drastic. They dropped the facial analysis feature entirely. And the reason why they dropped it is the most crucial part of this story.

They didn't just drop it because it was creepy or unpopular. No, they dropped it because of the math. Exactly.

When O'Neill Risk Consulting or CAA audited the tool, they found that the facial analysis data added almost nothing to the model's ability to predict actual job performance. Nothing. Almost nothing.

The biometric data was essentially noise. If we look at this through the lens of the burden shift framework we just discussed, it's a perfect lesson in liability. Walk us through how that would have played out in court.

Sure. Under Step 1, the feature carried obvious adverse impact risks, particularly for candidates with disabilities that affect facial mobility or neurodivergent candidates who make less eye contact. So a plaintiff easily proves adverse impact.

Step 1 is done. Then we move to Step 2. The employer has to prove validity. But the audit showed the facial data didn't actually predict job success.

Therefore, the employer would instantly fail Step 2. They would have zero validity defense. And let me be clear, possessing a tool that causes adverse impact, combined with having no valid business necessity to defend it, is the absolute worst possible legal combination. It is an unwinnable case.

This perfectly illustrates a massive confusion in the market right now. The confusion between accuracy and validity. I see vendors all the time standing at trade show booths, handing out pamphlets that say, our hiring AI is 95% accurate, and procurement officers think that protects them.

95% accuracy is a completely meaningless phrase in a disparate impact court case. Really? Meaningless? It will not save you. Accuracy, in the context of machine learning, often just means the model reliably copies its training data.

If you train a model on 10 years of historical hiring data where human managers consistently underpromoted women, the AI will learn that pattern. And it will execute that pattern with 95% accuracy. Exactly.

A 95% accurate model might just be copying human bias perfectly. It is accurately biased. Validity on the other hand, is a legal and psychological standard.

It proves the tool actually predicts future job success, independent of historical bias. When a plaintiff proves adverse impact in Step 1, your defense in Step 2 is job relatedness, which requires validity. A vendor shouting, the tool is accurate, is entirely irrelevant to the judge.

I like to compare the burden shift to a restaurant health inspection, because it makes the sequence of events very clear. Oh, I like this. Step 1 is the health inspector walking into your kitchen and finding mouse droppings under the stove.

That's the adverse impact. Now, Step 2 is not you throwing your hands up and saying, hey, we didn't put the mice there, we had no bad intent. Because intent doesn't matter to the health inspector, the droppings are there.

Step 2 is you having to prove that you have a rigorous sanitization process that is a business necessity for running a clean kitchen. That is a very precise analogy. And here is where the vendor contract fails.

If you try to pass Step 2 by just waiving a receipt from an exterminator, which is like waiving a vendor contract that says, software provided as is, but you have no actual proof the extermination worked, you fail the inspection. You get shut down. You need the validation study showing the sanitization protocol actually works.

The receipt is not enough. And to finish your analogy, Step 3 would be the inspector pointing out that you could have just sealed the hole in the wall where the mice were getting in a highly effective, less discriminatory alternative, but you didn't bother to look for it. Exactly.

Okay, so to survive Step 1 of this burden shift, to even know if you have mouse droppings in your HR data, you need to conduct an audit. But what software vendors call an audit is usually a trap for the unwary, which brings us to a harsh truth of procurement. A fairness statement is not a bias audit.

This is a critical distinction that trips up even sophisticated legal teams. A vendor might hand you a beautifully designed, glossy one-page PDF. It has their logo.

It's titled Fairness Statement, and it includes paragraphs saying the model was designed to be unbiased, built on ethical AI principles, and tested for equity. It usually features a very reassuring large green checkmark. Looks great on a desk.

That document is a marketing statement. It is legally weightless. It will collapse instantly under the first question and cross-examination.

So what does a real bias audit actually contain? If a glossy PDF is fake, what does the real thing look like? A real bias audit isn't pretty. It is a dense table of numbers. It must explicitly state the exact selection rate for every single protected group.

It must show the mathematical impact ratio comparing those groups. It must show the final four-fifths rule result. It must state the name and the institutional independence of the auditor who ran the test.

It must provide the date of the data snapshot, and it must define the total population size that was audited. If it doesn't have those raw numbers, it's not an audit. Let's actually do the math on the four-fifths rule right now, because this rule comes from the 1978 EEOC Uniform Guidelines.

It is the bedrock of compliance. And everyone listening to this needs to know how to calculate it on the back of a napkin. Gladly.

It is simple arithmetic, but it's essential. First, you calculate the selection rate for each group independently. Let's say your new AI resume screener evaluates 1,000 male applicants and advances 400 of them.

That's a 40% selection rate for men. Now let's say it evaluates 1,000 female applicants and advances 240 of them. Okay, so we have a 40% pass rate for men and a 24% pass rate for women.

Right. You always take the highest selected group's pass rate, in this case, the men, at 40%. That becomes your benchmark.

Then you calculate the impact ratio for the lower-scoring group by dividing their rate by the benchmark rate. So you divide 24 by 40. And 24 divided by 40 is .60? Correct.

.60, or 60%. The federal rule of thumb states that any group passing at less than 80%, which is four-fifths, or .80 of the highest group's rate, is flagged for having adverse impact. Since .60 is well below .80, this AI tool fails the test dramatically.

So if a vendor hands me a 40-page audit report that buries an impact ratio of .60 on page 38, but they slapped a big green pass-tested-for-equity stamp on the cover page? The green stamp is a lie. The arithmetic controls the legal reality, not the marketing stamp. You have just been handed written, documented evidence of your own company's adverse impact.

And if you file that away without acting on it, you have just documented your own negligence. Wow. Okay, so how do we learn to read these audits like a regulator? Because vendors know how to play with the numbers.

What are the specific tricks we need to spot when we are reviewing these tables? There are three main statistical tricks that vendors use to artificially inflate their fairness scores. Okay, hit me. The first is small or missing categories.

If a specific demographic group has very few applicants, say, Native American women, a vendor's auditor might just drop that row from the final table entirely, claiming they need to avoid computing an unstable statistical weight. But the result is that a tool can look like it passes every recorded category while quietly excluding the very minority group where the algorithm is causing the most harm. So you have to read for what's absent.

You have to ask, where do these 50 applicants go? Exactly. You cannot just read what is on the page. The second trick is intersectional blindness.

The four-fifths rule is typically run one dimension at a time. You run race alone, and then you run sex alone. An AI tool might pass perfectly for race overall, and it might pass for sex overall.

But if you combine those data points, it might be failing disastrously for a specific intersection, like older black women. Oh, because it gets masked by the broader averages. Precisely.

If the auditor never computes that combined intersectional rate, the failure remains completely hidden. You must demand the intersectional cuts in the data. And the third trick.

The third is the wrong denominator. This relates directly back to our funnel mapping earlier. The selection rate percentage depends entirely on who is counted as an applicant.

If your algorithmic targeted ad tool suppressed a specific demographic at the source stage, those people never applied. They never entered the denominator for the screening audit. So your screening tool looks perfectly fair, an impact ratio of 0.95, but only because the discrimination already happened upstream.

The denominator was poisoned from day one. Yes. And this is where we have to go one level deeper into the math, because passing the four-fifths rule with a bill 0.81 is a floor, not a ceiling.

It is just an administrative rule of thumb. When you get into federal court, judges also look at statistical significance. What does that mean in actual practice during a lawsuit? Let's cite the two bedrock Supreme Court precedents, Castaneda v. Partita and Hazelwood School District v. United States, both from 1977.

Federal courts have long treated a disparity of roughly two to three standard deviations as the point where a difference becomes statistically significant, meaning the gap is so large it is highly unlikely to be just random chance. OK, so how does standard deviation interact with the four-fifths rule? Because this sounds like conflicting math. It's the law of large numbers, and it cuts both ways.

Let's say you are a massive national retailer. You have a pool of 60,000 applicants for holiday seasonal work. The selection rate gap between men and women might be very small.

It might result in an impact ratio of 0.85, which easily passes the four-fifths rule. You think you are safe. But because the sample size of 60,000 is so massive, that tiny percentage gap might represent a difference of three or four standard deviations away from the expected norm.

It is mathematically statistically significant and therefore legally actionable in court, even though it passed the 0.80 floor. Because with 60,000 people, any deviation from the mean is clearly a systemic pattern, not a fluke. So a bare four-fifths pass is not a clean bill of health.

Not at high volumes, no. And conversely, it works in reverse. If your applicant pool for a highly specialized executive role is tiny, 15 people, an impact ratio of 0.50 looks terrifying on paper, it fails the four-fifths rule instantly.

But it might not be statistically significant at all because the sample size is just too small to draw a mathematical conclusion. One person withdrawing their application skews the entire table. In those cases, sophisticated practitioners use advanced models like Fisher's exact test for small samples.

So what's the takeaway here for an executive? The overarching point is this. You must ask your vendor for both the four-fifths result and the statistical significance test. Never let one substitute for the other.

So if my vendor pushes back and just gives me a beautiful PDF with a green checkmark that says tested for equity and refuses to provide standard deviations, sample sizes, or intersectional cuts, I effectively have nothing. You have a marketing statement that will collapse in court. And worse, you have exposed your organization to massive liability because you documented that you accepted a vendor's PDF as a substitute for real compliance.

Let's talk about how these passing or failing numbers actually get generated. Because it's not just the AI acting entirely on its own in a vacuum. The AI outputs a score, let's say 1 to 100.

It doesn't output a definitive yes or no or reject them. The human being in HR sets the cutoff threshold. The human says only candidates scoring 85 or above will advance to the interview stage.

That is a vital point that changes everything. The threshold choice is a human governance decision, not a technical default. Moving that single bar changes exactly who passes and who fails, which immediately changes every selection rate and every impact ratio in your entire audit.

So if your audit shows you failing the four-fifths rule with a 0.70 impact ratio to a threshold of 85, if you lower the bar from 85 to 80, more people from the lower scoring protected group might clear it, and your impact ratio might suddenly jump up to a passing 0.85. Precisely. And this is why the threshold setting is exactly the less discriminatory alternative that plaintiff's lawyers will zero in on during step three of the burden shift. If you could have lowered the threshold slightly, dramatically improved the fairness of the outcome, and still met your business needs for qualified candidates, but you didn't bother to test that lower threshold, you lose the case.

Because you ignored a less discriminatory alternative. You must document exactly who chose the bar, why they chose that specific number, and what a fairer bar would have cost in terms of candidate quality. Okay, let's unpack a completely different kind of trap.

Let's say an organization does everything right. You have a real bias audit. The numbers look good.

The threshold is carefully documented and defended. Are you safe? Not necessarily. Right, because of the inputs.

There is a defense I hear constantly from tech companies. Well, we scrubbed all the resumes. We deleted the names.

We deleted race. We deleted gender from the data. The AI is entirely blind to demographics, so it can't possibly be biased.

That defense is completely bankrupt mathematically, and courts see right through it. Proxy discrimination is why inputs are not enough. Proxy discrimination.

We mentioned this briefly with the Illinois law. What is it, mechanically? Proxy discrimination occurs when a complex machine learning model reconstructs a protected demographic trait using other seemingly neutral variables that it is allowed to see. It does this because those neutral variables are highly correlated with the protected trait in the real world.

Give me a concrete example of how an AI does that. The most famous and historically damaging example is ZIP codes, where a person lives in the United States is deeply, highly correlated with race due to decades of historical redlining and residential segregation. Right.

If you strip the word race from the resume, but the AI is allowed to look at the candidate's ZIP code to calculate commute times, the AI will quickly learn to penalize candidates from certain ZIP codes because they historically correlate with lower historical hiring rates in your biased training data. The algorithm is effectively discriminating by race, even though it was explicitly blind to the race category. It just found a backdoor.

And this isn't just theory. As we discussed, the law is acting on it. Illinois HB 3773 explicitly bans using ZIP codes as a proxy for a protected class for exactly this reason.

Yes. But the ZIP code is just the most obvious, well-known example. In a high dimensional neural network, the AI can weave together dozens of faint signals.

It can use the specific high school a candidate attended, the length of their employment gaps, their hobbies, or even specific syntactic phrasing patterns in how they write their cover letter. Just tiny little clues. Exactly.

It weaves these data points together to infer class, race, or age with terrifying accuracy. This completely obliterates the idea of fairness through unawareness. The idea that if the model is unaware of race, it must be fair.

Exactly. Fairness through unawareness is a myth. The only way to prove a model is fair is to audit the outcomes, the actual selection rates and impact ratios of the human beings exiting the funnel.

You cannot successfully police every single input a neural network uses to infer a demographic. You have to measure the output. But wait, if we have to audit the outcomes, what does this mean for groups of people that don't neatly self-identify on an EEOC form? Which brings us to a massive blind spot that almost no one is talking about.

The disability angle. The Americans with Disabilities Act, the ADA, is incredibly powerful in this space, perhaps more so than Title VII right now. The ADA strictly limits screening out candidates based on disability and requires employers to provide reasonable accommodations during the hiring process.

Right. But the unique challenge with algorithmic hiring is that disability is often completely invisible, and candidates understandably do not self-identify on their initial application due to fear of stigma. So if they don't check a box saying, I have a disability, they aren't tracked in your vendor's beautiful bias audit.

Correct. Your standard group-by-group audit will completely miss them. It tests for race and gender, but not for neurodivergence or medical history.

For example, a video interview AI might score a candidate poorly for having a flat affect or a lack of steady eye contact. Which heavily penalizes autistic candidates, even though eye contact has nothing to do with writing software. Yes.

Or an AI resume screening tool might heavily penalize a two-year gap in a candidate's employment history, ranking them in the bottom percentile. Which aggressively penalizes someone who took medical leave for cancer treatment or someone dealing with long COVID. Precisely.

The vendor's audit looks completely clean because it isn't specifically measuring for autism or cancer survivors. But the tool is actively, systematically penalizing traits and data points that are direct markers of a disability, rather than actual necessary job requirements. But wait, if we can't audit for an invisible disability because they never self-identified, how do we actually protect them? We can't just leave a massive legal and ethical blind spot.

Because you cannot perfectly audit for every intersectional class, and you cannot audit for an invisible disability, you must provide a procedural safety valve. And that safety valve is candidate notice. Notice is the cheapest obligation and the most skipped.

It is the primary legal protection for the ADA angle. Let's contrast a hollow notice with a truly defensible notice. Because I have seen the hollow ones, they are usually buried on page nine of a privacy policy hyperlink at the very bottom of the careers page.

And they say something vague like, by submitting your application, you consent to our use of technology and data analytics in our recruitment process. That is entirely hollow. It names no specific tool.

It says absolutely nothing about what traits or skills are being evaluated. It offers no alternative path. And realistically, a candidate will never see it before applying.

So what makes it defensible? A defensible notice one that will actually protect you under the ADA and NYC 144 does four specific things. One, it explicitly names the specific AI tool being used. Two, it states in plain non-technical words what qualifications or characteristics that tool evaluates.

Three, it offers a real accessible route to request a human alternative or an accommodation. And four, it is delivered before the tool runs. That timing is crucial.

Not hidden in the rejection email. We are sorry, an AI decided you are not a good fit. That's not a notice, that's an autopsy.

Exactly, it must be visible before the candidate hits the submit button. A defensible notice reads simply like this. We use an automated screening tool called Talent Sift to review your resume against this job's required technical skills.

If you would prefer a human recruiter to review your application first, or if you need an accommodation due to a disability, email us at this specific address within five business days. That is maybe three sentences. It costs $0 in software development to implement.

It is the least expensive, highest ROI compliance step we will discuss today. Which is why it is so frustrating that employers skip it. People think, well, New York City's local law 144 requires notice.

But I read that a 2025 New York State controller audit found that enforcement by the city was weak. Only a fraction of companies were complying. So we can just skip it and blend in with the crowd.

Just because the city agency is understaffed and busy doesn't mean the plaintiff's lawyer won't find it. A missing notice is an absolute smoking gun in civil discovery. It shows a jury that you didn't even try to provide a good faith process for disabled applicants.

And that leads directly to the concept of meaningful human oversight. The notice offers a human alternative, but that human has to actually be able to do something. Right, if the human alternative is just a tired recruiter who looks at the AI's rejection pile and says, well, the computer said no, so I agree.

That's not oversight. That is decorative oversight. A human must be able to see exactly why the tool rejected someone and have the actual organizational authority to override the algorithm.

If an AI auto declines 500 applicants, and a recruiter only ever logs in to look at the top 20 ranked survivors and rubber stamps them for interviews, the tool is the one making the employment decisions. The human is just decorative. And that is exactly the configuration that triggers agent liability under the workday precedent.

Okay, we have covered an immense amount of theoretical and legal ground. We've covered the laws, the disparate impact math, the proxy traps, and the notices. Now, I wanna bring this directly to the desk of the listener.

Let's look at exactly how a sharp professional handles this reality on a Tuesday morning. Let's dive into the immersive scenario. Let's imagine Kyle.

Kyle is the newly appointed AI governance lead at Northwind Freight, which is a mid-sized logistics and warehousing company. They hire hundreds of warehouse staff every month across three major sites, New York City, Chicago, and a European hub in Dublin. Okay, so Kyle sits down on a Tuesday morning and pulls up his AI systems inventory.

He sees a vendor tool called TalentSift that is handling all warehouse applications across all three locations. The software is currently considered to automatically decline the bottom 50% of resumes based on a keyword matching algorithm. How does Kyle handle this using an expert mental model? Kyle needs to silence the noise and ask three specific procedural questions.

Question one, where does this tool's output land geographically? This determines which rule books he is fighting. Let's map it. New York City means local law 144 applies.

He needs to publish bias audit and the notice. Chicago means the Illinois Human Rights Act applies. He needs to ensure no zip codes or proxies are being used.

And he faces strict outcome liability. Dublin means the EU AI Act applies. He needs Article 4 AI literacy training for his recruiters right now and conformity assessments by 2027.

One tool, three totally different rule books. Correct, the geography dictates the compliance roadmap, so he moves to question two. What can this tool prove about its own fairness right now? Kyle opens the vendor compliance folder on his shared drive and finds a glossy PDF titled Fairness Statement.

It says the model was tested for equity, but there are no selection rates, no impact ratios, no sample sizes, no mention of standard deviations. Kyle's holding a marketing brochure. He has no legal defense.

Yes. So Kyle gets on the phone with the vendor's account rep. The rep tries to dodge saying their algorithm is proprietary and they can't share the underlying data.

Kyle knows his defensible move. He officially demands the raw impact ratios in writing within 30 days, citing their contract's audit rights clause. If they refuse, he prepares to pause the tool.

Then he asks question three. Who signs the human decision and can they override it? So Kyle walks down the hall and talks to Sarah, the lead warehouse recruiter. He asks her how she handles the talent shift output.

Sarah admits she basically never looks at the auto declines because there are hundreds of them, and the VP of Logistics is screaming at her to just get bodies on the forklift by Monday. If she thought the AI made a mistake, she wouldn't know why or how to reverse it in the system. This is a classic corporate friction point.

The business side wants speed. The legal side needs safety. Sarah is providing decorative oversight.

So Kyle makes a highly defensible operational move. He drafts a memo switching the system's configuration from auto decline to a human-reviewed auto flag. How does that change the dynamic? Instead of the AI sending a rejection email at 3 AM, the AI flags the bottom 60% as requires review.

Sarah now has to physically log in, glance at the flagpile, and click a button to approve the rejections. It adds maybe 30 minutes to her week, but legally, a human is now formally signing off on the rejections before their final. Kyle is establishing a paper trail of human oversight.

He isn't panicking and ripping the software out of the wall, which would make the VP of Logistics furious. He is fixing the governance structure around the software. Exactly.

But Kyle must also remember one final, crucial reality. Audits drift. Even if the vendor finally sends Kyle a perfect, mathematically sound bias audit today, it isn't permanent.

Why not? If the algorithm's code hasn't changed, shouldn't the fairness remain the same? Because the world changes and the data changes. The applicant pool shifts based on the season or the macroeconomic job market. In November, Northwind Freight might get a massive influx of applicants from a different demographic group due to a local factory closing.

Furthermore, the vendor might silently update the model's neural weights in the background to improve efficiency. A passing 0.85 impact ratio from January can quietly fail with a 0.60 by autumn. The governance of hiring AI is not a one-time gate you pass through.

It is a continuous cadence. So you have to stay on top of it. It requires annual re-audits, triggers for off-cycle audits if the vendor pushes a major update, and continuous dashboard monitoring of those impact ratios.

It requires vigilance. Okay, we have covered an immense amount of ground today. We've looked at the workday precedent, the five stages of the funnel, the four legal rule books, the math of disparate impact, proxy traps, and the sheer necessity of candidate notice.

Let's bring this all home with the Monday morning move, the single most valuable concrete action you should take when you log in on Monday. Do exactly what Kyle did. Pull your AI system's inventory and aggressively map your hiring funnel.

Don't just look at the final interview. Look at all five stages. Source, screen, assess, interview, and decide.

Identify exactly which of your automated tools touch each of those stages. And for every single tool, check your files. Do you have a real quantitative bias audit with actual impact ratios and a validation study proving job relatedness? Or are you holding a vendor's marketing PDF? Check your careers page.

Do you have a defensible candidate notice that names the tool and offers an accommodation? Don't wait for a federal subpoena or an EEOC letter to find out you're holding a marketing brochure instead of legal defense. I want to leave you with a final thought to mull over. Something we haven't touched on yet, but that is rapidly approaching on the horizon.

The goal of AI governance isn't to reach zero AI. Hiring AI genuinely helps at scale. It cuts a mountain of applications into a workable short list.

The goal is to build a paper trail you would hand a regulator or a plaintiff's lawyer without flinching. But consider what happens next. What do you mean? We've focused entirely on AI rejecting candidates today.

But consider the liability when AI starts actively negotiating salaries based on historical market data. If an algorithm realizes it can consistently secure female candidates or minority candidates for 10% less pay because of historical market inequities, and it automatically lowers their offer letters to save the company money. Your next massive class action liability won't just be about who you hired.

It will be about how much the algorithm secretly decided they were worth. That is a terrifying frontier for equal pay legislation, because the day is absolutely coming when that black box will be open in civil discovery. Exactly, and when that day comes, the question the judge asks won't be whether you used an algorithm.

The question will be whether you can mathematically prove you were the one driving it, or if you were just along for the ride while it broke the law. To bring it all back to where we started, when you finally get handed the x-ray of your HR hiring process, you don't want to be surprised by the fractures you see on the film. You want to be the executive who built the machine, calibrated it, and knows exactly how to read the results.

Well said. That is our deep dive for today. Thanks for joining us as we unpack the complex high stakes realities of hiring AI.

We'll see you next time.

Real cases

Example 1: Mobley v. Workday (United States, 2023 to present), the agent theory made real. This is the anchor. Derek Mobley, an applicant who is Black, over forty, and has a disability, applied to more than one hundred jobs whose employers used Workday's applicant-screening platform and was rejected every time, sometimes within an hour. His suit targeted the vendor, not only the employers. Judge Rita Lin of the Northern District of California held that Workday's software could be liable as the employers' agent because it participates in hiring decisions rather than passively recording them, and in May 2025 she certified a nationwide ADEA collective action reaching applicants forty and older screened out since 24 September 2020, drawing roughly fourteen thousand opt-ins. In March 2026 she denied Workday's motion to dismiss. The case is unresolved on the merits (emerging), but the doctrine, that a hiring-AI vendor can be an agent and that a 1967 statute reaches a modern algorithm, is established on this record. Source: Holland & Knight, "Federal Court Allows Collective Action Lawsuit Over Alleged AI Hiring Bias" (2025); Seyfarth Shaw, "Mobley v. Workday: Court Holds AI Service Providers Could Be Directly Liable" (2024).

Example 2: New York City Local Law 144 (United States, effective 2023), the first hiring-AI-specific rule. New York City became the first United States jurisdiction to require, by law, an independent bias audit of hiring algorithms, a published summary of the selection rates and impact ratios, and advance notice to candidates. Enforcement began 5 July 2023. The instructive part is what happened next: a New York State Comptroller audit in 2025 concluded that enforcement had been thin and compliance uneven. The lesson for your organization is not "the rule has no teeth"; it is that a live legal obligation exists whether or not the city is aggressively policing it, and a plaintiff's lawyer or a European regulator will read your missing audit very differently from a busy city agency. Source: NYC Department of Consumer and Worker Protection, Local Law 144 pages; Office of the New York State Comptroller, "Enforcement of Local Law 144" (2025).

Example 3: Illinois House Bill 3773 (United States, effective 1 January 2026), discrimination without intent, and the ZIP-code trap. Illinois amended its Human Rights Act to make discriminatory use of AI in employment a civil-rights violation regardless of intent, and it specifically banned using a ZIP code as a proxy for a protected class. That ZIP-code clause is a tell that legislators now understand proxy discrimination: a model that never sees race can still discriminate by race if it leans on a variable tightly correlated with it. The unsettled notice rules (temporarily withdrawn in 2026) show something else worth teaching: even where the operational detail is still being written, the underlying prohibition is live, and "the rules were not final" is not a defense to using a discriminatory tool. Source: Illinois General Assembly, HB 3773; Seyfarth Shaw and Ogletree Deakins analyses (2024 to 2026).

Example 4: The EU AI Act's Annex III, point 4 (European Union), hiring AI named as high-risk by category. The Act does not wait to see whether a given screener discriminates; it classifies recruitment and selection AI as high-risk by its function (Regulation (EU) 2024/1689, Annex III, point 4). That is a different regulatory philosophy from the United States' outcome-first, litigate-after model: the EU front-loads the duty. The high-risk obligations for stand-alone Annex III systems were deferred to 2 December 2027 by the 2026 Digital Omnibus, so an organization has a real window to prepare, but the classification itself is settled and the direction is clear. Source: EUR-Lex, Regulation (EU) 2024/1689; Council of the EU, Digital Omnibus adoption (29 June 2026); European Commission draft guidelines on high-risk AI in employment (2026).

Example 5: HireVue drops facial analysis (United States, 2019 to 2021), the interview stage, and validity as the real test. HireVue, a widely used video-interview vendor, sold a product that scored candidates partly on facial expressions, and it marketed claims about measuring traits like "cognitive ability" and "emotional intelligence" from a recorded interview. In November 2019 the Electronic Privacy Information Center (EPIC) filed a complaint with the United States Federal Trade Commission arguing the practice was unfair and deceptive. HireVue commissioned an independent algorithmic audit from O'Neil Risk Consulting and Algorithmic Auditing (ORCAA), and in January 2021 it announced it would stop using facial analysis in its assessments. The detail governance teams should carry is why it dropped the feature: reporting indicated the facial-analysis component added only a tiny sliver to predicting job performance. Read against Section 3H, that is the whole lesson in one fact. A feature carrying obvious adverse-impact and privacy risk that also cannot show it predicts job success is indefensible: it fails the job-relatedness step with nothing to trade for the risk. Note the honest boundary of the story: HireVue kept analyzing other signals such as speech and intonation, so "dropped facial analysis" is not "solved bias." The teachable point is the test, not a clean ending: a hiring-AI feature that cannot prove validity is a feature you cannot defend. Source: EPIC, "HireVue, Facing FTC Complaint From EPIC, Halts Use of Facial Recognition" (2021); Fortune, "HireVue stops using facial expressions to assess job candidates amid audit of its A.I. algorithms" (19 January 2021).

A note on globally varied practice. These five examples span a city, a state, a federal system, a supranational regulation, and a vendor forced to change by public and regulatory pressure, and they represent two governance philosophies: the United States tends to prohibit outcomes and enforce after the fact through litigation and agencies, while the European Union front-loads duties on a named high-risk category before deployment. The practical contrast is worth holding as a pair of questions each philosophy asks of your tool:

  • The United States, outcome-first, asks after the fact: did this tool produce a worse outcome for a protected group, and if so, can you prove it was job-related and that no fairer alternative existed? You may deploy first, but you carry the burden the day a claim lands.
  • The European Union, duty-first, asks before deployment: is this tool in a named high-risk category, and have you completed the risk management, documentation, oversight, and conformity assessment the category demands? The duty attaches to the category, before any harm is shown.

A tool used in both worlds owes both answers at once, which is why your record must satisfy the stricter reading of each. Other jurisdictions are moving too, and this layer changes quickly; the discipline is to re-verify your own operating jurisdictions each hiring cycle rather than assume the map you learned last year still holds (see Topic 6.1 for the wider United States mosaic and other regions).

Where people go wrong

Mistake 1: "We do not use AI to hire." Most organizations that say this are wrong. They picture a robot reading resumes and miss the targeted-ad tool that decides who sees the job, the applicant-tracking system that auto-ranks, and the video-interview product that scores candidates. The fix: map the whole funnel from Section 3B, stage by stage, against your systems inventory, before you answer the question.

Mistake 2: "The vendor built the model, so the liability is the vendor's." Mobley v. Workday cuts this off in both directions. The employer delegated a hiring function and carries the liability of that delegation; the vendor's software participates in the decision and can be pulled in as an agent. "It is the vendor's model" is not a defense; it is a description of shared exposure. The fix: treat vendor hiring AI as your system, with your audit rights and your documentation.

Mistake 3: "The model never sees race, so it cannot discriminate." Proxy discrimination is the whole reason Illinois banned ZIP code as a stand-in for a protected class. A model can reproduce a protected trait through correlated variables it is allowed to see. Facial neutrality of inputs does not equal fairness of outcomes. The fix: audit outcomes (selection rates and impact ratios), not just inputs.

Mistake 4: "The audit passed, so we are fine." An audit that reports a checkmark can be hiding a failing impact ratio, an omitted small group, or an intersectional failure it never computed. A "pass" on race and a "pass" on sex can coexist with a severe failure for older women. The fix: read the actual numbers, demand the intersectional and full-funnel cuts, and treat a bare "pass" as unaudited.

Mistake 5: "New York City barely enforces Local Law 144, so we can skip the audit." Weak city enforcement is a real finding, but it protects you only from the city. It does nothing against a private plaintiff, a class action, or an EU regulator, all of whom will read your missing audit as a smoking gun. The fix: comply to the standard, not to the current enforcement temperature.

Mistake 6: "Human oversight means a person clicks approve." A recruiter who rubber-stamps a ranked list without being able to see or reverse the tool's rejections is not oversight; the tool is deciding. That is the configuration the agent theory targets. The fix: give the human the ability to see why the tool decided and the authority to override it, and log when they do.

Mistake 7: "Notice is a legal formality we can put in the privacy policy." A buried, vague notice is not a notice. The rules want a specific tool named, in plain words, before it runs, with a real alternative offered. The fix: post a plain, specific notice in advance; it is the cheapest obligation you have and the easiest to prove.

Mistake 8: "High-risk under the EU Act means it is banned, so we must stop." High-risk is not prohibited; it is heavily governed, and the stand-alone obligations are deferred to December 2027. Panic-pulling a useful tool can create its own harm. The fix: govern it to the high-risk standard on a timeline, do not confuse a duty with a ban (see Topic 5.3).

Mistake 9: "It passes the four-fifths rule, so it is legally clean." The four-fifths rule is a rule of thumb, not the ceiling of the law. With a large applicant pool, a gap too small to fail the four-fifths rule can still be statistically significant and actionable; courts look at standard deviations, not only the eighty-percent line. The fix: report both the four-fifths result and whether the disparity is statistically significant for your sample size, and never treat a four-fifths pass as the end of the analysis.

Mistake 10: "The tool is accurate, so it is defensible." Accuracy and validity are not the same thing, and validity is what the law asks for. A model can be internally accurate at reproducing whatever it was trained on, including past biased decisions, while having no evidence that it predicts actual job success. When adverse impact appears, the defense is job-relatedness (criterion, content, or construct validity), and "it is accurate" is not that evidence. The fix: hold a validation study that shows the tool predicts performance on this job, and if the vendor cannot supply one, treat that gap as a live liability, not a formality.

Mistake 11: "We can use the vendor's own bias audit to satisfy our obligation." New York City's Local Law 144 requires an independent bias audit, and "independent" is doing real work in that sentence: an audit the vendor commissioned of its own general product, on its own applicant pool, is not automatically independent of your specific use, your specific applicant pool, and your specific configuration of the tool (your threshold, your feature set). A vendor's marketing-adjacent audit of "the product" is a different document from an audit of the tool as you actually deploy it. The fix: confirm the audit was performed by an auditor independent of the vendor, on data that reflects your own applicant pool and configuration, not just accept whatever PDF the vendor already had on file.

Questions people ask

What is automated employment decision tool (AEDT)?
The term New York City Local Law 144 uses for a computational tool that substantially assists or replaces a hiring or promotion decision. If a tool meets this definition for a New York City role, it triggers the bias-audit, publication, and notice duties. More on Automated employment decision tool (AEDT)
What is adverse impact (disparate impact)?
A selection procedure that is neutral on its face but produces a worse outcome for a protected group. Under United States federal law it can be unlawful even without intent to discriminate, unless the employer shows the procedure is job-related and consistent with business necessity.
What is agent theory (agent liability)?
The legal idea, accepted in Mobley v. Workday, that a hiring-AI vendor can be liable under anti-discrimination statutes as the employer's agent, because the software carries out an employment function rather than passively recording the employer's own decision.
What is Annex III, point 4 (EU AI Act)?
The part of Regulation (EU) 2024/1689 that classifies AI used for recruitment, selection, and employment-relationship decisions (promotion, termination, task allocation, performance monitoring) as high-risk.
What is article 6(3) exemption (EU AI Act)?
A narrow carve-out that can drop an Annex III system out of high-risk if it performs only a narrow procedural task and does not materially influence the outcome. It never applies to a system that profiles people, which a candidate screener does by design.

Keep going