Skip to main content

Risk classification: which of your organization's systems is high-risk and proving why

The short answer

Classification is the hinge of the whole Act

The tier a system lands in decides everything after: banned, heavily regulated, transparency-only, or largely free. This topic sorts every system you run into a tier and proves each sort, because every later obligation attaches to the right systems only if the classification is right.

What you will be able to do

  • Distinguish the four risk tiers of the EU AI Act (prohibited, high-risk, limited or transparency risk, and minimal risk) and place a given AI system in the correct tier with a reason.
  • Analyze an AI system against the two doors into high-risk classification: Article 6(1) (a safety component of a product covered by the Union harmonisation legislation listed in Annex I) and Article 6(2) (a use case listed in Annex III).
  • Map a system to the eight Annex III areas, and recognize when a use case such as evaluating eligibility for public benefits (Annex III, point 5(a)) puts a system squarely in high-risk territory.
  • Apply the Article 6(3) filter correctly: judge whether an Annex III system escapes high-risk because it performs only a narrow procedural, confirmatory, pattern-detecting, or preparatory task and poses no significant risk, and recognize that any system that profiles natural persons is always high-risk regardless.
  • Separate a prohibited system from a high-risk one, using the Denmark case to see why the same welfare model can raise both the Article 5 social-scoring ban and the Annex III high-risk classification.
  • Identify who bears the classification duty across the provider, deployer, and developer roles, and when a deployer becomes a provider by modifying a system or putting its own name on it (Article 25).
  • Produce a risk classification record for your own AI systems inventory that names each system's tier, cites the Annex III point or Article 5 practice that decides it, and, for anything you call not-high-risk, documents the Article 6(3) reasoning, so the placement is provable and not merely asserted.

The lesson

In Denmark, the National Welfare Authority, known as UDK, and the pension administrator ATP, deployed dozens of algorithmic models to monitor millions of residents. Two of these models were named Really Single and Model Abroad. Their function was to flag citizens for fraud investigation based on perceived anomalies.

Really Single flagged unusual living arrangements, while Model Abroad flagged beneficiaries with ties outside the European Economic Area. In November 2024, Amnesty International released an investigation into this digital apparatus. They found that these systems were pulling together disparate threads of public data, residency records, citizenship status, place of birth, and family relationships, and feeding them into predictive models.

The output was a targeted list for human caseworkers to investigate. Amnesty found these algorithmic flags disproportionately singled out people with disabilities, low-income households, and marginalized racial groups. Denmark's legal failure occurred when the state operated a high-stakes, rights-affecting system under the same governance standards as a low-risk internal spreadsheet.

Under the EU AI Act, every AI system is routed through a strict classification ladder with four distinct tiers, prohibited, high-risk, transparency, and minimal. You evaluate a system by starting at the top rung and stopping at the first tier that applies. This top-down sequence ensures a dangerous system cannot be quietly filed away on a lower rung.

The top tier covers practices prohibited outright by Article 5. This includes manipulative techniques and, crucially, the social scoring of individuals based on their behavior or characteristics. Amnesty argues that Denmark's data-merging welfare models function as a social scoring system. If a legal test confirms that, the models are banned.

But even if they clear that bar, evaluating a person's eligibility to keep a public benefit places a system squarely in the second tier, high-risk. These two tiers operate under completely different regimes. A high-risk system is permitted if it meets extensive compliance obligations.

A prohibited system cannot be saved by paperwork, it must be dismantled. If not prohibited, check the high-risk tier. It has two doors.

Door 1 covers product safety. AI inside regulated physical products inherits high-risk status automatically. Door 2 covers software in eight sensitive Annex 3 areas.

This is where most enterprise systems land, covering employment, law enforcement, and public benefits. Placement requires reading the exact subpoint. Point 5B classifies consumer credit scoring as high-risk, but explicitly carves out AI used to detect financial fraud.

Point 5A, which covers eligibility for public benefits, contains no such exception. Because welfare fraud detection evaluates whether a person remains eligible for a public benefit, it triggers Point 5A. It does not benefit from the credit scoring carve-out.

Door 2 is designed to capture any algorithmic model that makes high-stakes decisions about a person, regardless of how harmless the software looks. The AI Act provides a specific exemption filter for Annex 3 systems. Under Article 6.3, a system may be classified as not high-risk if it performs genuinely narrow, procedural, or preparatory tasks that pose no significant risk to fundamental rights.

Organizations frequently attempt to invoke this filter by arguing that their AI only makes recommendations, leaving the final decision to a human caseworker. In December 2023, the Court of Justice of the European Union addressed this exact dynamic in the SCHUFA credit scoring case. The court ruled that if a human decision-maker draws strongly on an algorithmic score, that score constitutes the decision in substance.

You cannot disclaim automated decision-making simply by placing a human reviewer at the end pipeline. The AI Act reinforces this with a decisive limit. The final subparagraph of Article 6.3, the profiling override, explicitly prevents certain systems from using the exemption.

Under the Act, any system performing profiling, the automated processing of data to predict aspects of a person, is always high-risk. A model that builds a picture from their citizenship status and residency data and family relationships to predict fraudulent behavior is profiling by definition. This profiling trigger nullifies the human-in-the-loop defense.

If a system evaluates behavior to predict risk, it remains high-risk regardless of who reviews it. Classification depends on your role. The Act separates the provider, who builds the system, and the deployer, who uses it.

But under Article 25, if your team fine-tunes a vendor model on your data or wires it into an Annex III decision, you trigger a role switch, becoming the provider. Modifying a system legally transforms a buyer into a provider, meaning the organization instantly inherits all heavy compliance and classification duties. Article 6.4 transforms risk classification into a documented legal assessment that must withstand regulatory inspection.

The law builds the proof burden directly into the process. If you determine an Annex III system is not high-risk, you must document the assessment and produce it on request. Listing a system as low-risk without the specific Article 6.3 justification represents an immediate legal exposure.

Relying blindly on a vendor's assurance that a tool is minimal risk without verifying it against your specific use case transfers their legal liability to you. Conversely, defaulting every basic software tool, like a customer service chatbot, to high-risk drains the exact compliance resources you need to secure your critical systems. An undocumented low-risk classification on a consequential system is the exact point of failure the AI Act was written to catch and penalize.

The Council's digital omnibus deferred the heavy compliance obligations for standalone high-risk systems to December 2nd, 2027. That timeline provides the runway to get compliant. It is not permission to wait.

You cannot prepare a system you have not classified, and the rules governing prohibited systems and transparency disclosures are in effect right now. The immediate task is to construct a rigorous risk classification record. You must evaluate every system in your inventory, assign it a tier, specify the exact legal hook, define your role, and write out the justification.

Prove your risk classification before the regulator asks. It's the only way to ensure an intended safety net never operates as an undocumented surveillance machine.

The ideas, one by one

Four tiers, tested strictest first

Prohibited (Article 5), high-risk (Article 6), transparency (Article 50), minimal. Test the top rung first and stop at the first that decides, so a banned system never hides as merely high-risk and a high-risk system never falls to minimal.

High-risk has exactly two doors

Article 6(1) (a safety component of an Annex I product needing third-party conformity assessment) and Article 6(2) (a use in one of the eight Annex III areas). Most organizations meet the use-case door, and the everyday high-stakes uses (hiring, credit, benefits, policing) are behind it.

The Article 6(3) exit is narrow, and profiling closes it

An Annex III system escapes high-risk only if it does a narrow procedural, confirmatory, pattern-detecting, or preparatory task and poses no significant risk, and any system that profiles people is always high-risk regardless. The "there is a human in the loop" defense does not lower the tier of a profiling system.

Prohibited and high-risk are categorically different

A prohibited system may not be used at all; no conformity file rescues it. The Denmark welfare models are argued to be prohibited social scoring and are separately high-risk on their benefit-eligibility function. Classify to the strictest tier the use triggers, and never build a compliant record for something that is simply banned.

Classification attaches to a role, and the role can switch

The provider makes the primary classification; deployers must still know the tier to meet their own duties and to catch an under-classifying vendor. Under Article 25, a deployer that rebrands or substantially modifies a system becomes its provider, inheriting the heavy obligations. Record your role, and re-run when you modify a system.

"Proving why" is the second half of the job

Article 6(4) requires a provider that calls an Annex III system not-high-risk to document the assessment and produce it on request. A tier is not a decision until it is a documented decision; never record a tier without its reason, and never claim not-high-risk on an Annex III system without the written 6(3) argument.

Classify to the sub-point, not the area

The precise wording carries the carve-outs: Annex III point 5(b) credit scoring exempts fraud detection, while point 5(a) benefits eligibility does not. Reading the general area misses the exact line the legislator drew.

The high-risk timeline moved, and it is runway, not a reprieve

The Digital Omnibus (Council adoption 29 June 2026) deferred stand-alone Annex III high-risk obligations to 2 December 2027 and embedded Annex I high-risk to 2 August 2028. Classification is due now, the prohibitions and transparency and literacy duties already apply, and you cannot make compliant a system you have not classified.

Over-classification has a real cost

Calling a limited-risk chatbot high-risk steals the scarce compliance effort the genuinely high-risk systems need and trains the organization to treat the label as noise. Correct classification cuts both ways, up for the welfare model, down for the chatbot.

The record is a living artifact that travels

Your conformity file is built on it (see Topic 5.6), the deployer impact assessment attaches to it (see Topic 10.4), the dossier inherits it (see Topic 13.1), and a hostile review will attack it (see Topic 11.1). Build it from the real use of each system and a written reason, and it holds; label by gut and it fails at the worst moment.

Classify by use, not by system, and split what does more than one thing

A single system can carry several uses in several tiers, and a single use can answer to several Annex III points at once. Split a system by its distinct uses, classify each, record every applicable hook, and govern the whole to the strictest tier any use triggers. The model layer (a general-purpose model's own obligations, owned by (see Topic 5.5)) and the use layer (the Article 6 tier) are separate questions; the use sets the tier, not the model.

Courts already test the "we only assist" defense, and it fails when the human leans on the output

In the SCHUFA credit-scoring ruling the Court of Justice of the EU held that an agency generating an automated repayment score is itself making an automated decision when lenders draw strongly on it, rejecting the claim that it only performed a preparatory act (Case C-634/21, 7 December 2023). It is a data-protection case, not an AI Act one, but it lands where the profiling override lands: you cannot score a person and disown the decision. "A human decides" is a factual claim about reliance, not a label that lowers the tier.

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 36 of the podcast.

Read the full conversation

Imagine being interrogated by the government, like really interrogated, forced to prove that your family is legitimate, that your relationships are real, and that your living arrangements are standard. And not because a human investigator actually found something suspicious, and definitely not because you committed a crime, but because a mathematically opaque algorithm, one literally named Model Abroad, scoured through these merged public databases and just decided your life looked, well, unusual. Yeah, unusual, which is terrifying.

It sounds like a dystopian novel, right? But that is the actual reality documented recently in Denmark. And legally speaking, for the executives and the governance leaders that we're speaking to today, it represents exactly what happens when a massive organization fails at one foundational regulatory task, and that is risk classification. It really is.

And the fallout from that specific failure in Denmark is just staggering. I mean, if you look at the November 2024 report by Amnesty International, it's titled Coded Injustice. They documented how Denmark's Welfare Authority, which is UDK, along with the Pension Administrator, ATP, they deployed up to 60 of these automated fraud detection models, 60 of them.

And they put marginalized people, specifically people with disabilities, migrants, low income earners, under this intense automated scrutiny. Wow. And the root cause of this, I mean, from a compliance perspective under the EU AI Act, it wasn't just, you know, a flawed line of code.

It was a complete failure to understand what legal bucket their technology actually belonged in. They just categorized it wrong. Exactly.

And that changes everything. So we are going to dissect exactly how that happens, because this whole conversation today is a deep dive into mastering the legal hinge of the entire EU AI Act, which is the classification system or Topic 5.3. If you oversee technology or if you procure software or manage risk for your organization, your entire regulatory fate is determined by how you sort your AI inventory into the Act's legal tiers. That's the hinge.

Right, the hinge. If you get the tier wrong, every single compliance control you build afterward is applied to the wrong system. So we need to walk through the mechanics of the four tiers, the exact legal doors into the high-risk category, the traps embedded in the vendor purchasing process, and the timeline, because the law is already actively enforcing these rules today.

To really understand the stakes for your own organization, we have to look closer at the data architecture that UDK and ATP were actually running in Denmark. I mean, these algorithms with names like Really Single and Model Abroad. I mean, what a name for a government algorithm.

I know, right. But these were not simple spreadsheet macros. They were designed to ingest and cross-reference massive amounts of merged public data.

So we are talking about pulling citizenship records, analyzing family ties, tracking physical movements and residency histories. Just vacuuming up everything. Exactly.

And then running all of that through predictive models to flag benefits recipients for potential fraud. And when a system flags someone for, you know, unusual living arrangements, the burden of proof suddenly shifts to the citizen. The amnesty report includes these testimonies of people feeling like they are, quote, sitting at the end of the gun.

That's a powerful phrase. It is. People were forced to submit highly personal documentation just to prove the shape of their own households, simply because the algorithm flagged an anomaly in the data structure.

Which brings us to the legal synthesis of this whole tragedy. Amnesty's central legal argument under the EU AI Act wasn't merely that the algorithms were biased. I mean, bias is bad, but the legal argument was that UDK and ATP were running these highly consequential rights affecting models as if they were harmless, minimal risk internal analytics tools.

They underclassified them. Severely. They categorized a massive surveillance and evaluation apparatus as a low tier administrative aid.

And by doing so, they completely bypassed the mandatory safeguards. They bypassed the fundamental rights impact assessments, the transparency requirements, basically everything the law demands for systems that actually alter human lives. So if I am listening to this and I'm looking at my own organization's AI inventory right now, the immediate question is, how do I prevent that exact operational blind spot? You called classification the hinge of the entire act.

How is that hinge actually structured in the legal text? Well, the EU AI Act is built on a very rigid risk based pyramid, and it consists of four distinct tiers. At the very top, you have prohibited AI systems. Those are governed by Article 5. OK.

Prohibited is at the top. Right. Just below that are high risk AI systems governed by Article 6. Then moving further down, you have limited risk or transparency systems which fall under Article 50.

And finally, at the very base of the pyramid, you have minimal risk systems. But here is the most critical operational mandate for your governance team. It is the methodology of how you actually assign a system to one of those tiers.

The methodology being what we call the strictest first ladder. Precisely. You do not get to just look at a system, gauge its general vibe and drop it into whichever tier feels appropriate.

You can't just guess. No, absolutely not. You must test every single system in your inventory strictly in order.

You start from the absolute top prohibited and you work your way down and you stop at the very first tier that legally applies. You have to definitively rule out the most severe legal category before you are allowed to even consider the next one down. So it's like a medical triage.

You don't check a patient for a scraped knee before checking if they're having a heart attack. If you don't test strictest first, you might end up putting a Band-Aid on a fatal wound. That is a perfect analogy.

If you skip a step, a ban system might quietly hide as a permitted high risk system or worse. A high risk system hides as minimal risk, which is what happened in Denmark. But let's look at the business reality of that for a second, because if I'm the general counsel at a Fortune 500 company, right, and I'm staring down the barrel of massive regulatory fines for under classifying a system like the Denmark disaster, my immediate legal instinct might be to just stab high risk on literally everything.

The over classification trap. Right. I mean, if I broadly label my internal HR chatbot and my customer service automated routing tools as high risk, I am technically safe from the fines associated with under classification.

Why isn't that a valid conservative legal strategy? Because that strategy will operationally bankrupt your compliance department. Really bankrupted completely and it will create catastrophic alert fatigue across your entire enterprise. If you broadly categorize a limited risk customer service chatbot as a high risk system under Article six, you are legally triggering a massive, inescapable avalanche of obligations.

What kind of obligations? You are forcing your engineering teams to build exhaustive technical documentation. You are requiring continuous, robust quality management systems, continuous automated logging, formal conformity assessments. And you have to physically register the system in an EU database.

Wow. So you are spending millions of dollars mitigating a risk that doesn't actually exist. Exactly.

And the secondary damage is even worse. By drowning your compliance budget and your engineering bandwidth and paperwork for completely harmless tools, you are actively stealing resources and executive attention away from the genuinely dangerous systems. The ones that actually matter.

Yes. The ones actually making credit decisions or screening resumes or evaluating biometric data. Look, when everything in an organization is classified as an emergency, nothing is an emergency.

You train your internal teams to view the high risk label as just bureaucratic noise rather than a signal of critical legal exposure. Precision is your only survival strategy here. You cannot guess low and you really cannot afford to lazily guess high.

Which means we need to deeply understand the exact boundary lines between these tiers, starting with the top two, because when you read the act, the difference between a prohibited system under Article 5 and a high rate system under Article 6. I mean, it isn't just a matter of how many forms you have to fill out. These are categorically different legal universes. They are entirely, fundamentally different.

Article 5 is an outright unconditional ban. The act identifies a very specific set of AI practices that the European Union has deemed inherently unacceptable because they pose a threat to fundamental rights. And you can't fix them.

No. No amount of technical safeguarding or human oversight or logging can mitigate that threat. We're talking about practices like the social scoring of natural persons based on their social behavior or the use of subliminal manipulative techniques to materially distort behavior causing harm or certain uses of real time remote biometric identification in publicly accessible spaces.

Meaning you cannot build a robust compliance file to somehow rescue an Article 5 system. You cannot. You may not place it on the market.

You may not put it in service and you may not use it. Full stop. It's banned.

Okay. And Article 6. Article 6, which governs high risk systems, is a framework of permission. High risk systems are legally allowed to exist and operate in the European market, but only if they carry the full weight of the act's mandatory requirements.

Permitted, but with heavy rules. Exactly. The legislator is essentially saying this technology is dangerous, but it has legitimate utility.

So you may use it provided you execute a formal conformity assessment, guarantee human oversight, ensure high quality data governance and maintain detailed logs. Okay. Let's apply that stark boundary back to the models in Denmark.

Amnesty International's central argument is that UDK and ATP's models, by merging personal data to flag unusual behavior, function as a social scoring system. Right. And the implications of that specific legal argument are completely binary.

If an auditor or regulator or a court agrees that evaluating welfare recipients based on personal characteristics and social behavior to subject them to disproportionate investigations, if that constitutes social scoring under Article 5, then the system is prohibited. And so game over. The legal analysis ends immediately.

The welfare agency cannot negotiate a mitigation strategy. They cannot say, oh, we'll implement stronger human in the loop protocols. The system must be dismantled.

But let's play this outright, because lawyers love to argue. Let's say the agency's legal counsel mounts a really aggressive defense. They argue in court, listen, this is absolutely not generalized social scoring.

This is a highly targeted, mathematically proportionate financial fraud investigation model. It's strictly limited to the context of welfare payouts, and it does not evaluate general social trustworthiness. Okay.

The standard defense. Let's assume just for the sake of the exercise, they win that argument. The system clears the Article 5 ban.

Does it then drop all the way down to a minimal risk analytics tool? This is exactly why the strictest first ladder is the fundamental rule of the act. If a system clears the prohibited tier, you do not get to just jump to the bottom. You step down exactly one rung to the high risk tier.

Why automatically high risk? Because the AI Act explicitly lists systems used by public authorities to evaluate the eligibility of natural persons for public assistance benefits and to grant, reduce, revoke or reclaim those benefits as inherently high risk. So even if the agency successfully proves the system is not banned under Article 5, it lands squarely in the heaviest regulatory bucket available for permitted systems. But if I'm advising the board, my immediate concern is how to handle the operational gray area while the lawyers are actually arguing about that boundary.

Say our internal team builds a novel fraud detection model. The general counsel looks at it and says, this is right on the line. It might trigger the Article 5 ban on social scoring or it might just be high risk under Article 6. We need to commission an external legal opinion and review recent case law.

A very realistic scenario. Right. So what do we do with the software in the meantime? Do we just run it as high risk, keep the business operations moving and build the compliance file while we wait for the final legal verdict? Absolutely not.

You must immediately restrict or pause the system. Pause it completely. Yes.

You tag it internally as prohibited pending. The operational mandate here is critical. If a system has a credible possibility of being prohibited under Article 5, running it while you assemble an Article 6 compliance file is an active, ongoing legal violation.

Oh. Every single day that system operates, you are unlawfully subjecting individuals to a banned practice. No amount of retroactive paperwork or delayed conformity assessments will cure the fact that you operated a prohibited system.

When you are navigating the boundary between banned and high risk, you halt the machine first and you answer the legal question second. Contrast that with uncertainty lower down the ladder. If we are debating whether a tool is high risk or minimal risk, the operational approach is different, isn't it? Completely different.

If the uncertainty is between Article 6 high risk and Article 50 transparency or minimal risk, the safest operational default is to classify it upward to high risk while you resolve the doubt. Because in that scenario, the system is at least permitted to operate. Treating a permitted system with a higher degree of safety and oversight than legally required doesn't violate a fundamental ban.

But at the absolute top of the ladder, the mandate is clear. Do not run a potentially prohibited system. OK, so we run the strictest first test.

The system clears the top rung. We are confident it is not prohibited. Now we need to determine if it is high risk.

And the Act doesn't leave this up to interpretation or general themes, does it? It establishes very specific, hard-coded legal gates. It does. When we move from the outright bans of Article 5 to the permissions of Article 6, we are looking at exactly two doors into the high risk category.

If your A.I. system fits through either of these two doors, it is legally high risk. There is no third door. Walk us through Door 1. Door 1 is defined in Article 6, Paragraph 1. This door relies on the concept of inherited risk for safety-critical product components.

And it points directly to Annex I of the A.I. Act. Annex I? Yes. Annex I is essentially a comprehensive list of existing European union harmonization legislation.

We're talking about the established product safety laws that govern physical, tangible items. So medical devices, civil aviation, machinery, toys, marine equipment, elevators. The physical infrastructure of the economy.

Things that can literally physically injure you if they fail. Exactly. And the legal mechanism for Door 1 is incredibly elegant.

The rule states that if your A.I. system is intended to be used as a safety component of a product or is itself a product covered by Annex I, and that product is already required under EU law to undergo a third-party conformity assessment before being placed on a market, then the A.I. system automatically inherits high risk status. Ah, I see. So the EU isn't reinventing the wheel here.

If the existing regulations for medical devices already mandate that a new MRI machine requires an independent safety audit because a failure could harm a patient, the A.I. Act simply piggybacks on that logic. That's right. If you embed an A.I. system into that MRI machine to, say, interpret the scans, the A.I. is doing a safety-critical job inside a highly regulated physical product.

So it just automatically becomes high risk. That is exactly the mechanism. It is inherited risk based on existing European product safety frameworks.

But while Door 1 is critical for manufacturers of physical goods, the reality is that the vast majority of software companies, enterprise I.T. departments and corporate data teams will never touch Annex I. They're dealing with enterprise software, human resources, tools, financial algorithms. For them, the legal pathway is Door 2. Let's open Door 2. Door 2 is Article 6, Paragraph 2. And this points to Annex 3 of the A.I. Act. Annex 3 is a definitive list of eight highly sensitive use case areas.

The legislative intent here wasn't to regulate specific underlying technologies like neural networks or large language models. The intent was to regulate high stakes decisions about human lives. What are the eight areas defined in Annex 3? The eight broad areas are biometrics, critical infrastructure management, education and vocational training, employment, workers management and access to self-employment.

OK, that's four. Then you have access to and enjoyment of essential private services and essential public services and benefits, law enforcement, migration, asylum and border control management. And finally, the administration of justice and democratic processes.

I mean, when you list them out like that, you realize these are the foundational pillars of participation in modern society. Can I get a job? Can I get a loan? Can I cross a border? Will I be arrested? Will I receive medical or welfare benefits? Exactly. If your standalone software system makes or materially supports a decision in one of these eight areas, it comes through door two.

But this is where organizations make one of the most common and financially devastating analytical errors. What's the error? Knowing that your software operates generally within the financial services or public benefit space is not enough. You cannot classify your systems by their general theme.

You must read the specific sub points within Annex 3 because the exact wording of the sub point dictates your regulatory burden. OK, give me a concrete example of how classifying by theme or vibe falls apart in practice. Let's look really closely at Area 5 in Annex 3. That's the one covering access to and enjoyment of essential private services and essential public services and benefits.

Under this broad umbrella heading, there are highly specific sub points. Let's compare Point 5A and Point 5B. Point 5A explicitly covers AI systems used by public authorities to evaluate the eligibility of natural persons for public assistance, benefits and services.

This is the exact legal hook that catches the models used in Denmark. OK, so welfare, public pensions, government assistance. Right.

Now look at Point 5B. This sub point covers AI systems intended to be used to evaluate the credit worthiness of natural persons or establish their credit score. OK, so both sub points deal with evaluating individuals' financial standing.

One is just for public government money and the other is for private bank money. Yes, but the legislative text treats them vastly differently. Point 5B, the credit scoring clause, contains an explicit, hard-coded legislative carve out.

It states that credit scoring systems are high risk, with the exception of AI systems put into service by small scale providers for their own use and crucially, except for AI systems used for the purpose of detecting financial fraud. Wait, really? Yes. The legislators specifically negotiated a carve out so that transaction fraud detectors used by commercial banks would not be blanketed as high risk.

But when you read Point 5A, the public benefits clause, is there a parallel carve out for welfare fraud? There is absolutely no fraud carve out in Point 5A. That is a massive distinction. So if I am a B2B software vendor and I develop a highly sophisticated anomaly detection algorithm, and I sell that exact algorithm to a commercial bank to detect fraudulent credit card transactions, I look at Annex 3, Point 5B.

I see the explicit carve out for financial fraud. I document that my system is exempt from the high risk category, and it likely drops down to minimal risk. My compliance burden is virtually zero.

Correct. You have an exemption. But if I take that exact same underlying code base, tweak the data parameters, and sell it to a government agency to detect welfare fraud, I am caught squarely under Point 5A.

There is no carve out. My system is now legally high risk. It is.

I now have to build a risk management system, ensure human oversight, maintain automated logs, write extensive technical documentation, and register the system in the EU database. You have articulated the trap perfectly. The exact same underlying technology sitting on the same shelf put into two different operational bottles results in two entirely different legal realities.

If you classify by the general vibe, if your compliance team just says, oh, it's just a fraud detector and we know fraud detectors are exempt under Area 5, you commit a fatal regulatory error. Because welfare fraud detection is not credit scoring fraud detection. Exactly.

The exemptions and carve outs live entirely within the microscopic wording of the subpoints, not in the general area headings. If I am an executive staring down a million dollar compliance bill because my system hit one of these specific Annex 3 subpoints without a carve out, my immediate instinct is to call my legal team and tell them to find a loophole. I mean, I am going to ask them to find a way to argue that the A.I. doesn't actually make the final decision, that it just supports our human staff.

Does the act leave a back door open for that kind of defense? The law does anticipate that exact executive reaction and it does provide an exit pathway. It's found in Article 6, Paragraph 3, but it is an incredibly narrow exit and one that corporate legal departments routinely misunderstand, which leads to massive exposure. Let's dissect Article 6, 3. How does a system clearly land in a high risk Annex 3 area, but then legally escape the high risk obligations? Article 6, 3 acts as a highly restrictive filter.

It states that an A.I. system whose intended purpose falls under Annex 3 shall not be considered high risk if two things are true. First, the system must pose no significant risk of harm to the health, safety or fundamental rights of natural persons, including by not materially influencing the outcome of decision making. OK, that's the first hurdle.

And second, the system must meet at least one of four strict procedural conditions. What are those four conditions? The system must either perform a narrow procedural task, improve the result of a previously completed human activity, detect decision making patterns or deviations without replacing or influencing human assessment, or perform a purely preparatory task to an assessment relevant to the Annex 3 use case. I need to see what an honest, legitimate use of this exit looks like before we talk about how companies abuse it.

Give me an example of a system that sits in a sensitive Annex 3 area, but genuinely qualifies to use the 6, 3 exit. OK, imagine an A.I. tool deployed in a government benefits office. Yeah.

Its sole function is optical character recognition. It converts scanned, handwritten benefits applications into structured digital text. OK, OCR, just reading handwriting.

The human caseworker then sits down, looks at both the original scanned document and the digital text, and the human independently evaluates the information to decide if the applicant is eligible for the benefit. The A.I. is technically operating within the Annex 3 Area 5A space public benefits, but it only performs a narrow procedural task. It's just transcribing.

Exactly. It does not evaluate the applicant. It generates no probability scores, and it does not materially influence the outcome because the human is evaluating the source information themselves.

If you can meticulously document those facts, that system can use the 6, 3 exit and escape the high risk classification. That makes total sense. It is essentially acting as an administrative assistant, not an evaluator.

But here is where the corporate loophole hunt begins, right? Companies build highly complex predictive scoring algorithms that ingest massive amounts of data and spit out a stark recommendation like a red flag saying investigate this person for welfare fraud or deny this mortgage application. Yes. And then the corporate lawyers draft a memo saying, well, the A.I. doesn't actually finalize the decision.

A human caseworker has to physically click the approve or deny button on the screen. Therefore, the A.I. is just performing a preparatory task. It is just decision support.

This is the classic human in the loop defense. It is the most common defense in the industry, and legally it is built on sand. To understand why the just support myth fails, we have to look outside the text of the A.I. Act itself and look at how the Court of Justice of the European Union, the C.J. E.U., interprets automated decision making.

We have a definitive precedent from December 7, 2023. It's a landmark case involving Shufa. Shufa is the major credit scoring agency in Germany, correct? On Equifax.

Similar to Equifax or Experian in the U.S.? Correct. The mechanics of the case were exactly what we are discussing. Shufa generates automated probability scores about individuals, predicting their likelihood to repay a loan based on historical and personal data.

Shufa argued in court that they merely produce a preparatory score. They send it to the bank? Yes. They argued that they send this score to the bank, and the bank's human loan officer makes the actual decision to grant or deny the credit.

Therefore, Shufa claimed they were not engaged in automated decision making about natural persons because the human was the ultimate decision maker. They were trying to claim the exact equivalent of an Article 6.3 preparatory task exemption. Exactly.

But the C.J. E.U. rejected that argument entirely. The court analyzed the reality of the business process. They ruled that when a human decision maker draws strongly on an automated probability score to make their final choice, the automated score is the decision in substance.

The score is the decision? Yes. The court recognized that you cannot generate a highly consequential predictive score, hand it to a human who routinely relies on it, and then legally disown the final decision. This gets into the cognitive psychology of automation, doesn't it? If a human caseworker has a quota to review 50 welfare fraud flags an hour, they do not have the time or the cognitive bandwidth to conduct 50 independent Menovo investigations from scratch.

They don't. They are going to trust the machine. They're going to draw strongly on the A.I.'s recommendation.

Yes. It is a well-documented phenomenon known as automation bias or automation fatigue. Think of it like relying on GPS navigation.

If you are driving and the GPS tells you to turn left into a lake and you do it, you cannot blame the GPS for driving the car into the water. You are holding the steering wheel. But the C.J. E.U. is legally recognizing the reality of automation fatigue.

If the GPS has been right the last 500 times, the human brain stops looking out the windshield and just turns the wheel when prompted. The law is saying if you build a system that practically forces a human to stop looking out the windshield and just accept the prompt, the system is the actual driver. Wow.

If the human materially relies on the A.I. flag, the A.I. is materially influencing the decision. The moment that happens, it fails the Article 6.3 exit conditions. It snaps immediately back to high risk.

But the A.I. Act actually goes one step further than just relying on the specific case law, doesn't it? There is a structural failsafe built into the text of Article 6.3 to prevent this exact loophole. Yes, there is. It is perhaps the most important sentence in the classification rules.

The final subparagraph of Article 6.3 contains what we refer to as the profiling override. This is the ultimate unconditional backstop against the human-in-the-loop defense. Walk us through the mechanics of the profiling override.

The override explicitly states that if an A.I. system performs profiling of natural persons, the Article 6.3 exit is closed completely. It does not matter if you meet all four of the procedural conditions. It does not matter if a human hits the final button.

If the system profiles humans, it is unconditionally high risk. Profiling is a word that gets thrown around casually in tech and marketing, but it has a very specific, hard-coded legal definition here. How does the law define it? The A.I. Act does not invent a new definition.

It borrows the established definition directly from the GDPR, specifically Article 4, Paragraph 4. Profiling is defined as any form of automated processing of personal data to evaluate certain personal aspects relating to a natural person. In particular, to analyze or predict aspects concerning that person's performance at work, economic situation, health, personal preferences, interests, reliability, behavior, location, or movements. Let's connect that GDPR definition directly back to the Amnesty International report on Denmark.

The UDK models, like Model Abroad, were taking a citizen's residency data, cross-referencing their family ties, tracking their physical movements, and analyzing their citizenship status to predict the likelihood that they were committing welfare fraud. Which is the textbook definition of automated processing of personal data to analyze and predict a person's behavior, economic situation, and reliability. It is profiling by the exact letter of the law.

Absolutely. And because the models were profiling natural persons within an Annex 3 area, the profiling override dictates that the Article 6.3 exit is legally sealed shut. UDK could argue all day that the algorithm only performs a preparatory task, and they could point to the human caseworkers who conducted the final interviews.

It is legally irrelevant. The profiling override dictates that any system profiling natural persons in an Annex 3 area is high risk, period. So the takeaway for the executive listening is this.

If your software is generating predictive scores about human behavior to evaluate them, you are high risk. You cannot lawyer your way out of it by pointing to the intern who hits the enter key on the final screen. You cannot.

And realizing that the exit is closed leads to the next massive realization for a corporate leader. Once you accept that a system is firmly classified as high risk, the immediate existential question becomes, who has to do all the compliance work? Who is legally and financially on the hook to build the conformity files, run the risk management systems and register with the EU? Right. Because the instinct of every software buyer is to push the liability onto the seller.

I mean, if I am the head of H.R. at a logistics company and I buy a highly sophisticated A.I. resume screening tool from a massive enterprise software vendor, my assumption is that the vendor built the A.I. The vendor coded the algorithm. So the vendor is the regulated entity. I am just the customer paying a subscription fee.

That assumption is one of the most dangerous misconceptions in the market today, and it will get your organization into immense trouble. In the EU A.I. Act, classification is not a free-floating label that applies to the software in a vacuum. The classification and the burdens that come with it attach to specific legal roles.

To understand who does the work, you have to understand the difference between the provider and the deployer. Let's define the provider first. The provider is the actor, whether a natural or legal person who develops an A.I. system or has it developed and places it on the market or puts it into service under their own name or trademark.

The provider bears the heaviest regulatory burdens of the entire A.I. Act. So what exactly do they have to do? If it is high risk system, the provider must design and implement the continuous risk management system. They execute the formal conformity assessment.

They draw up the exhaustive technical documentation and they physically register the system in the EU database. OK, and what is the definition of the deployer? The deployer is the actor who uses the A.I. system under their own authority in the course of their professional activity. In your example, the corporate H.R. department that purchases the resume screening tool is the deployer.

Deployers do not have to write the core technical documentation or run the initial conformity assessment, but they are not exempt from the law. They have their own specific set of duties outlined in Article 26. What does Article 26 require a deployer to do? A deployer must use the system strictly according to the provider's instructions.

They must ensure that human oversight is actually functioning and assigned to competent personnel. They must monitor the system's operation for anomalous behavior or risks, keep automated logs generated by the system, and crucially, they must inform the provider if they identify any serious incidents or risks to fundamental rights. OK, so as the H.R. executive, I am the deployer.

The software vendor is the provider. As long as I follow the instruction manual and keep my logs, the vendor has to do the heavy lifting of the high risk conformity assessment. Normally, yes, that is the standard balance of obligations.

But here is the legal trapdoor that every procurement team and data science leader needs to memorize. It is found in Article 25. Article 25 contains the role switch rules.

Role switch, meaning my legal status transforms from deployer to provider? Exactly. Under Article 25, a deployer legally transforms into a provider if they do one of three things. If they put their own name or trademark on a high risk system, or if they modify the intended purpose of a non-high risk system so that it becomes high risk.

Or, and this is the most common operational trap, if they make a substantial modification to a high risk AI system. Let's unpack the data science mechanics of a substantial modification. What does that actually look like on the ground in a normal corporate environment? OK, let's stick with your HR department.

Your team procures a generic vendor provided large language model. Out of the box, it is classified as a general purpose model. It's perhaps subject to some transparency rules, but it is not inherently an annex the high risk system.

It just processes text. Right. It's just a generic tool.

But your internal data science team looks at it and says, we can make this much more useful for our specific company. So they take the vendor's base model and they fine tune it. They ingest 10 years of your company's proprietary historical hiring data.

Who got promoted? Who got fired? What keywords correlated with long term retention at your specific firm? And they train the model on that exact data. They are adjusting the weights and biases of the model to optimize it for our specific corporate environment. And by doing so, they are wiring it directly into an Annex 3.4 use case.

Employment, workers management and access to self-employment. Exactly. By fine tuning the model on your proprietary data and wiring it into a high risk decision making process for hiring, your data science team has not just configured the software.

Under the law, they have substantially modified it and fundamentally changed its intended purpose. Wait, so. The moment the fine tuning is deployed, Article 25 triggers instantly.

Legally, your HR department is no longer just the deployer of a general purpose tool. You have legally transformed into the provider of a newly created high risk AI system. That is a staggering shift in liability.

I thought I was just a software buyer customizing my settings to get better ROI on my subscription. But by letting my data team fine tune the algorithm, Article 25 just dumped the entirety of the high risk compliance burden onto my desk. I have to design the risk management system.

I have to execute the conformity assessment. I have to register it with the EU. The vendor is no longer responsible.

I am. You are. And ignorance of this role switch does not shield you from the lie at Ruby.

A regulator will not accept the defense that your IT team didn't know fine tuning constituted a substantial modification. This is why risk classification cannot be a static one time checklist you fill out at procurement. It must be a living, breathing record.

You classify the tool when you buy it. But the moment your data team says, hey, we improve the vendor tool by training it on our internal data. The governance team is instantly rerun the classification analysis to see if your role just switched from deployer to provider.

This also highlights a massive blind spot during the initial software purchase. If a vendor or sales team comes to me and says, don't worry, you can buy our AI tool. We've already done the legal analysis and classified it as minimal risk.

I cannot just accept their marketing brochure as legal cover, can I? Because if they intentionally under classified the system to make the sale easier and I deploy it. You are liable as a deployer. Under Article 26, a deployer must independently know the true tier of the system to meet its own operational duties and to ensure they are not operating an unmitigated high risk system.

If you deploy a system that actively profiles natural persons for employment decisions and you fail to maintain human oversight or keep logs because the vendor lied to you and said it was low risk, the regulator is going to ask why you didn't do your own independent annex third analysis. You cannot outsource your foundational classification judgment entirely to a vendor sales pitch. You have to verify the tier.

OK, so the path is clear, but incredibly rigorous. We have to run the strictest first test. We had to find the exact sub point in annex three.

We have to recognize that the six three exit is sealed shut if the system profiles natural persons and we have to continuously monitor for the Article 25 rule switch. How does a governance leader actually operationalize this? What is the tangible artifact that proves we did all of this correctly? The artifact is the mandatory documentation of your reasoning. Risk classification under the EU Act is not just a label you mentally assign in a meeting is a documented defensible statutory decision.

This brings us to the requirements of Article six for the duty to prove why on the record. Exactly. Article six four establishes that if you evaluate a system that falls within the domains of annex three, but you conclude that it is not high risk, for example, because you believe you genuinely qualify for the six three exit because the tool only performs a narrow procedural task, you must formally document that assessment.

You have to show your work. You have to write down exactly why it performs a narrow task, why it poses no significant risk to fundamental rights, and you must explicitly prove in writing why it does not profile natural persons. So if a regulator or an auditor knocks on the door and says, we notice you are using an A.I. system to process benefits applications, but it isn't registered in the EU database as high risk.

Why? Under Article six four, you are legally obligated to provide that written documentation to the national competent authorities upon request. If you reply by saying, oh, our internal risk committee looked at it last year and we just decided it was decision support. So it's low risk and you hand them nothing.

An undocumented decision is not a classification. It is a massive legal exposure. The legislative drafters specifically anticipated that companies would be heavily tempted to creatively label Annex three systems as not high risk to avoid the compliance costs.

So the law demands that you show your work. You have to write the essay proving you deserve the exemption. And if you can't write a coherent essay, you don't get the exemption.

And that documented essay, that classification record becomes the load bearing foundation of your entire compliance architecture. You cannot build a conformity file. You cannot draft technical documentation and you cannot train your deployers on human oversight.

If you haven't definitively proven to yourself that the system is high risk in the first place. Which brings up the most urgent issue for our listeners. The timeline.

Yeah, because there's a very dangerous myth circulating in corporate networks right now. I hear executives say all the time the AI Act was passed, but the high risk rules don't actually kick in until late 2026 or 2027. We have plenty of time.

We will worry about classifying our systems next year. We are recording this deep dive in the September 2026. What is the actual current legal timeline that executives need to operate against today? It is vital we correct this misconception because the timeline changed recently and operating on old advice is incredibly dangerous.

Initially, the high risk obligations were said to apply 24 to 36 months after the Act's entry into force, which would have meant late 2026. However, on the 29th of June, 2026, the Council of the EU formally adopted an amendment known as the Digital Omnibus on AI. What exactly did the Digital Omnibus change? The Omnibus deferred the application date of the heavy high risk obligations in Article 6. For standalone Annex 3 systems, which is the software we spent most of our time discussing today, like HR screening, welfare evaluation and biometric categorization, the high risk obligations are deferred and now apply from the 2nd of December 2027.

OK, December 2027. For high risk AI systems that are embedded as safety components and annex the physical products, the obligations are deferred even further to the 2nd of August 2028. So the executive who says we have until December 2027 is technically quoting the right date for Annex 3 systems.

But what are they getting wrong operationally? They're mistaking a runway for a reprieve. December 2nd, 2027, is not the date you start thinking about compliance. It is the date you must be completely finished with compliance.

It takes organizations months and in complex environments over a year to build a functioning, continuous risk management system to overhaul their data governance architecture, to meet the Act's rigid standards, to conduct fundamental rights impact assessments and to complete the formal conformity assessments. You cannot build a comprehensive conformity file in November 2027 for a system you haven't even bothered to classify yet. I look at it like a hurricane forecast.

If the meteorological models tell you a Category 5 storm is going to hit your coast in December 2027, you don't wait until November 2027 to start pouring the concrete for your seawall. You start today. Because if you wait until the month before, the storm is going to hit while you are still trying to read the blueprints.

That is exactly the reality. But to extend your metaphor, the rain has already started falling. The deferral and the digital omnibus only applies to the high risk operational obligations in Article 6, the absolute prohibitions in Article 5. Those are already in force.

They have applied since early 2025. Wow. OK.

The transparency duties under Article 50, those already apply. So if you haven't run the strictest first classification ladder on your systems yet, you do not even know if you're currently running a banned Article 5 system today. You cannot use a December 2027 deadline for high risk systems to excuse an ongoing active violation of a prohibited practice in September 2026.

So classification is not a future project. It is due right now. Yes.

The classification record is the map that tells you which of your systems need the seawall bill before December 2027, which systems just need a transparency notice today and which systems you need to unplug immediately because they are actively prohibited. Let's synthesize everything we have covered because the framework is dense but absolutely critical. Classification is the hinge of the entire AI act.

If you get it wrong, you are either operating a dangerous tool illegally or you are drowning your compliance budget in unnecessary paperwork. You must test your systems strictest first. Start at the top with prohibited systems under Article 5 because that is an outright ban.

Then move to high risk under Article 6, which permits the technology but applies massive operational duties. And when you are analyzing Article 6 high risk, remember the two doors. Door 1 is Annex I inherited risk for physical products.

Door 2 is Annex 3 for the 8 sensitive use cases. When you're looking at Door 2, you must classify to the exact sub point. The microscopic difference between .5A for public benefits and .5B for private credit is the difference between having a fraud carve out and having no carve out at all.

And beware the temptation of the narrow Article 6.3 exit. A human in the loop rubber stamping an AI flag does not save you. As the CJEU proved in the Schieffer ruling, if a human draws strongly on an automated score, it is a decision in substance due to automation bias.

More importantly, if the system profiles natural persons, if it predicts their behavior, economic situation or reliability based on their data, the profiling override unconditionally slams the 6.3 exit shut. It is high risk. Furthermore, you must track your legal role continuously.

Watch out for the Article 25 role switch. If your internal data science team takes a generic vendor tool and fine tunes it on your proprietary company data to make high stakes decisions, you have substantially modified the system. You have legally transformed from a deployer into a provider and all the heavy regulatory obligations instantly transferred to your desk.

And finally, you have to prove all of this on paper. Article 6.4 demands that if you claim an Annex 3 system isn't high risk, you must have the documented reasoning drafted and ready to hand to a regulator upon request. Because an undocumented decision is not a legal classification.

It is just a massive exposure waiting to be discovered. So how do we translate all of this framework into immediate action? What is the single most valuable move our listeners can make at the start of their work week? The Monday morning move. For your Monday morning move, open your organization's AI systems inventory.

Find just one system operating in an Annex 3 area like employee management, public benefits or credit evaluation that your internal team currently labels as not high risk, decision support or minimal risk. Sit down and try to write the Article 6.4 justification document that the law requires. Try to explicitly prove in writing why that system does not profile natural persons under the GDPR definition and why it does not materially influence the human decision maker.

If you cannot write a legally defensible paragraph proving those two specific points, you have just found a major regulatory exposure hiding in plain sight. You need to fix it before the regulator knocks on the door and asks for the documentation. And as a final thought to leave you with, consider the long term reality of that documentation.

The EU AI Act forces organizations to rigorously expose and document the inner workings, the data flows and the decision making logic of their AI systems. Think about what happens when those mandatory compliance logs and technical documents become discoverable in civil litigation. You aren't just writing a compliance file for a regulatory auditor today.

You might be drafting Exhibit A for a discrimination lawsuit tomorrow. The rigor you apply to your classification process right now will dictate how defensible your entire organization is in the future. It is about clearing the muddy waters before you are forced to navigate them in court.

You need the X-ray and the classification record is that X-ray. Thank you for joining us on this deep dive. Take that Monday morning move, build your definitive record, and we will catch you next time.

Real cases

These examples show risk classification done well and badly, with the legal reasoning made explicit. The deep anchor is Denmark's welfare AI; the others sharpen one point each and are owned in depth by their own topics.

Example 1 (the anchor): Denmark's welfare fraud models and the tier that was skipped. Between the pension administrator ATP and the welfare authority UDK, Denmark ran as many as sixty fraud-detection models over the merged personal data of millions of residents, including "Really Single" (flagging "unusual" living arrangements in pension and childcare schemes) and "Model Abroad" (flagging beneficiaries with strong ties to non-EEA countries), and Amnesty International's November 2024 report "Coded Injustice" documented that these systems risk discriminating against people with disabilities, low-income people, migrants, and marginalized racial groups (Amnesty International, "Coded Injustice: Surveillance and Discrimination in Denmark's Automated Welfare State," November 2024). Read against this topic's four questions, the classification these systems demanded is stark. Question 1: Amnesty argues the apparatus functions as social scoring, which Article 5 prohibits; if that argument holds for a given model, the model is banned and the analysis ends. Question 2: even setting the prohibition aside, a system a public authority uses to evaluate benefit eligibility and to reduce, revoke, or reclaim benefits is high-risk under Annex III point 5(a) by name, and because these models profile people from residency, citizenship, and family data, the Article 6(3) profiling override keeps them high-risk regardless of the "we only support caseworkers" defense. The lasting lesson: the harm did not begin with a bad model, it began with a system operated as if it were a minimal-risk internal analytics tool when its use put it on the top two rungs of the law. Classification was the control that would have forced the safeguards, and it was the control that was skipped.

Example 2 (the same area, the credit sub-point): consumer credit scoring under Annex III point 5(b). Annex III point 5(b) makes AI used to evaluate the creditworthiness of natural persons or set their credit score high-risk, with a narrow carve-out for AI used to detect financial fraud. A bank running an AI model to approve or price consumer loans is squarely high-risk on this hook, and the classification is usually uncontested. The instructive subtlety is the fraud carve-out: an AI used purely to detect financial fraud in transactions is excepted from the credit-scoring category, which is a deliberate line the legislator drew, and it is a reminder that classification is exact wording, not theme. Note the contrast with the welfare case: welfare-benefit fraud detection is not credit scoring, so the 5(b) fraud carve-out does not rescue it; it lives under 5(a), which has no such carve-out. Reading the precise sub-point, not the general area, is the skill. The deep treatment of specific enforcement in these areas belongs to other Module 5 topics; here the point is that neighboring sub-points carry different carve-outs and you must classify to the sub-point.

Example 3 (the employment door, referenced): AI hiring screens under Annex III point 4. AI used to recruit, screen, or select candidates, or to make promotion and termination decisions, is high-risk under Annex III point 4. The 2025 litigation over Workday's hiring-screening AI, where a federal court allowed an age-discrimination collective action to proceed and treated the vendor's tool as an agent of the employers, shows why the tier fits: a screening system materially influences who gets a job, a textbook fundamental-rights effect (deep treatment and the case citation are owned by Topic 5.4 (see Topic 5.4)). For classification, the lesson is two-fold: the employment area is one of the most common ways an ordinary company's AI becomes high-risk, and the role question bites hard here, because an employer that fine-tunes or configures a vendor screen to its own criteria may cross Article 25 into being the provider of a high-risk system, not merely its deployer.

Example 4 (the filter used honestly): a genuinely procedural Annex III system. Not every system that touches an Annex III area is high-risk, and Article 6(3) exists for real cases. Consider an AI that does nothing but convert scanned benefit application forms into structured text for a human assessor who then independently decides eligibility, reading the original documents. If the extraction does not evaluate the person, does not profile them, and does not materially influence the outcome (the assessor decides from the source documents, not the AI's summary), it can qualify for the Article 6(3) narrow-procedural or preparatory-task exit and be classified not-high-risk. The discipline is that this classification is a claim you document under Article 6(4), and it holds only as long as the facts hold: the moment the extraction starts scoring applicants or the assessors start deferring to it, the profiling and significant-risk tests reengage and the system climbs back to high-risk. The honest use of the filter is narrow and evidenced; the dishonest use is calling a scoring system "just support."

Example 5 (transparency risk, not high-risk): a customer-service chatbot. A retail company runs a generative chatbot that answers product questions. It does not decide anything in an Annex III area, so it is not high-risk. But it interacts with people who might think they are talking to a human, so Article 50 requires it to disclose that it is an AI system. Classify it as transparency-risk, record the Article 50 hook, and the obligation is a clear notice, not a conformity file. This example matters because it is the opposite error from Denmark: over-classifying a genuinely limited-risk system as high-risk wastes the scarce compliance effort you need for the systems that are actually high-risk. Correct classification cuts both ways, up for the welfare model and down for the chatbot.

Example 6 (minimal risk, and why "minimal" is still a decision): a demand-forecasting model. A logistics firm runs an AI model that forecasts warehouse demand to optimize stock. It touches no Annex III area, interacts with no consumer, and generates no content, so it is minimal risk with no mandatory obligations under the Act. The reason to still record it, with the reason "no Annex III use, no Article 50 trigger," is that classification is a living decision: if the firm later repurposes the same model to, say, score gig workers for shift allocation, it has entered Annex III point 4 and must be reclassified. A minimal-risk record with its reason is what makes that later change visible; an unclassified system is one whose creep into a higher tier no one notices until an incident. Minimal is a tier you assign on purpose, not a synonym for "we did not look."

Example 7 (a court closes the "we only prepare, the human decides" defense): the SCHUFA credit-scoring ruling. The Court of Justice of the European Union ruled in December 2023, in the SCHUFA case, that when a German credit-reference agency generates an automated probability score of whether a person will repay, and lenders "draw strongly" on that score to grant or refuse credit, the agency is itself engaged in automated individual decision-making under Article 22 of the General Data Protection Regulation (Court of Justice of the EU, Case C-634/21, Judgment of 7 December 2023, SCHUFA Holding). SCHUFA had argued that it merely produced a score and that the bank made the decision, so the rules on automated decisions fell on the bank, not on it. The Court rejected that: where the human decision-maker draws strongly on the score, the score is the decision in substance, and the scorer cannot disclaim it as a mere preparatory act. This is a data-protection case, not an EU AI Act ruling, but read next to the Article 6(3) filter the lesson is exact. The "we only support a human who really decides" defense is not a magic phrase; it is a factual claim that a court will test against how much the human actually relies on the output. For classification it reinforces two things at once: creditworthiness scoring of natural persons is the textbook example of a consequential, profiling decision (the Annex III point 5(b) hook from Example 2), and the "the AI only assists, the human decides" argument fails exactly when the human leans on the AI, which is the very test that closes the Article 6(3) exit under the AI Act. Two different laws, data protection and the AI Act, arrive at the same place the profiling override reaches: you cannot score a person and then disown the decision.

Where people go wrong

  • "There is a human in the loop, so it is not high-risk." This is the Denmark defense, and it fails twice. Article 6(3)'s pattern-detection exit requires that the AI not replace or improperly influence the human assessment, and a caseworker who defers to the flag is being influenced; and the profiling override makes any system that profiles people high-risk regardless of human review. A nominal human does not lower the tier of a profiling system.
  • "We decided it is low risk." A tier is not a decision until it is a documented decision. Article 6(4) specifically requires a provider that calls an Annex III system not-high-risk to document the assessment and produce it to authorities on request. An undocumented "low risk" is not a classification; it is the exposure the law was written to catch.
  • "It is just internal analytics, so the Act does not care." The Act grades use and effect, not where the system sits on your org chart. Denmark's models were internal fraud analytics; their use, deciding who keeps a benefit, is what put them on the top two rungs. An internal tool that evaluates people for a consequential decision is classified by that use, not by the label "internal."
  • "The model is simple, so it cannot be high-risk." Classification grades consequence, not sophistication. A plain statistical model that decides benefit eligibility is high-risk; a frontier model that suggests recipes is minimal. Do not let "it is only a logistic regression" talk you down a tier the use case forces you up to.
  • "We bought it, so classification is the vendor's job." The provider makes the primary classification, but a deployer must know the tier to meet its own Article 26 duties and to catch an under-classifying vendor, and under Article 25 a deployer that modifies a system or puts its name on it can become the provider. Buying does not outsource the classification decision; it just changes whose primary duty it is, and modification can hand it to you.
  • "Annex III means automatically high-risk." Almost, but not quite: Article 6(3) genuinely exits a narrow-procedural, confirmatory, pattern-detecting, or preparatory system that poses no significant risk and does not profile. The mistake in both directions is real: treating every Annex III touch as automatically high-risk over-classifies harmless systems, while treating the 6(3) exit as a general hatch under-classifies dangerous ones. The exit is narrow and evidenced, and profiling closes it.
  • "High-risk does not apply until December 2027, so we can wait." The deferral (Annex III high-risk to 2 December 2027, Annex I high-risk to 2 August 2028) is runway to get compliant, not permission to ignore. You cannot prepare a system you have not classified, and the prohibitions and transparency and literacy duties already apply. Classification is due now; the obligations bite later.
  • "Prohibited and high-risk are basically the same, both mean lots of rules." They are categorically different. A prohibited system may not be used at all; no conformity file, no oversight, nothing rescues it. A high-risk system is permitted with obligations. Classifying a prohibited system as merely high-risk and building it a conformity file is building a compliant record for something that is simply banned.
  • "We classified everything last year, so we are done." Classification is a living record. A modification can change the tier (a minimal model wired into a hiring decision becomes high-risk) and the role (Article 25 can turn a deployer into a provider), and the law itself moves (the Omnibus changed the timeline). Re-run classification when a system changes purpose, when you modify it, and when the instrument changes.
  • "Reading the general Annex III area is enough." The precise sub-point carries the carve-outs. Point 5(b) credit scoring exempts fraud detection; point 5(a) benefits eligibility does not. Classifying to the area rather than the sub-point misses exactly the line the legislator drew, and a fraud model can be excepted under one sub-point and squarely high-risk under its neighbor.
  • "Over-classifying is the safe error, so when in doubt call it high-risk." Over-classification has a real cost: it drowns the scarce compliance effort you need for the genuinely high-risk systems in paperwork for harmless ones, and it trains the organization to treat the high-risk label as noise. The customer-service chatbot classified as high-risk is effort stolen from the eligibility model. Correct classification, up and down, is the goal, not maximal caution. (The one exception is a genuine borderline near the high-risk line, where defaulting up while you resolve the doubt is right, as long as you record the open question rather than freeze a guess.)
  • "One system gets one tier." A single deployed system can do several jobs, and different jobs fall in different tiers. A platform that both schedules shifts (minimal) and scores workers for promotion (Annex III point 4, high-risk) is not one tier; split it by use, classify each use, record every applicable hook, and govern the whole to the strictest. Treating a multi-use system as one tier hides its highest-risk use, which is usually the one that carries the obligations.
  • "The foundation model sets the tier." The general-purpose model layer and the use layer are separate questions. A GPAI (general-purpose AI) model carries its own obligations (owned by (see Topic 5.5)); the Article 6 tier attaches to the system and its use, not to the underlying model. Wiring a general-purpose model into a benefits decision makes that system high-risk regardless of the model's own status, and using the same model to draft marketing copy does not. Do not let "it is just the vendor's model" or "it is a powerful model" set the tier; the use sets the tier.
  • "An AI that only produces a score cannot be the decision." The SCHUFA ruling shows a court rejecting exactly this: when the human decision-maker draws strongly on the score, the score is the decision in substance (Court of Justice of the EU, Case C-634/21, 7 December 2023). A scoring system is not saved from high-risk by the presence of a human who defers to its score; that is the same fact the Article 6(3) profiling and influence tests turn on.
  • "A pilot or an internal proof of concept is not classified yet." The tier follows the use and its effect on real people, not the project's formal status. A pilot that scores real benefit recipients or screens real candidates is doing the high-risk thing during the pilot. Classify it by what it does to people now, not by "it is not on the market yet," so the safeguards are in place before the harm rather than after the launch.

Questions people ask

What is risk classification?
The analysis this topic produces for every AI system you run: sorting each into one of the EU AI Act's four risk tiers, naming the exact legal hook that decides it, recording your role, and documenting the reason, so the placement is provable and not merely asserted.
What is the four risk tiers?
The EU AI Act's risk pyramid: unacceptable risk (prohibited, Article 5), high risk (Article 6), limited or transparency risk (Article 50), and minimal risk (everything else). Obligations follow the tier, and you classify to the strictest tier a system's use triggers.
What is prohibited (unacceptable risk)?
A practice banned outright under Article 5 (for example social scoring of people, manipulative or exploitative techniques causing significant harm, certain biometric uses). For most of the list, above all social scoring, a prohibited system may not be placed on the market or used at all and no conformity file rescues it; real-time remote biometric identification for law enforcement carries its own narrow, conditioned exceptions, so read the exact sub-clause rather than assume every prohibition is unconditional. In force since 2 February 2025.
What is high-risk AI system?
A permitted system carrying the Act's full requirements because it can seriously affect health, safety, or fundamental rights. Defined by Article 6 through two doors: Article 6(1) (safety component of an Annex I product needing third-party conformity assessment) and Article 6(2) (a use in one of the eight Annex III areas). More on High-risk AI system
What is Annex III?
The list of eight areas whose use cases make a system high-risk through Door 2: biometrics; critical infrastructure; education; employment and worker management; access to essential private and public services and benefits; law enforcement; migration, asylum, and border control; and administration of justice and democratic processes. More on Annex III

Keep going