Skip to main content

Who owns the output: IP, training-data provenance, and the liability chain when AI work goes wrong

The short answer

"Who owns the output" is three questions, not one

Did the input infringe, is the output ownable, does the output infringe a third party. They have different laws, different answers, and different exposed parties. The first expert move is to say which one you are actually facing.

What you will be able to do

  • Separate the three distinct ownership questions that hide inside "who owns the output": does training on the input infringe, is the output itself protectable, and does the output infringe someone else's rights.
  • Trace the training-data provenance of a specific AI system from original source to final output, using the provenance file your organization built earlier (see Topic 2.6).
  • Analyze the liability chain of a deployed AI system, naming each party (foundation-model provider, integrator, deployer, end user) and the basis on which each could be held responsible.
  • Distinguish what a contract allocates (indemnity, warranty, ownership assignment) from what the law imposes (copyright, product liability, professional duty) and from what the evidence can actually prove.
  • Evaluate a vendor's intellectual-property indemnity for the carve-outs that decide whether it protects you when a claim actually lands.
  • Apply the reasoning of Thomson Reuters v. Ross Intelligence and the 2025 United States Copyright Office guidance on human authorship to your own organization's outputs without overstating what either one settled.
  • Produce a plain-language ownership-and-liability map for one AI system that a successor, an auditor, or opposing counsel could read and follow.

The lesson

When a cease and desist letter arrives, or a competitor copies an AI-generated asset, organizations typically scramble to find a single point of failure. In that moment, leadership focuses on one urgent question, who owns the output? Addressing the problem this way obscures the fact that AI risk is actually a chain of three distinct legal exposures. Each column carries its own laws, requires its own fixes, and targets different parties.

To resolve the dispute, you have to move past the top-line question and start a structured evidentiary process to separate those three exposures. Column 1 covers the input question, did training the model on underlying data infringe a third party's rights? Column 2 focuses on output protectability, is the generated asset legally protectable, and who holds the rights? Column 3 tracks output infringement, does the deployed AI output violate someone's rights or cause harm? Proving compliance across these columns depends on a strict evidentiary discipline, a paper trail of retained records that connect specific outputs back to their origins. Constructing this map and filling it with concrete evidence provides a defense before a court, regulator, or auditor demands one.

In Thomson Reuters v. Ross Intelligence, a federal court ruled that training a legal AI tool on copyrighted Westlaw headnotes constituted infringement. The ruling turned on the first factor of fair use, whether the use was commercial or transformative. The judge applied the reasoning from the Supreme Court's 2023 Warhol v. Goldsmith decision.

The Warhol precedent limits how companies claim transformation. Consequently, the court decided that using protected content to build a direct commercial competitor is not fair use. This specific ruling applied to non-generative AI.

The copyright rules for generative models remain unresolved, with several high-profile cases currently moving through the appeals process. Executives often misinterpret the Ross decision as a blanket ruling against all AI training. Overstating and narrow-holding misrepresents the actual state of the law.

While the U.S. courts deliberate, the regulatory standard is shifting. Article 53 of the EU AI Act requires general-purpose AI providers to publish detailed training content summaries. Because upstream data choices create downstream exposure, retrieving and filing that provenance summary is the baseline defense for any organization deploying a model.

On the output side, product teams often assume that a highly detailed prompt grants them automatic copyright ownership over the resulting image or text. In early 2025, the United States Copyright Office confirmed that purely prompt-generated output lacks human authorship. A prompt acts as an instruction to a machine.

It does not provide the creative expression required for copyright protection. Securing ownership in the U.S. requires a documented chain of human creative selection, arrangement, or modification applied after the AI produces the initial draft. Standards diverge internationally.

In 2023, the Beijing Internet Court granted copyright to an AI-assisted image by crediting the human's documented intellectual inputs during the generation process. Securing rights to a valuable AI asset requires engineering the ownership record based on the specific laws of each jurisdiction where you operate. Output infringement involves the exposure created when an AI system produces content that directly harms a third party.

There are two distinct risk profiles here. The first is regurgitation, where the model emits near-verbatim chunks of its copyrighted training data. Regurgitation originates with the vendor's data choices and requires anti-memorization safeguards built into the model itself.

The second risk is independent similarity. Here, the output resembles a protected style or mark without directly copying a specific training file. Independent similarity is a downstream problem that can only be mitigated through a human review step before the asset reaches the market.

Governance controls only work if you identify which of these two infringement types your specific system is likely to generate. Liability in an AI deployment attaches at four specific nodes, the foundation model provider, the integrator, the deployer, and the end user. The deployer is typically the most exposed entity because they are responsible for reproducing and distributing the output to the public.

Article 25 of the EU AI Act further complicates this role, allowing a deployer to absorb a provider's heavy legal obligations under certain conditions. This shift happens the moment a deployer white labels, modifies, or repurposes a high-risk AI system for a new use. Simultaneously, the revised EU Product Ledebra Lee Directive brings AI software under a strict no-fault liability regime.

Under strict liability, your internal logs, warnings, and design records become the primary evidence used to determine the organization's responsibility for a defect. Organizations often assume that standard AI vendor indemnities provide blanket protection against third-party copyright claims. These contracts contain specific carve-outs.

They may require that safety filters remain enabled, exclude free tiers, or cover only the output while ignoring the data you input. Limitation of liability caps can also shrink a massive indemnity payout down to a mere multiple of the fees you paid the vendor. Legal exposure outlives the software itself.

Ross Intelligence shut down in 2021, but the court's infringement ruling didn't arrive in 2025. An indemnity is only useful if you can produce the technical logs required to prove you met its specific conditions. Before shipping any AI-generated asset, you must execute five tactical moves.

Name the specific claim you want to make. Verify the input provenance. Define the required ownership level.

And insert a human screen for third-party harm. Finally, review your completed map to identify your chain's single weakest link, the exact point where your evidence will likely fail first. Organizations survive AI litigation by designing the evidence before the question is ever asked.

The ideas, one by one

Training-data provenance is now legal exposure, not hygiene

Thomson Reuters v. Ross Intelligence (D. Del., February 2025) showed that training on someone else's protected content, to build a competitor, without a license, can be infringement and not fair use. The EU AI Act (Article 53) now requires GPAI providers to publish a training-content summary. The provenance file you built in Module 2 is your only way to answer the input question (see Topic 2.6).

Ross decided one narrow thing; do not overstate it

It was non-generative, commercial, and directly competing. It did not rule on generative AI, and it is on appeal. Telling a board that "AI training is illegal" is a falsification. State exactly what was held and what remains open.

Prompts do not earn you copyright

Under the 2025 United States Copyright Office guidance, raw prompt output is not protectable; only recorded human creative contribution is. If you want to own your AI-assisted work, engineer and record a real human editorial step. That same step also lowers your infringement risk.

Output ownership changes by jurisdiction

The United States denies copyright to prompt-only output; the Beijing Internet Court granted it to an AI-assisted image (see Topic 10.6). A multinational maps ownership per country, not once.

Liability attaches along a chain, on more than one basis

Foundation-model provider, integrator, deployer, end user, each reachable through copyright, product liability, professional duty, or misrepresentation. The deployer who ships the output is usually the most exposed. A tool never absorbs a professional's duty of care.

Contract, law, and evidence are three layers that can disagree

Law sets the baseline; contracts allocate risk between the parties who signed; evidence decides what you can actually prove. A right or a defense you cannot document is worth little, which is why this lives in the evidence module.

Read every indemnity for its carve-outs

Vendor IP indemnities (Microsoft, OpenAI, Adobe, Google, AWS) are real but conditional: safety features on, paid tier, output claims only, capped payout. Record the conditions and whether you can prove you meet them.

Strict liability raises the value of your records

The revised EU Product Liability Directive brings software and AI under no-fault liability with claimant-friendly presumptions. When fault is no longer the question, your logs, warnings, and design records become the decisive evidence (see Topic 10.2).

Exposure outlives the system

Ross shut down in 2021; the ruling landed in 2025. Retiring a system does not retire its liability, so preserve the evidence rather than deleting it. Your finished map feeds the evidence annex and the capstone dossier (see Topic 10.6) (see Topic 13.1).

Name which kind of output infringement you fear

Regurgitation of training data is defended upstream (vendor assurances, anti-memorization safeguards); independent similarity to a protected work is defended downstream (human review, records of independent creation). Aim your controls at the right end of the chain.

The input side has an ownership question too

What your organization puts into a model can cost you IP: confidential material entered into an uncontrolled tool can lose its trade-secret protection. Govern inputs by tool and tier, and record that the policy was followed.

"We only deploy" is not a permanent shield

Under EU AI Act Article 25, renaming, substantially modifying, or repurposing a system can turn a deployer into a provider with heavier duties. Record the role you actually occupy for each system.

Code and real people are special cases

AI-generated code can carry open-source license obligations (scan it before shipping); AI output about a real person can trigger defamation, right-of-publicity, and data-protection claims that have nothing to do with copyright (never ship it without human sign-off).

You read it. Now prove it.

Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.

The conversation

The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.

Listen to it as episode 80 of the podcast.

Read the full conversation

Imagine, your company goes bankrupt today. You shut the doors, you sell off the servers, dissolve the LLC, and you send everyone home. Right.

It's totally over. Exactly. It's over.

Yeah. But then, four years from now, long after that business is literally nothing but a memory, a federal judge pulls you back into court, and they hand down a ruling that attaches a multi-million dollar liability to the ghost of your company. All because of an AI tool you deployed today.

Yes, exactly. Which is, I mean, that is the exact nightmare scenario that just played out in federal court. It proves that the consequences of poorly deployed AI systems, they absolutely outlive the organizations that built them.

Yeah. That is exactly what happened to a company called Ross Intelligence. And you know, we are kicking off this executive deep dive with their story, because I think it shatters the illusion that AI liability is just like a theoretical problem for tech giants.

Right. It's not just a Google or open AI problem. No, not at all.

For you, the executive listening to this, the person deciding how your company uses AI, this is your immediate reality. So our mission for this deep dive today is to equip you with a really precise framework to trace ownership and blame. From the original training data, all the way to the final AI-generated output.

And you really have to do this before a court or a client or, you know, a regulator forces you to do it, because the story of Ross Intelligence is basically a masterclass in what happens when you don't map that chain in advance. So true. So let's look at the mechanics of what they actually did.

Yeah. So Ross Intelligence was this startup, and they were trying to build a generative search tool for lawyers. The pitch was actually great.

I mean, instead of a lawyer spending hours hunting for case law using those clunky Boolean search terms, they could just ask the Ross AI a natural language question. Like, has a judge ever ruled on this specific contract clause? Exactly. And the AI would point them straight to the relevant passages in past court cases.

But well, to build a model that smart, you need highly structured, perfectly labeled training data. Right. You can't just feed it the raw internet.

They needed legal text that had already been digested by human experts. So they went to Thomson Reuters, which is the parent company of the massive legal database Westlaw, and they asked for a license. Yeah.

And Thomson Reuters flat out refused. Flat out. So Ross needed a plan B. They went to a third party, this legal research firm called Legally Solutions, and they purchased about 25,000, what they called, bulk memos.

And here, I mean, this is where the legal architecture just starts to completely crumble because legalese didn't just write those memos out of thin air. They built them in part by relying on Westlaw's editorial summaries of court opinions. Right.

So Ross took those memos, fed them into their machine learning pipeline, trained their model, launched their tool, and just started directly competing with Westlaw. Which is a bold move. Very.

And predictably, Thomson Reuters sued them for copyright infringement. Okay. We really have to pause here and define exactly what was allegedly copied.

Because it gets at the very heart of AI data provenance. Because they didn't just copy the law. Like the law is public domain.

Right. The judicial opinion itself, the actual 50-page ruling written by a federal judge, that is a government work. No one owns it.

You literally cannot copyright the law. Exactly. But Westlaw doesn't just publish the raw opinions, right? Yeah.

They employ all of these attorneys to read those dense rulings and write what are known as headnotes. Yes. Headnotes.

And a headnote is basically a short, highly synthesized editorial summary of a specific point of law found in that case. And that synthesis, like that editorial choice of what matters and exactly how to phrase it, that involves human creativity. It does.

Therefore, the headnote is entirely protectable by copyright. Got it. So Ross is using an AI trained on memos, which were based on headnotes, which were written by Westlaw.

Thomson Reuters sues. Ross Intelligence fought the lawsuit. But I mean, the legal fees just bled them dry.

Yeah. They officially shut their doors in January of 2021. But the liability completely outlived the company.

Because in February 2025, Judge Stephanos Beavis in a Delaware federal court handed down a ruling. It was copyright infringement. Wow.

Four years later. Four years after the company died, the judge declared that the foundational data they used basically poisoned the entire well. And that timeline.

I mean, that is what every executive needs to internalize. The company was gone, but the legal exposure remained. Which brings us to the really central theme we are unpacking today, which is liability attaches along a chain on more than one basis.

Liability attaches along a chain on more than one basis. Right. Westlaw's editors wrote the headnotes, legalese turned them into memos, Ross bought the memos and fed them to a model.

The model produced answers for end users. At every single link in that chain, ownership passed from someone to someone else or, you know, to nobody. And every link has an owner who basically answers for it when the chain inevitably breaks.

Exactly. So keep that phrase anchored in your mind. If you are deploying AI in your enterprise right now, you are sitting at the end of a very long, very complex chain.

So you know, when a product manager walks into your office, drops a new AI generated marketing campaign on your desk and asks, hey, who owns this AI output? They really think they're asking a simple single question. Yeah. And answering it as if it is a single question is honestly the most dangerous analytical error you can make in AI governance right now.

Really? Yes. Who owns the output is three questions, not one. Conflating them into a single bucket called ownership completely masks the distinct legal exposures hiding in your system architecture.

OK, let's break those three questions apart. What is the first hurdle we have to jump? Question one is the input question. Did training on the data infringe someone else's rights? This is entirely about the raw ingredients that went into the model.

Yeah. Like if a foundation model was trained on copyrighted text, images or, say, proprietary code without permission or a valid legal exception, the very act of training might be an infringement. The exposure here usually sits upstream with whoever did the training or commissioned the data scraping.

Right. So that's the input. And if we clear that hurdle, we face the second one.

Yeah. Question two is the output ownership question. Is the output itself something anyone can actually own? Like can it be protected by intellectual property laws? And this is where your exposure sits if you are the deployer.

Right. Exactly. Think about the business value.

If your flagship marketing asset or, I don't know, the core code for your new app was generated by AI and it turns out the law says no human being can own it, your competitor can legally copy it verbatim. Oh, wow. Yeah.

You have lost a massive business advantage. So if the first hurdle is how we got the data and the second is whether we even own the result, what happens if that result we don't own actually goes out and like ruins a competitor's reputation? That is question three, the output infringement question. Does the output step on someone else's rights? Like defaming someone.

Right. Does it elucidate a defamation against a real person? Does it accidentally leak a trade secret you fed it yesterday? Here the exposure sits squarely with you, the deployer who actually ships the output to the market because you are the publisher. It's kind of like asking who owns this cake.

You can't just look at the cake and say, I do. It's on my plate. We actually have to ask three totally different things to understand our liability.

I love that analogy. All right. First, were the ingredients stolen from the grocery store? Second, can you actually patent the recipe so the bakery down the street can't copy it? And third, did the final cake accidentally poison a guest? Yes.

Those are three different problems carrying three totally different liabilities, potentially involving three entirely different parties. That analogy holds up perfectly, especially when you realize that these three questions don't rank by difficulty in some fixed order. They run in sequence along a chain.

And depending on your specific AI use case, a completely different link will be the one that snaps. Give me an example of that. Well, if you're building a massive image generator for an ad agency, question two, can we own this recipe, is your biggest headache.

Because you need to protect the art. Right. But if you're building a spiralized legal or medical research tool like Ross did, question one, were the ingredients stolen, is the live wire.

And if you're building a customer-facing support chat bot, question three, did we just poison a guest with hallucinated advice, is what keeps your general counsel awake at night? Okay. That makes total sense. So let's follow that chain from the very beginning.

We start in the kitchen with the ingredients going into the bowl, the input question. And I want to push back on a historical assumption here. Because a few years ago, asking a vendor where they got their training data was basically treated as a corporate social responsibility issue.

It was a hygiene question, like, are we being good citizens? Yeah. That era is definitively over. We really need to establish a new reality today.

Training data provenance is now legal exposure, not hygiene. There is real existential financial risk attached to it. It is no longer just an ethical nice to have.

So let's look closer at the mechanics that Ross Intelligence Ruling to see exactly how that exposure plays out in a real courtroom. The court examined the 2,830 Westlaw headnotes that were at issue. And they found that 2,243 of them were effectively copied into the training memos that Ross bought.

But Ross didn't just surrender, right? They mounted a very specific defense. They argued that even if they copy the headnotes, using them to train a machine learning model is protected under fair use. Yes.

And fair use is the most vital and honestly, arguably the most misunderstood copyright defense in the United States. Oh, for sure. It permits the unlicensed use of copyright protected works under certain specific circumstances, like commentary, search engines or parody.

When a judge evaluates a fair use defense, they are legally required to weigh four specific factors. The purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used, and finally, the effect of the use upon the potential market. OK, so how does a federal judge actually balance those four factors when dealing with AI training data? Because that feels unprecedented.

In the Ross case, it was a split decision, but it was not an equal one. Let's look at what Ross won first. Factor two evaluates the nature of the copied work.

Like are we talking about a highly creative novel or a factual list? The judge found this favored Ross because legal headnotes, while protectable, are heavily factual. They just summarize real events. OK, that makes sense.

Then factor three looks at the amount used. This also favored Ross because the copied material didn't actually appear in Ross's public facing output. It stayed entirely behind the scenes, basically locked inside the intermediate training phase.

Wait, wait. If the copied text never even made it to the end user, if it was just intermediate processing, how is that infringement at all? That seems like a massive technicality. I get why you'd say that.

But intermediate copying is a very well established concept in copyright law, particularly in software. Even if the final product doesn't contain the copied code, the act of making a temporary copy in your RAM or on your hard drive to build the competing product still violates the reproduction right. Oh, I see.

You made a copy without permission. That is the fundamental trigger. OK, so Ross is winning on factors two and three.

They used factual data and they kept it behind the scenes. Why did they lose the entire case? Because they lost factors one and four, and in the current legal landscape, those carry the heaviest weight by far. Factor one asks if the use is commercial and, crucially, if it is transformative.

Transformative. Meaning what, exactly? Did you turn the original work into something with completely new meaning or purpose? The judge, relying heavily on a recent Supreme Court precedent involving Andy Warhol, found that Ross's use was highly commercial and completely non-transformative. OK, hold on.

How does Andy Warhol fit into an AI legal research case? It's fascinating. The Supreme Court recently ruled that when Andy Warhol took a photographer's portrait of the musician Prince and turned it into a silkscreen print, he wasn't transforming the purpose of the photo. Because it was still just a picture of Prince.

Essentially, yes. Both the photo and the silkscreen were serving the exact same commercial purpose, illustrating a magazine article about Prince. So the judge in the Ross case applied that exact logic.

Ross didn't transform the purpose of the headnotes. They used Westlaw's legal summaries to build a tool that directly competed with Westlaw to provide legal summaries to lawyers. It was the exact same commercial purpose.

Ah, and that failure to transform the purpose bleeds directly into Factor 4, I assume. Precisely. Factor 4 analyzes the effect on the potential market for the original work.

The judge literally called this undoubtedly the single most important element of fair use. Right. Follow the money.

Exactly. Ross, his tool didn't just threaten Westlaw's core subscription market, it also threatened a brand new emerging market, the market for licensing data for AI training. By taking the data without paying, Ross harmed Thomson Reuters' ability to sell that data to other AI developers.

So because Factors 1 and 4 went against Ross, the entire fair use defense collapsed. I need to jump in here with a massive warning for you, the listener. When you take this specific ruling back to your general counsel or your board, you have to be surgically precise.

Ross decided one narrow thing. Do not overstate it. Oh, right, sure.

Right. Like, do not walk into a board meeting and say, hey, a federal judge just ruled that all AI training is a legal copyright infringement. That is a dangerous falsification of the law.

That distinction is paramount. You have to look at the architecture of the system Ross actually built. Their AI was non-generative.

It was an extractive search system. It analyzed a user's question, searched a database, and retrieved existing text. It did not generate entirely new expressive paragraphs of text from scratch the way a large language model does.

The judge explicitly noted that this ruling does not bind the massive generative AI copyright cases that are currently working their way through the federal courts. Yeah, we should definitely mention that the Ross case isn't even fully settled law yet, is it? Like, it is currently on interlocutory appeal to the Third Circuit. But if Ross is the non-generative benchmark, what is the generative counterpart? What's the anchor case we should be tracking to understand how foundation models like ChatGPT or Claude will be judged? The anchor case for the generative side is the New York Times v. OpenAI and Microsoft.

The Times is suing over the unlicensed scraping and training on millions of their journalism articles. They're basically alleging that these generative models aren't just learning abstract concepts. They are capable of essentially reproducing their articles verbatim if prompted correctly.

Yeah, it's a massive, complex piece of litigation, and it is still open and unresolved. But honestly, regardless of who wins, it proves our central point for the enterprise deployer. A foundation model providers upstream data choices create findable massive exposure for downstream deployers.

But it isn't just federal judges making these rules retrospectively, is it? Because we are seeing global legislatures step in proactively to force transparency. Oh, absolutely. We are in the middle of a massive global legal shift, and the epicenter is the European Union.

If you look at the EU AI Act, specifically Article 53, Paragraph 1 of Regulation 2024-1689, it fundamentally rewrites the rules of engagement. This regulation places a direct, unavoidable transparency duty on providers of general-purpose AI or GPAI. Let's define GPI technically real quick.

A general-purpose AI model is basically a foundation model trained on a broad volume of data at scale, typically using massive self-supervision, that can perform a wide range of distinct downstream tasks rather than, say, a narrow AI trained to do one specific thing like detect tumors in an X-ray. That's exactly right. And under Article 5-3-1, providers of these GPAI models must publish a sufficiently detailed summary of the content used for training.

So they can't just be vague. Right. They don't just get to write a blog post saying, oh, we use the open internet.

They have to use a mandatory, standardized template that the EU AI Office published in July of 2025. And this compliance duty takes effect in August 2025. This isn't just some voluntary safety standard.

It is written into the law of the world's largest single market. OK, let's unpack this for the executive listening right now. If I'm running a logistics company or a marketing firm, I am not training my own foundation models from scratch.

It costs $100 million in compute. I am just buying an enterprise API license to use a major vendor's GPAI model. So why is this my problem? Why do I care if the vendor fills out an Article 53 template in Brussels? Because like we said, the input question rolls downhill and gravity always wins.

Gravity always wins. Always. You didn't choose the training data.

You didn't scrape the internet. But you chose to ship the output of that model to your end customers. If that upstream model was trained on infringing data, the output you ship may carry that lead legal taint.

Makes sense. And when a major rights holder, say, a massive stock photo agency or a consortium of software developers looks for someone to sue, they don't always go after the AI vendor. They look for the party actually distributing the output and making money off it.

You are the findable solvent target. So the transparency mandate is actually a weapon for us to use in our own defense. Exactly.

That Article 53 training summary is a document you are legally entitled to request and read when you buy the model license. If you don't ask for it, file it in your compliance registry and check it against your specific use case. Your defense of we assumed the vendor had the rights will completely evaporate in court.

Willful blindness is not a valid legal defense. Wow. Okay, that perfectly encapsulates the input question.

The ingredients matter and you are basically responsible for checking the label. Well said. So let's move down the chain.

The data has been processed, the prompt has been sent, and the model has spit out a result on your screen. It's a new corporate logo, a block of back-end code, a massive marketing strategy. We arrive at question two.

The output ownership question. Can you actually own this recipe? And this is where we have to deliver some very hard truths to product teams. Do it.

In the United States, the fundamental rule you must operate by is completely unyielding. Prompts do not earn you copyright. I can just imagine the marketing directors face when legal tells them their multi-million dollar Super Bowl ad campaign is public domain because the creative team used mid-journey and didn't save the proms logs.

They will scream, but I wrote a 500 word, highly detailed prompt. I specified the lighting, the camera angle, the exact color palette. It took me an hour of iterative engineering.

Oh, I hear that exact argument from creatives every single day. It feels like authorship because it requires effort, right? But in January 2025, the U.S. Copyright Office issued part two of its official binding guidance on AI, and they were unequivocal. What did they say? The Copyright Office views a purely prompt-generated output as an instruction to a machine, basically akin to telling a commissioned artist what you want them to paint.

The prompt is the idea, but copyright only protects the specific, final expression of the idea. So the machine is executing the expression, not the human. Exactly.

When you prompt a model, the AI determines the traditional elements of authorship. The exact placement of the pixels, the specific flow of the syntax in a paragraph. To earn copyright protection in the U.S., there must be a genuine, recorded, human, creative contribution to the final expressive output.

What does that mean in practice, though? Give me a real-world example of where this line is actually drawn. Let's look at the famous State Fair Prize AI image. An artist named Jason Allen won a prize with a breathtaking AI-generated image called Teatro Adopera Spatial, and he tried to register the copyright.

He argued he spent hundreds of hours refining the prompts and tweaking the output. And what did they say? The Copyright Office refused the registration flat out. They ruled that because the final expressive elements, the actual pixels forming the image, were generated by a machine, the work totally lacked human authorship.

But there is a way to blend human and AI work, right? Yeah. What happens if you use the AI image, but you build something larger around it? Ah, that brings us to Christina Kashinova's comic book, Zarya of the Dawn. This case is basically the perfect blueprint for understanding the boundaries.

How so? Well, Kashinova used mid-journey to generate the images for the comic book, but they didn't just hit generate and publish. They wrote the overarching narrative text, they designed the page layouts, and they arranged the AI images in a specific sequence to tie a story. So they applied human authorship to the arrangement, but not the wrong images.

How did the Copyright Office handle that one? They narrowed the registration. It was a very surgical decision. The Copyright Office granted copyright for the human-authored text.

They granted copyright for the human compilation and arrangement of the pages, but they explicitly denied copyright for the raw AI images themselves. Yeah. So if you take one of the mid-journey images out of that comic book and put it on a t-shirt, Kashinova cannot sue you for copyright infringement because they don't own the raw image.

That is fascinating. And this framework doesn't just apply to creative arts like text or images, right? It applies to hardcore industrial applications too. It applies to inventions and patents.

It does. And the states there are arguably much higher. In a landmark 2022 case, Thaler v. Vidal, a computer scientist named Stephen Thaler tried to list an AI system he created, named Debus, as the sole inventor on a patent application for a new type of food container.

The AI was the inventor? Yes. The U.S. Patent and Trademark Office rejected it. Thaler sued, arguing that AI can independently invent.

How did the federal courts react to that? The federal circuit held that under the plain text of the U.S. Patent Act, an inventor must be a human being. The statute literally uses words like individual and himself or herself. The Supreme Court even declined to hear the appeal, cementing the ruling.

Okay, the implications of that for the pharmaceutical and material science industries are just staggering. If your enterprise deploys an AI system that independently discovers a revolutionary new chemical compound, and there is no human inventor to legally claim it, that compound might be entirely unpatentable. Your massive R&D investment just became a gift to the public domain.

Exactly. It's a huge risk. This sounds brutal for any company trying to build an asset base using AI.

So in the U.S., if we just use a prompt and ship the output, we own absolutely nothing. But is that the rule everywhere? Like, if I'm a global enterprise, is this a universal law of gravity? No. And this is a crucial pivot for your international strategy.

Output ownership changes by jurisdiction. The philosophy of human authorship is not some single, harmonized global rule under treaties like the Breening Convention. It varies drastically depending on local judicial interpretation.

Give me an example of a major market where the rule flips entirely. Look at China. In November 2023, the Beijing Internet Court ruled on a copyright infringement case involving an AI-generated image of a young woman.

The facts were very similar to the U.S. cases. A human user spent time iterating prompts and adjusting parameters in a text-to-image model. But the Beijing court reached the exact opposite conclusion of the U.S. Copyright Office.

They actually granted copyright in the image to the human prompter. Really? How did they justify that legally? They reasoned that the prompters documented intellectual inputs like the specific selection of negative prompts, the iterative refinement of the seed weights, the deliberate artistic choices to achieve a specific aesthetic. They said all of that constituted original human authorship.

They viewed the AI not as an independent creator, but merely as a highly sophisticated tool wielded by the human. No different than a camera. So if you are running a multinational enterprise, your legal strategy has to be geographically fragmented.

You cannot just walk into a global town hall and tell your entire workforce, we don't own our AI assets. In your Chinese operations, you might own them completely. In your U.S. operations, you own nothing unless you actually intervene.

Precisely. A single global answer offered with total confidence by a consultant is usually a wrong answer in at least half of your operational markets. So what does this actually mean for the executive trying to operationalize this? If I am building a software platform in the U.S. and my engineers are using AI to write boilerplate code, how do I secure my IP? What is the actionable insight here? The actionable insight is that output ownership is something you engineer into your workflow.

It is not something you argue for after the fact. You cannot wait until the output is generated, slapped onto a product, and challenged by a competitor, and then try to argue really hard with the copyright office that your prompt was super creative. It's too late by then.

Way too late. You have to design a recorded, human editorial step into your production pipeline before you generate the final product. You need tangible, provable evidence of human selection, human arrangement, or human modification.

If your lead designer uses an AI tool to generate 50 conceptual variations of a logo, selects three of them, composites them together in Photoshop, paints over the lighting, and adds proprietary typography, that is undeniable human authorship. Yes, exactly. That final asset is protectable.

But, and this is the critical failure point for most enterprises, it is only protectable if you keep the receipt. Keep the receipt. You must log the entire workflow.

Who made the creative choices? What exactly were those choices? When did they happen? What did the raw AI output look like before the human touched it? That provenance record is the only thing that turns the abstract claim of AI helped us into the legally defensible claim of a human authored these specific protectable parts. Okay, that comprehensively covers question two. If you don't engineer that human step, you don't own the output, which means you lose a potential business asset.

It's bad. It hurts the balance sheet. But fundamentally, it's just a lost opportunity.

Right. But question three, question three is a different beast entirely. This is where your AI doesn't just fail to be an asset, it actively harms someone else.

This is the output infringement question. And this is where we have to map out the liability chain, because this is where the lawsuits actually happen. To understand who pays the damages when things go wrong, we have to formally name the players in the ecosystem.

Okay, let's list them out so we have a shared vocabulary. First, at the top of the chain, you have the foundation model provider. This is the tech giant, OpenAI, Google, Anthropic, Meta.

They spent hundreds of millions of dollars to scrape the data and train the massive base model. Makes sense. Next, you have the integrator or fine tuner.

This is a middle vendor, or sometimes it's an internal engineering team, who takes that base model and adapts it. They might fine tune it on specific industry data, host it on a private cloud, or wrap it into a specialized SaaS product. Third, you have the deployer.

Let's define deployer really clearly, because this is the bullseye for our audience. The deployer is the organization that puts an AI system into operational use and ships its outputs to the market, to clients, or to employees. If you are listening to this deep dive to figure out your corporate strategy, you are almost certainly the deployer.

Exactly. And finally, at the very end of the chain, you have the end user. The individual employee or customer actually typing the prompt into the interface.

Okay, got the players. Now here is the legal reality that most people completely miss. Liability doesn't just pass cleanly down this chain like a hot potato until it lands on one unlucky person.

It can, and often does, attach at multiple links simultaneously on totally different legal bases. So when we talk about output infringement from the deployer's perspective, what exactly are we talking about? How does the AI step on someone's rights in a way that actually gets us sued? There are two distinct technical categories of output infringement, and you absolutely must know the difference because your legal defenses for each are totally different. What's the first one? The first category is regurgitation.

This is exactly what it sounds like. This is when the model emits near verbatim chunks of its original training data. It spits out an exact copy, or a derivative so close it is practically identical, of a protected work.

Like the New York Times alleging that if you prompt a model with the first sentence of an exclusive investigative article, the model will spit out the next five paragraphs word for word, completely bypassing their paywall. Exactly. Regurgitation ties the output directly back to the input question.

The copied material was memorized in the training set, therefore this is largely an upstream vendor problem. The model memorized something it shouldn't have. As a deployer, your defense against regurgitation claims relies heavily on vendor assurances, indemnification contracts, and ensuring you have actually turned on the anti-memorization safeguards built into the API.

And the second category. If it isn't regurgitating memorized data, how is it infringing? The second category is independent similarity. And this is far more insidious.

This is when the output wasn't lifted directly from any single piece of training data, but the mathematical combination of concepts lands so close to a third party's protected work that it violates their rights. Can you give an example? Sure. For example, here, it might generate an image that perfectly mimics a living artist's unique trademark style.

Or it might generate a brand name that causes massive market confusion with an existing trademark. Ah, so it didn't copy a specific file, but the end result looks exactly like a copy. Right.

And this is exclusively a downstream deployer problem. Your upstream vendor cannot perfectly prevent this because the model didn't memorize a file. It generalized a concept way too well.

So we hold the bag. You do. Your only defense against independent similarity is a robust human screening and clearance check before you ship the output to the public.

If you mix up these two risks, like if you rely on vendor safety filters to prevent trademark infringement, or if you blame your human reviewers for not catching verbatim regurgitation of a random article they've never even read, you will end your safety controls at the completely wrong end of the liability chain. Wow, that's critical. Okay, we've been talking heavily about copyright and IP, which is usually just a financial penalty.

But here is where it gets incredibly serious. We aren't just talking about intellectual property anymore. We're talking about physical harm, property damage, or severe financial ruin caused by relying on AI.

And the European Union has just fundamentally rewritten the rules of gravity on this. They really have. In November 2024, the EU published the final text of the revised product liability directive.

Member states have until December 2026 to transpose this into their national laws. This directive does something completely revolutionary. It explicitly, for the first time, brings software and AI systems within the legal definition of a product.

Why is that a big deal, though? Software has been running the world for 30 years. Why does calling it a product suddenly change our liability? Because in European law, classifying something as a product subjects it to a regime called strict liability. Let's break down strict liability mechanically, because this concept terrifies software developers.

How is it different from the way software liability normally works? Under a normal fault-based legal rule, which is how software has historically been treated, if someone sues you because your software caused them harm, they have to prove that you were careless. They have to prove negligence. They have to show that your engineers failed to meet an industry standard of care.

Right, like you coded this badly. Exactly. Under strict liability, negligence is removed from the equation entirely.

The claimant only needs to prove three things. That the product was defective, that they suffered damage, and that there is a causal link between the defect and the damage. Your carefulness as a developer is entirely irrelevant.

Wait, let me make sure I understand this. If we spend $10 million on red teaming, safety alignment, and rigorous testing, and we follow every single industry best practice, if the AI still hallucinates a dosage recommendation in a medical app and harms a patient, we are liable. Yes.

It does not matter how hard you tried to make the AI safe. If it was objectively defective and caused physical or severe material harm, you are on the hook. And the revised directive goes even further.

How could it go further? It introduces a mechanism to presume defectiveness. If the AI's technical complexity, the whole black box nature of a neural network, makes it excessively difficult for a consumer to prove exactly how the defect occurred, the court can legally presume the AI was defective. The burden of proof shifts to you, the deployer or provider, to prove it wasn't defective.

Oh my god. That shifts massive, almost incalculable exposure onto providers and deployers. If your carefulness doesn't matter, and the court presumes the software is defective because it's too complex to understand, what is your defense? How do you even survive that lawsuit? Your technical logs.

When fault doesn't matter, evidence is literally everything. The pristine, immutable record of exactly what the system did, what version of the model weights were running at that exact millisecond, what the user prompted, and exactly how the user was warned about the system's limitations becomes your only shield to prove that the harm was caused by user misuse, not a system defect. We also have to talk about how this intersects with professional duty.

Let's look at a scenario we've actually already seen play out in real life. A lawyer uses an AI chatbot to draft a legal brief for a federal court. The AI hallucinates beautifully formatted citations for cases that do not exist.

The lawyer submits the brief, the judge realizes the cases are fake, and all hell breaks loose. Who gets sanctioned by the judge? Does the judge haul the AI vendor into court? Absolutely not. The lawyer faces the sanctions.

We have seen this happen repeatedly in New York and other jurisdictions. A technological tool never, under any circumstances, absorbs a human professional's legal duty of care. That makes sense.

Yeah. If you are a doctor, a structural engineer, a certified public accountant, or a lawyer, the state has granted you the professional license to practice. The AI does not hold a license.

If you rely on an AI output without conducting the independent verification check that your professional duty requires, the liability for professional malpractice or negligence lands squarely on you. The tool is just a tool. Okay, let me play the role of the frustrated executive here for a second.

Go for it. The law says I'm exposed as a deployer. My professionals are exposed to malpractice.

Strict liability is waiting in Europe. This sounds like an unmanageable nightmare. But wait, I have a massive legal department.

Can't we just sign a contract to make the vendor pay for all of this? Like if I sign a massive multimillion-dollar enterprise agreement with Microsoft, OpenAI, or Adobe, and they give me an ironclad IP indemnity, aren't I safe? I've transferred the risk. This is the most dangerous assumption in corporate AI adoption right now. To understand why you aren't perfectly safe, we have to look at the three layers that decide who actually pays when the lawsuit lands.

You are focusing entirely on layer two, but you have to understand how all three interact. Okay, break down the three layers. Layer one is what the law imposes.

This is copyright law, the EU product liability directive, consumer protection statutes. This sets the non-negotiable baseline. A vendor's contract cannot waive away a citizen's statutory rights.

If your AI defames someone, they sue you based on layer one. Got it. Layer two is what the contract allocates.

This is the master service agreement between your enterprise and your AI vendor. Through indemnities, warranties, and liability caps, you attempt to legally move the financial risk upstream back to the vendor. Which sounds great.

It does, until you hit layer three. Layer three is what the evidence can prove. And this is the layer where almost every single enterprise fails.

A generous contractual promise in layer two is completely useless if you lack the technical evidence in layer three to prove you actually qualify for that promise. Okay, let's drill into layer two, because these indemnities are the major selling point right now. Every major vendor offers one to get enterprise adoption.

OpenAI has Copyright Shield, Microsoft has the Copilot Copyright Commitment, Adobe has the Firefly indemnity. Let's use our cake analogy again. An indemnity is essentially the vendor promising to pay the hospital bill if the cake they helped you make ends up poisoning a guest.

So if I have that written promise, why aren't I safe? Because you have to hunt for the carve-outs. Carve-outs. Yes.

A carve-out is a specific condition buried in the contract that, if triggered, entirely voids the vendor's promise to pay the hospital bill. What do these carve-outs actually look like in these AI contracts? They are very specific and heavily engineered by the vendor's lawyers. First, you almost always must have used the vendor's built-in guardrails.

If your engineering team tweaked the API settings to turn off the content filkers, or bypass the safety mitigations because they wanted less restricted or more creative output for a marketing campaign, the indemnity evaporates instantly. Ouch. So, if a developer flips a toggle to false in the code, our multi-million dollar legal shield disappears.

What else? Free and preview tiers are frequently excluded. You usually have to be paying for a generally available enterprise-tier license. If an employee is using a free beta version of a new model, there is no indemnity.

Makes sense. Third, the claim usually must be about the output, not your inputs. If your employee feeds the model infringing material, like a copyrighted article, and the model predictably spits out a summary of that infringing material, the vendor won't cover you.

They're indemnifying the model's behavior, not your bad inputs. Okay. Any others? And finally, you have territory and use case limits.

The indemnity might only apply if you use the AI for the specific use cases outlined in your terms of service, and only if the lawsuit is filed in certain approved jurisdictions. So if an infringement claim lands on my desk, I don't just casually waive the contract at my general counsel and say, oh, Microsoft is paying for this. I actually have to prove to the vendor that my engineers didn't disable the safety filter six months ago.

Precisely. Okay. This is what we call the evidence imperative layer three.

Let's return to the cake. The vendor promised to pay the hospital bill, but the carve-out says they only pay if we bake the cake at exactly 350 degrees. When the guests get sick, it's not enough to say, we always bake at 350.

You have to produce the oven's immutable digital temperature logs for that exact baking session. Wow. Okay.

If your indemnity requires the safety filters to be on, and your IT system doesn't generate unalterable technical logs proving they were actively engaged during the exact session that produced the disputed output, you don't have an indemnity, you just have a comforting sentence on a piece of paper. That is a massive operational gap for most companies. But there's a silent killer in these contracts too, isn't there? Something even worse than failing a carve-out.

Yes. The limitation of liability cap. Right.

Let's walk through the math on this one. Let's say my marketing team generates a logo using an enterprise AI tool. We clear all the hurdles.

The filters were on. We have the logs. We didn't feed it bad inputs.

We put that logo on our global packaging. A month later, it turns out the AI hallucinated a design that perfectly infringes on a major brand's trademark. They sue us, and we have to pull millions of products off the shelves.

Our total financial exposure is $10 million. Okay. Okay.

Bad day. Very bad day. But we didn't violate any carve-outs.

The vendor's model is entirely at fault. Will the vendor write us a check for $10 million? Almost certainly not, because somewhere in that master service agreement, completely separate from the indemnity clause, is a limitation of liability clause. It sets a hard mathematical ceiling on what the vendor will pay out, regardless of fault, regardless of the indemnity.

Really? Very often, especially in standard enterprise agreements, that cap is tied to your spend. It is limited to fees paid in the preceding 12 months. Let me do the math.

If my annual enterprise license fee for that specific AI tool is $50,000. The vendor's maximum legal payout for your $10 million disaster is exactly $50,000, even if your real exposure is 200 times that amount. An indemnity sitting underneath a tiny liability cap is a very generous promise about a very tiny number.

It transfers almost none of your real systemic risk. This is exactly why mapping this entire chain is so critical for an executive. We've covered the standard chain from the input data to the output liability, the strict liability waiting in Europe, and the contractual illusions of safety.

But there are some quiet, almost invisible ways you can accidentally multiply your exposure without even realizing it. Let's talk about the special risks. Yeah, there are four special cases where the standard risk profile fundamentally shifts and your traditional IT governance will just completely fail to catch them.

You need to govern these entirely differently. Okay, the first special risk is inputting trade secrets. We've spent a lot of time talking about the liability chain running downwards from the vendor's training data down to our outputs, but the chain runs both ways.

What happens when an employee takes our proprietary, highly confidential Q3 acquisition strategy or a massive block of our unreleased proprietary source code and pastes it into a consumer tier AI chatbot just to summarize it or debug it? Two catastrophic things happen. First, depending on the specific tool's terms of service, that input might be ingested to train or improve the vendor's future models. Your secret has literally left your secure perimeter and entered a third-party neural network.

That's a nightmare. It gets worse. The second issue is much deeper, and it involves the legal definition of a secret.

Under laws like the U.S. Defend Trade Secrets Act, a piece of information is only legally protected as a trade secret as long as the owner takes reasonable measures to keep it secret. So if you willingly paste it into a third-party web server that you do not control, governed by terms of service that allow data reuse. You may have legally destroyed the trade secret status of that data entirely.

A federal judge might look at that action and say, you didn't treat this information as a secret, so the court will not protect it as one. If your competitor subsequently gets ahold of that strategy or code, you might have absolutely no legal recourse to stop them from using it. That is terrifying.

Yeah. Just by trying to save five minutes summarizing a document, an employee can destroy millions of dollars of intellectual property value. Okay, what is the second special risk? AI-generated code.

Generating software code carries a unique hazard that generating text or images just does not. It introduces the risk of open-source license contamination. Specifically, it triggers copyleft obligations.

Let's define copyleft for the non-technical executives. What does that mean? Open-source code is a wonderful resource, but it isn't just free to do whatever you want. It comes attached to specific legal licenses.

Some licenses, like MIT or Apache, are very permissive. But copyleft licenses, like the famous GNU General Public License or GPL, are viral. Yes.

A copyleft license legally requires that any new software that incorporates or derives from that open-source code must also be released to the public under that exact same open license. Oh boy. So, walk me through the nightmare scenario here.

My developers are under a tight deadline. They are using an enterprise AI coding assistant. The AI assistant was trained on millions of repositories on GitHub, including GPL license code.

The AI regurgitates a 10-line block of that copyleft code into my developer's IDE. And your developer, assuming the AI generated it for scratch, integrates that 10-line block into your company's proprietary, closed-source, flagship software platform. By incorporating that viral code, you might unknowingly be legally forced to open-source your entire proprietary codebase.

You might have to publish the source code that runs your entire business for anyone, including your competitors, to download and use for free. How do we possibly stop that? We can't ban AI coding assistants. If we ban them, the developers will just use them secretly on their personal phones because the productivity gains are too massive.

You are exactly right. Banning them is a failed strategy. You don't ban the generation.

You govern the integration. You must scan the outputs. Scan them how? Running software composition analysis, or SCA, on all AI-generated code before it is allowed to merge into your production codebase is completely non-negotiable.

An SCA tool acts like a spellchecker for legal licenses. It scans the code, compares it against a massive database of known open-source code, and flags those viral license matches before they contaminate your proprietary product. Okay, third special risk.

Real people and outputs. When an AI output depicts, describes, or speaks about a real, identifiable human being, you step completely outside of copyright law and trigger a total minefield of personal rights. Give me an example.

If your marketing AI fabricates a false, damaging quote and attributes it to the CEO of a rival company, you face a massive defamation claim. If it generates a synthetic voice or a deepfake likeness of a celebrity to endorse your product without permission, you trigger right-of-publicity laws. And if it processes identifiable data about European citizens, you immediately trigger data protection laws like the GDPR, which carry fines of up to 4% of your global revenue.

And as we established earlier, the AI hallucinated it and we didn't notice, is not a legal defense. Absolutely not. You publish the output to the world, you are the publisher, and you are fully liable for the harm it causes to real people.

Which brings us to the fourth and perhaps most complex special risk, role shifting under the EU AI Act. Yes. If we connect all of this to the bigger regulatory picture, the EU AI Act draws a very hard fundamental line between a provider and a deployer.

The provider is the entity that develops the system. They carry the heaviest, most expensive compliance duties, risk assessments, conformity declarations, quality management systems. The deployer is the party that just uses the system under their own authority.

They have much lighter obligations. Right. So as an enterprise, we are usually just deployers.

We pay the license fee and let OpenAI, Anthropic, or Microsoft carry those heavy, expensive provider duties. We just want to use the tool. Usually yes, but the law contains a trapdoor.

Under Article 25 of the EU AI Act, a deployer can legally transform into a provider, instantly inheriting all of those massive, expensive compliance duties if they take certain actions with the AI system. Give me an example. What action triggers that transformation? Well, if you take a high-risk AI system from a vendor and put your own trademark on it, white labeling it as your own proprietary creation, you become the provider in the eyes of the law.

Or, more commonly, if you make a substantial modification to the system. Let's define substantial modification precisely. Under Article 3, Paragraph 23 of the Act, a substantial modification is a post-market change that affects the system's compliance with the law or changes its intended purpose in a way that makes it high-risk.

Okay, so if my engineering team adjusts a basic prompt setting. If your engineering team adjusts a basic prompt setting or a temperature parameter, that is fine. You remain a deployer.

But if you take an open-source foundation model, heavily fine-tune the model weights on your own proprietary dataset of customer financial records, and deploy it to make automated credit decisions a high-risk use case. You didn't just deploy a tool. You built a new system.

So we become the provider. You might have just stepped into the provider's shoes. We only deploy is not a permanent legal shield.

It is a conditional status that you can accidentally void. Okay, let's bring all this abstract legal theory down to execution. We've mapped the horror stories.

We know the three questions. We understand the liability chain, the carve-outs, the strict liability, and the special risks. So what does this all mean for the executive listening today? How do we actually map our enterprise systems, govern this chaos, and protect our balance sheet? We need to borrow a discipline from two very different industries, Hollywood and commercial real estate, and apply it to AI data provenance.

It's a legal concept called chain of title. I love this concept. In the film industry, no major studio or distributor will release a movie unless the producer hands them a massive binder proving they have the chain of title.

It is an unbroken chain of legal ownership. You need the rights to the underlying book, the screenwriter's contract, the actor's image releases, the music licenses, a perfect unbroken chain. If you are missing one single release from a background actor, it can kill the distribution of a $50 million movie.

Exactly. Basically, training data provenance and output governance is simply chain of title for AI. You need to formally log the sequence showing where every piece of data came from and on what legal basis it was used in your system.

What specifically do we record? You need to record the source of the data. You need to record the legal basis. Was it owned by us, licensed from a vendor, scraped from the public domain, or generated synthetically? Yeah.

You need the date you recorded it, and you need the actual verifiable evidence. And we are starting to see technical standards emerge to automate this, right? It's not just lawyers with spreadsheets anymore. We have standards like C2PA for content metadata and SPDX for software.

Yes. C2PA stands for the Coalition for Content Provenance and Authenticity. It is an open technical standard that securely binds metadata to a piece of digital media, like an image or a video.

Oh wait, I need to interrupt. A securely bound metadata file sounds like a blockchain buzzword that a startup pitches to get funding. How does that actually prove to a federal judge that I didn't steal an image? Why can't a pirate just delete the metadata? Because it is cryptographically hash-linked.

When the AI generates the image, the C2PA standard creates a digital manifest detailing how it was made, who made it, and what tools were used. It then uses cryptographic hashing to bind that manifest to the image file itself. If anyone alters the image or tries to strip the metadata, the cryptographic hash breaks.

It becomes mathematically evident that the file was tampered with. In a courtroom, a mathematically tamper-evident metadata file survives a legal dispute infinitely better than a paragraph written by a product manager two years ago saying, we think we have the rights to this picture. Ah, okay.

That is brilliant. Right. And SPDX, the Software Package Data Exchange, does the exact same thing for code.

It creates a standardized software bill of materials, an SBBOM, so you know exactly which open source licenses are buried inside your compiled software. Okay, so we have the tools. Give us the playbook.

What is the concrete decision framework, the checklist, that an enterprise team should run before they launch any AI-generated work to the market? I break this down into five pre-ship moves. You need to integrate these as a mandatory gate in your product review cycle. Okay, let's hear them.

Move one, name the output and the claim. State exactly what the AI produced and what legal claim you want to make about it. If you want to claim this is our exclusive, copyrighted brand logo and no one else can use it, you need a mountain of evidence of human authorship.

If you state these are internal draft emails that nobody owns and they delete after 30 days, you need very little evidence. The strength of the legal claim you want to make sets your evidence burden. Makes complete sense.

Move two. Check the input. Where did the training data come from? Do you have provenance? If you are deploying a vendor's GPAI model, request that EU article 53 training content summary.

File it in your compliance system. If you built the model internally, run a chain of title audit on the data set. If you simply don't know where the data came from, state it clearly in your risk registry.

Input risk unquantified. Do not pretend it is safe. Love that.

Move three. Decide the ownership you actually need. Okay, I have to push back on this one.

You're telling me an enterprise engineering team of 5,000 people needs to log every single human keystroke, every iterative prompt, just to prove copyright. They would spend more time logging their workflows than actually coding or designing. It sounds impossible to scale.

You are right. Logging everything is impossible. That is why move three is about deciding what you need to own, not trying to own everything.

Ah, I see. If you are in the US and the asset is your core intellectual property, your flagship software, your global brand identity, then yes, you must engineer that human editorial step and log the workflow rigorously. But for 90% of what AI generates, internal memos, meeting summaries, generic boilerplate code, you don't actually need to own it.

Except that it's public domain. Except that it is unprotectable. And stop wasting time and money trying to engineer a chain of title for a disposable asset.

Triage-ing your assets is the key to scale. That makes perfect sense. Triage the risk.

Move four. Screen the output for third-party harm. This is how you prevent the strict liability lawsuits.

Put a mandatory human review step between the raw AI output and the public market. Size the rigor of the review to the risk of the deployment. High stakes outputs like professional legal advice, medical summaries, or public brand assets get rigorous, multi-layered human clearance checks.

Internal drafts get a light review. But crucially, you must securely log that the human screen actually happened. And the final move? Move five.

Line up the backstop and name the weak link. Before you ship, confirm exactly which vendor indemnity applies to this specific use case. Confirm using technical logs that you currently meet all the specific carve-outs, like the safety filters being engaged.

Check with your risk department to see if your current cyber insurance actually covers this specific AI risk, or if AI hallucinations are excluded from your policy. And finally, figure out where the chain breaks first. I want to hammer that last point home.

Every ownership and liability map you build for a new system must end with one brutal honest sentence from your team. The exact place the chain is going to break first. Is it the missing provenance file for the training data? Is it the thin human contribution that will cost you the copyright registration? Is it the unprovable indemnity condition because your IT department isn't actively logging the status of your safety filters? You must name the weakest link because you would much rather find it internally today than have opposing counsel expose it in a deposition three years from now.

That entire five-step process is the essence of evidence engineering. You are not just building a product. You are proactively building the evidentiary record your lawyer will need when the dispute inevitably arrives.

Let's synthesize this journey. We started with the realization that asking who owns the output is a trap. It is actually three distinct sequential questions.

Input infringement, output ownership, and output infringement. We established that training data provenance is no longer a corporate hygiene issue. It is a live existential legal exposure that outlives companies.

We learned the hard truth about U.S. copyright. Prompts alone do not earn you ownership. You must have documented genuine human creative contribution.

We saw that massive strict liability is waiting downstream in Europe, making immutable technical logs your only defense against presumed effectiveness. And we uncovered the reality of vendor contracts, the carve-outs, and liability caps that demand hard, unalterable evidence before they ever pay out a dime. We have covered a massive amount of ground, but there is one final dynamic we haven't discussed, and it completely alters the strategic landscape.

Okay, here is something else to mull over. A thought that builds on everything we've discussed today. We talk about how in the U.S., if your AI output lacks sufficient human authorship, it falls into the public domain.

You own nothing. Well, if your AI-generated code, your AI-generated marketing copy, and your AI-generated strategic analysis are all legally in the public domain, what stops your biggest competitor's AI agent from legally scraping your entire proprietary platform tomorrow to train its next model? Right, if no human authored it, there is no copyright to infringe. By relentlessly optimizing your workforce with AI to save costs, you might be legally stripping the intellectual property rights from your own products.

The very tools you are using to build your business today might be actively destroying the legal concept of a competitive moat. It is the ultimate paradox of AI adoption. You deploy it to move faster, but you inadvertently open your blueprints to the world.

It is a chilling, vital reality to consider. Alright, it is time for the Monday Morning Move. We end every deep dive with a single, zero-filler mandate, the most valuable action you should execute when you sit down at your desk on Monday morning.

Here it is. Open the internal documentation for your company's most valuable, highest-risk AI deployment. Then pull the Vendor's Master Service Agreement and read the IP Indemnity Clause.

And look closely. Yes, identify the exact carve-outs, like the explicit requirement that safety filters must be enabled and inputs must not infringe. Then go directly to your lead engineering team and verify, right then and there, if your current technical logging architecture can actually prove you are compliant with those specific conditions for every single generation.

If you cannot mathematically prove you met the conditions, you do not have an indemnity. You have a wish. Fix the logs.

That is what true, proactive governance looks like in action. Exactly. Thank you for joining us for this Executive Deep Dive.

We hope it gives you the clarity, the frameworks, and the foresight you need to protect your organization's future. Stay sharp, govern your evidence, and we will see you on the next Deep Dive.

Real cases

These examples show the three questions and the liability chain in real, documented cases. Each is used here to illustrate one link; the deep case treatment lives with its owning topic where noted. As you read, tag each example to the part of the chain it lights up:

  • The input question: Examples 1, 2, 10.
  • The output-ownership question: Examples 3, 4, 8.
  • The liability chain and its layers: Examples 5, 6, 7, 9.

Reading them this way is itself the skill: given any AI story in the news, an expert first asks which link it is really about.

Example 1: Thomson Reuters v. Ross Intelligence (the input question, decided). The anchor of this topic. A federal court held that training a competing legal-research tool on Westlaw's copyrighted headnotes, sourced through third-party memos after a license was refused, was infringement and not fair use, on the narrow ground of non-generative use for a directly competing commercial product (Thomson Reuters v. Ross Intelligence, D. Del. Feb. 11, 2025). The chain lesson: ownership of the editorial headnotes stayed with Thomson Reuters through every hop (Westlaw editor to LegalEase memo to Ross model). Liability reached the party that commissioned and used the copying, even after that party had dissolved.

Example 2: The New York Times v. OpenAI and Microsoft (the input question, unresolved). The Times sued over training on its journalism without permission, alleging the models can reproduce its articles closely. The case is ongoing and settles nothing yet, but it is the clearest live illustration that a foundation-model provider's data choices create exposure that can reach downstream deployers of that model. This case is the anchor for Module 5.5; here it stands only as the open counterpart to the decided Ross case (see Topic 5.5).

Example 3: The United States Copyright Office human-authorship guidance (the output-ownership question). In January 2025 the Copyright Office confirmed that purely prompt-generated output is not copyrightable, while human creative selection, arrangement, or modification of AI output can be (US Copyright Office, Part 2, 2025). Real consequence for organizations: a media company that wants to own its AI-assisted work must build and keep a record of the human creative contribution, or accept that the raw output sits in the public domain and can be freely copied.

Example 4: The Beijing Internet Court AI-image ruling (output ownership, a divergent jurisdiction). China's first ruling granting copyright in an AI-generated image, based on the prompter's documented intellectual inputs (2023), shows that the output-ownership answer flips across borders. A multinational cannot assume the United States "no copyright in prompts" rule applies in every market. Deep treatment belongs to Module 10.6 (see Topic 10.6); cited here only to make the divergence concrete.

Example 5: Vendor IP indemnities in practice (the contract layer). Microsoft's Copilot Copyright Commitment, OpenAI's Copyright Shield, and Adobe's Firefly indemnity all promise to defend paying customers against third-party copyright claims over covered AI output, and all condition that promise on the customer using the product as intended, with safety features enabled, on paid tiers (Lexology, 2024; Proskauer, 2024). Real consequence: a firm that disabled content filters to get "less restricted" output may have voided the very indemnity it was counting on.

Example 6: The revised EU Product Liability Directive (the liability-basis layer). Directive (EU) 2024/2853, in force from late 2024 and to be transposed by 9 December 2026, brings stand-alone software and AI within strict product liability and eases a claimant's burden of proving defect and causation for complex systems (Hogan Lovells, 2024; Orrick, 2025). Real consequence: an EU deployer of a defective AI feature can face a no-fault claim, which raises the value of every design record, warning, and log to the level of core legal evidence (see Topic 10.2).

Example 7: The AI-insurance retreat (what happens when the chain has no backstop). In 2025, several insurers filed to exclude AI-chatbot and agent liabilities from coverage, signaling that the residual risk at the end of the liability chain may not be insurable on ordinary terms. This is the anchor for Module 8.4 and is referenced here only to close the loop: when contract and insurance both thin out, the deployer's own evidence and controls are the last line (see Topic 8.4).

Example 8: A denied and a partly-cancelled registration (the human-authorship line, drawn in real files). The United States Copyright Office refused to register an AI-generated image that had won a state-fair art prize, on the ground that its expressive elements were produced by a machine rather than a human. The artist has challenged that refusal in federal court, where the case remained pending as of 2026, so the settled point is the Copyright Office's refusal itself, not a final court judgment. In a separate matter, the Office reconsidered the registration for Kristina Kashtanova's comic "Zarya of the Dawn," which combined human-written text and arrangement with AI-generated images, and narrowed the protection to the human-authored elements, excluding the individual AI-generated images. Real consequence for organizations: these are the human-authorship rule from Section 3D applied to concrete works, and they show the line is not theoretical. Recorded human contribution is registrable; the raw machine output is not (US Copyright Office decisions, 2023 to 2025).

Example 9: The EU AI Act value-chain roles (who becomes responsible). Regulation (EU) 2024/1689 places the heaviest obligations on the provider of an AI system and lighter ones on the deployer, but Article 25 provides that a deployer who puts its own name on a high-risk system, substantially modifies it, or repurposes it into a high-risk use takes on the provider's obligations. Real consequence: an organization that white-labels or heavily fine-tunes a vendor's system can move up the liability chain without noticing, which is why a map records the role you actually occupy, not the one on the purchase order (European Commission; AI Act Article 25).

Example 10: The training-content summary as a usable document (provenance, in your hands). From August 2025, providers of general-purpose AI models placing them on the EU market must publish a summary of training content using the AI Office's mandatory template, which asks for model and provider metadata, the main data-source categories (public datasets, licensed datasets, crawled or scraped content, user data, synthetic data), and the copyright and data-protection measures taken (European Commission AI Office template, 24 July 2025). Real consequence: when you deploy a major model, this summary is a real artifact you can request, read, and file against the input question, turning "we do not know what it trained on" into a documented answer or a documented gap.

Where people go wrong

  • "Who owns the output is one question." It is three: did the input infringe, is the output ownable, does the output infringe someone else. They have different laws, different answers, and different exposed parties. Mashing them together is the root error; separating them is the first expert move.
  • "The AI made it, so no human is responsible." The opposite is often true. When AI output causes harm, the law reaches a human or an organization: the deployer who shipped it, the professional who relied on it without the check their duty required, or the provider who supplied a defective product. "The model did it" is not a liability shield; a tool cannot hold a professional's duty of care.
  • "Ross v. Ross Intelligence decided that AI training is illegal." It did not. It held that one company's copying of copyrighted editorial headnotes, to build a directly competing tool, without a license, was infringement and not fair use, and it expressly limited itself to non-generative AI. The generative-AI questions are still open and on appeal. Overstating Ross to a board is a falsification you will be caught on.
  • "We prompted it in detail, so we own the copyright." In the United States, per the Copyright Office's 2025 guidance, prompts alone (however detailed) do not confer copyright in the output. Ownership attaches to human creative selection, arrangement, or modification, and only to those human contributions. Detailed prompting is instruction, not authorship.
  • "Copyright works the same everywhere, so we only need one answer." No. The United States denies copyright to purely prompt-generated output; the Beijing Internet Court granted copyright to an AI-assisted image on the strength of the human's inputs (see Topic 10.6). A multinational must map output ownership per jurisdiction, not once.
  • "The vendor indemnifies us, so we are covered." Only if you meet the carve-outs. Common conditions: you kept the built-in safety filters on, you used a paid and generally available tier, the claim is about the output and not your inputs, and the payout is not capped below the exposure. An indemnity you cannot prove you qualify for is not protection; it is a sentence.
  • "Provenance is an ethics nicety." As of 2025 it is a legal obligation and a litigation defense. EU GPAI providers must publish a training-content summary under Article 53 of the AI Act, and your ability to answer the input question at all depends on the provenance file you built in Module 2 (see Topic 2.6). No provenance, no defense.
  • "Strict liability, fault liability, same thing." Under fault-based rules a claimant must show you were careless; under strict liability (which the revised EU Product Liability Directive extends to software and AI) they need only show a defective product caused harm. The shift to strict liability, with presumptions that ease the claimant's proof, makes your own records the decisive evidence, because your carefulness is no longer the question.
  • "Liability ends when the company or the project ends." Ross Intelligence shut down in 2021; the infringement ruling landed in 2025. Winding down a system does not wind down its exposure, which is one more reason the evidence must be preserved, not deleted, when a system is retired.
  • "The model does not store the training data, so the output cannot infringe." How the model represents information internally is not the legal test. Courts look at what actually came out. A model can produce output substantially similar to a protected work whether or not that work is "stored" anywhere, and regurgitation of near-verbatim training data is a documented phenomenon. Do not offer "it only learns patterns" as a blanket defense.
  • "We only deploy other people's models, so provider duties are not ours." Not necessarily. Under Article 25 of the EU AI Act, a deployer that puts its own name on a high-risk system, substantially modifies it, or repurposes it into a high-risk use inherits the provider's obligations. White-labeling and heavy fine-tuning can quietly move you up the chain. Record which role you actually occupy.
  • "Confidential data is safe to paste into any AI tool because the vendor promises privacy." Read which tier you are on. Enterprise contracts often promise inputs are not used for training; consumer and free tiers frequently make no such promise. And disclosing a trade secret into a system you do not control can destroy its legal secret status entirely, not merely risk a leak. Input governance is an ownership question too.
  • "Regurgitation risk and look-alike risk are the same problem." They are not, and they need different controls. Regurgitation (the model emits near-verbatim training data) is defended upstream, through vendor assurances and anti-memorization safeguards. Independent similarity (the output merely resembles a protected work) is defended downstream, through human review and records of independent creation. A map that does not name which risk it faces aims its controls at the wrong end.

Questions people ask

What is fair use?
A defense in United States copyright law that permits some unlicensed use of a protected work, weighed across four factors (purpose and character of the use, nature of the work, amount used, and effect on the market). In Thomson Reuters v. Ross Intelligence the defense failed because the use was commercial, not transformative, and harmed the market. More on Fair use
What is headnote?
A short editorial summary of a point of law in a court opinion, written by a publisher's editors. The underlying opinions are government works and not copyrightable, but the editorial headnotes can be, which is what Thomson Reuters owned and Ross was found to have copied.
What is transformative use?
In fair-use analysis, a use that adds a new purpose or different character to the original rather than substituting for it. A use that builds a directly competing product, as in Ross, is generally not transformative.
What is training-data provenance?
The documented record of where an AI system's training and input data came from and on what legal basis it could be used. It is the artifact that lets you answer whether training on the input infringed (see Topic 2.6).
What is general-purpose AI (GPAI) model?
Under the EU AI Act, a model trained on broad data that can perform a wide range of tasks and be integrated into many systems. GPAI providers carry specific duties, including publishing a training-content summary. More on General-purpose AI (GPAI) model

Keep going