Fair Use vs. Infringement: The AI Training Debate in Plain English

Every day, AI models get better at writing like you, drawing like you, and sounding like you. To learn those tricks, they were fed enormous piles of text and images — books, articles, artwork, photos, code. A lot of that material was created by people who never got asked, never got paid, and never got a heads-up. That’s the fight in a nutshell: was training an AI on copyrighted work “fair use,” or was it infringement?

You’ve probably seen both sides shouted at full volume. “It’s theft!” versus “It’s just learning, like a human!” The real answer lives in the messy middle, and in 2025 a handful of courts finally started drawing lines. Here’s the debate in plain English — and, more importantly, what it means for you as a creator.

Fair Use vs. Infringement — context image

Fair use, in one paragraph

Copyright gives you the exclusive right to copy, distribute, and adapt your work. Fair use is the big exception. It lets people use copyrighted material without permission in certain situations — commentary, criticism, teaching, research, parody — because society benefits and the original creator isn’t really harmed. The catch is that “fair use” isn’t a checklist you can pass. It’s a judgment call a court makes after weighing four factors. AI companies say training is fair use. Creators say it isn’t. Both are really arguing about how those four factors apply.

Track the cases yourself

Every current AI-copyright ruling — and the ones still being fought — lives in our AI Copyright Ruling Tracker. Filter by your creator type to see, in plain language, how each case affects your rights.

The four questions a court actually asks

U.S. law (17 U.S.C. § 107) tells judges to weigh four things. Think of them as four questions:

  • 1. Why, and how, is the work being used? Is the new use transformative — does it add new meaning or purpose — or does it just repackage the original? Commercial use counts against fair use, but far less than most people assume.
  • 2. What kind of work was copied? Copying from factual works (a phone book, a news report) is treated more leniently than copying from highly creative works (a novel, a painting, a song).
  • 3. How much was taken? A short quote leans toward fair use; copying the whole thing leans against it — though sometimes you need the whole work to make the new use work.
  • 4. Does it hurt the market for the original? This is the heavyweight. If the new use competes with the original or kills its licensing market, that’s the strongest argument against fair use.

No single factor wins automatically. A judge weighs all four together, which is exactly why two smart people can look at the same AI model and reach opposite conclusions.

“Transformative” is the word everything hangs on

AI companies pin their whole case on factor one. Their argument: a model doesn’t store your painting and hand out copies — it studies millions of examples to learn patterns, then generates something new. That, they say, is transformative in the same way a search engine or a book-scanning project is: a genuinely different purpose from the original.

Fair Use vs. Infringement — detail image

Creators push back on two fronts. First, “learning patterns” can still produce outputs that mimic a specific artist’s style or reproduce chunks of a specific text. Second — and this is the sharper point — even if the training is transformative, that doesn’t automatically excuse how the material was obtained. Transformation is about purpose. It doesn’t give anyone a free pass to grab the raw materials from a pirate site.

How courts have actually ruled so far

For years this was all theory. In 2025, three rulings turned it concrete — and, tellingly, they didn’t all land the same way.

  • Thomson Reuters v. Ross Intelligence (Feb 2025). A startup used Westlaw’s editorial “headnotes” to train a legal-research tool that competed directly with Westlaw. The court said not fair use. The use wasn’t transformative — it served the same purpose as the original — and it threatened Thomson Reuters’ actual market. Note: this was a non-generative search tool, and the head-to-head competition mattered a lot.
  • Bartz v. Anthropic (June 2025). Judge William Alsup ruled that training an AI on books was “exceedingly transformative” and fair use — a big win for the AI side. But he split the question in two: Anthropic had also downloaded and stockpiled millions of pirated books to build a permanent library, and that was straightforward infringement, fair use or not. Anthropic later agreed to a landmark settlement over those pirated copies.
  • Kadrey v. Meta (June 2025). Judge Vince Chhabria sided with Meta — but went out of his way to say the authors lost because they made weak arguments, not because Meta was clearly in the right. He flagged “market dilution” (AI flooding the market with cheap imitations that undercut human creators) as potentially the strongest argument against AI training. He practically wrote the next lawsuit’s playbook.

Read together, the message is: training can be transformative, but that’s not the end of the story. How you got the data, and whether you’re hurting the creator’s market, can still sink you.

The two things that keep tripping AI companies up

If you strip away the jargon, most of the risk for AI companies boils down to two issues — and they’re the two places creators have the most leverage.

Where the data came from. “The training was fair use” and “we were allowed to download it” are separate questions. Scraping a pirate library of stolen books is a copyright violation on its own, no matter how clever the model is. Legitimately licensed or lawfully accessed data is a much stronger position — which is exactly why you now see AI companies signing licensing deals with publishers and stock libraries.

Market harm. Factor four is where creators land their best punches. If an AI trained on your work then generates cheap substitutes that eat your commissions, or produces near-copies of your style on demand, that’s real economic damage — and courts have signaled they’ll take it seriously when it’s actually proven with evidence.

Fair Use vs. Infringement — concept image

What this means for you as a creator

You don’t need a law degree to act on this. A few practical takeaways:

  • The law is unsettled, and that cuts both ways. No court has blessed all AI training as fair use, and none has banned it. Don’t let anyone tell you the fight is over — it isn’t.
  • Keep your receipts. If you can show an AI output closely mimics your specific work or style and hurt you commercially, that’s the kind of evidence that matters. Save dated originals, timestamps, and any provenance you can.
  • Opt out where you can. Robots.txt, “no AI training” settings on platforms, and ai.txt files won’t stop everyone, but they document that you said no — useful if a fight ever happens.
  • Watch the outputs, not just the training. Your strongest complaint is often not “you trained on me” but “your model produced something that copies me and cost me work.”
  • Licensing is becoming the norm. As courts lean on factor four, more AI companies will pay for data. If you have a body of work, collective licensing may eventually put money in your pocket.

The debate isn’t really “AI good vs. AI bad.” It’s a decades-old copyright test being applied to a brand-new machine — one factor at a time, one case at a time. And so far, the courts are saying the same thing to creators that this whole area has always said: the details decide everything.

This article is general information, not legal advice. If an AI company has used your work in a way that’s cost you money, talk to an intellectual-property lawyer about your specific situation.


Sources & further reading:

Related Reading

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top