It starts as a hunch. You type your own name, or your art style, into an AI tool and something comes back that feels a little too familiar. Or a headline tells you a model was trained on “the whole internet,” and you think: wait — that includes me. So here’s the practical question nobody answers clearly: can you actually find out whether your work was used to train AI — and if it was, what can you do about it?
The honest answer is: often yes, you can check, and there are real steps to take afterward. No law degree required. Let’s walk through it.

Why you can’t just “look inside” the model
First, a reality check that makes everything else make sense. Once a model is trained, it doesn’t store your painting or your paragraph like files in a folder. The training data gets blended into billions of numerical weights, so there’s no “open the AI and find my picture” button. Asking a chatbot “were you trained on my work?” is useless — it doesn’t know, and will happily make up an answer.
Track the cases yourself
Every current AI-copyright ruling — and the ones still being fought — lives in our AI Copyright Ruling Tracker. Filter by your creator type to see, in plain language, how each case affects your rights.
What you can often inspect is the dataset — the giant pile of scraped material the model was built from. Many of these datasets are public or have been leaked, catalogued, or reconstructed. That’s the door we’re going to walk through.
Know the usual suspects: where the data comes from
Almost all the big models draw from a short list of massive sources. Knowing their names tells you where to search:
- Common Crawl — a free, enormous archive of billions of web pages, re-crawled constantly. If your writing, art, or photos have ever been on a public website, they may be in here.
- LAION-5B — a dataset of roughly 5.8 billion image-and-caption pairs scraped from the web, used to train many image generators like early Stable Diffusion. This is the big one for visual artists and photographers.
- Books3 / LibGen / “the pirate libraries” — collections of digitised books (many pirated) used to train large language models. This is the one that matters for authors and writers.
You don’t need to memorise the tech. You just need to know: images live in LAION-style datasets, and books/text live in Common Crawl and the book collections. Different search tools cover each.

How to check if your images were scraped
For visual work — art, illustration, photography, design — the single most useful tool is Spawning’s Have I Been Trained. It lets you search the LAION datasets (the same billions of images used to train major generators) two ways: by typing keywords, or by uploading one of your own images so it finds visual matches.
Here’s how to use it well:
- Search your name and any usernames or watermarks you sign with.
- Upload a few of your actual images and let it find near-matches — this catches reposts where your name was stripped.
- Search distinctive titles or series names you’ve used publicly.
If your work shows up, it means that image (or a copy of it) was in a dataset used to build AI models. That’s your confirmation. While you’re there, you can also flag those works for opt-out, which brings us to the next question — but first, authors, your turn.
How to check if your writing or books were used
For books, the reporting has been unusually good to creators. Journalist Alex Reisner and The Atlantic built searchable databases of the pirated-book collections (like LibGen and the Books3 set) that companies used for training. You can type in an author or a title and see whether that specific book was in the trove. If you’ve published a book, this is the fastest way to get a yes-or-no.
For shorter writing — blog posts, articles, forum answers — it’s harder to get a per-item result, because that content mostly rode in through Common Crawl at web scale. The practical test: was it ever publicly posted on the open web? If yes, assume it was almost certainly crawled. That’s not paranoia; it’s how these datasets are built.
Found a match? Here’s exactly what to do next
Confirming your work was scraped is unsettling, but it also hands you options. Work through these in order:
- 1. Document it. Screenshot the search result, save the URL, note the date, and keep your original files with their creation dates. This evidence file is the foundation for everything else — it proves the work is yours and that it appeared in a dataset.
- 2. Opt out going forward. Register your work in a Do Not Train registry and lock down your own site so future crawls skip you. It won’t undo past training, but it closes the gate ahead. (We cover this step-by-step in our guide to opting out of AI training.)
- 3. Send a DMCA takedown where it fits. If your work is being hosted or reproduced without permission — say, a dataset copy or a site redistributing it — a DMCA notice can force its removal. DMCA is aimed at copies being distributed, not at the abstract act of training, so it’s a targeted tool, not a cure-all.
- 4. Check whether a settlement or class action already covers you. This is the big one — see below.

You may already be owed something: settlements and class actions
Here’s what most creators don’t realise: you don’t have to launch your own lawsuit to benefit. Creators have banded together, and some of those cases have already produced results.
The landmark example is Bartz v. Anthropic. In 2025, Anthropic agreed to a roughly $1.5 billion settlement with a class of authors whose books were downloaded from pirated-book libraries and used in training — one of the largest copyright recoveries on record. Crucially, there’s an official settlement site with a searchable Works List: authors and rights-holders can look up their titles to see if they’re covered and file a claim.
So if you’re an author, your action item is concrete: go to the settlement’s official Works List and search your books. If they’re listed, you may be eligible to claim a share.
Beyond that one case, several major suits are still moving through the courts — visual artists in Andersen v. Stability AI, Getty Images against Stability, and others. You generally don’t need to do anything to be part of a certified class except keep your evidence and watch for official claim notices (never pay a “fee” to join a legitimate class action — real ones don’t work that way).
The realistic bottom line
Can you find out if your work was used to train AI? For images and published books, very often yes — the tools above will tell you. For scattered web writing, the honest answer is “if it was public, assume so.” What you can’t do is surgically remove your work from an already-trained model; that technology doesn’t exist yet.
But finding out isn’t the end of the story — it’s the start of your leverage. It lets you opt out going forward, document a clean record, fire off takedowns where they apply, and step into settlements and class actions that are increasingly landing in creators’ favour. The creators who benefit most from this shifting landscape are the ones who checked, kept the receipts, and were ready.
A light note: this is general information to help you understand your options, not legal advice. If you’ve found your work in a dataset and real money or a specific claim is on the line, talk to an IP attorney about your situation.
Sources & further reading:
- Spawning — Have I Been Trained (search the LAION training datasets)
- The Atlantic — Search LibGen, the pirated-books dataset used to train AI (Alex Reisner)
- Bartz v. Anthropic — Authors’ copyright class-action settlement information
- LAION — About the LAION-5B image-text dataset
- U.S. Copyright Office — DMCA takedown notices (section 512)