The OpenAI Author Case Now Turns on Downloads, Not Training
Unsealed briefs shift the fight from fair use to how the books were obtained — and Anthropic's $3,000-a-book settlement is the number everyone is doing math against.
What actually became public
Redacted portions of the summary judgment briefing in the consolidated authors' litigation against OpenAI and Microsoft are now on the public docket. The headline most outlets ran was some version of "internal documents show AI firm knew." The more useful observation is structural: the plaintiffs' brief spends far more of its argument on how books were acquired than on whether training is fair use. That is not a rhetorical choice. It is what the case law since 2025 has forced.
Reality check one: a brief is a party's best version of its own case, written by lawyers paid to win. Unsealing makes it visible, not true. The exhibits attached to a brief are evidence. The brief itself is advocacy. Any story that quotes the argument and not the exhibit is reporting a press release with a docket number.
Why acquisition became the whole ballgame
To build a large language model you need a pretraining corpus — hundreds of billions of words of running prose. Web scrapes give you volume but poor quality: comment threads, SEO filler, boilerplate. Books give you long-form, professionally edited argument and narrative, which is disproportionately valuable for teaching a model to hold a thought across thousands of words. Every major lab wanted books. The question is where they came from.
There are three routes. Buy and scan physical copies. License from publishers. Or download an existing "shadow library" — bulk pirate collections like the ones referenced in filings across several of these cases — because the entire corpus arrives in an afternoon, indexed and de-duplicated, for nothing.
The 2025 rulings split precisely along that line. In the Anthropic case, Judge Alsup held that training itself was transformative fair use, while retaining pirated copies in a permanent internal library was not. Anthropic settled for roughly $1.5 billion covering about 500,000 works — on the order of $3,000 a book. That figure is now the reference price in every settlement conversation in this field, and it was set by the acquisition question, not the training question.
So when you read that plaintiffs "proved OpenAI infringed," ask which claim. Training on lawfully obtained text has survived challenge more than once. Downloading a torrent of 200,000 novels is a separate act with its own liability, and it happened years before any product shipped.
Where Microsoft sits, and where it doesn't
Microsoft is not alleged to have run the scrapers. Its exposure runs through what it shipped and what it powered: Azure compute for training runs, and consumer products built on the resulting models. That is a contributory and vicarious theory, and it is genuinely harder to prove than direct copying. Treat any claim that Microsoft "has been found liable" as false until a court says so — nothing has been decided on the merits.
The other claim worth deflating is regurgitation. Plaintiffs in these cases have shown that models can be prompted into reproducing passages of specific books. What that demonstrates is memorization under adversarial prompting, not that ordinary users receive infringing output. Defendants respond that filters block it in the shipped product. Both things can be true, and they carry different damages.
Questions You Should Be Asking
- When a vendor says its model is "trained on licensed data," does that cover the base model or only the fine-tuning layer added on top of a base model trained years earlier?
- What does your indemnification clause actually cover — output-level infringement claims against you, or the vendor's own upstream acquisition liability? Those are almost never the same clause.
- If statutory damages landed at the willful maximum rather than the settled rate, could your vendor absorb it, and what happens to your deployment if it cannot?
- Has anyone on your team read the exhibits in these filings, or only the briefs summarizing them?
- Can your vendor produce a provenance record for its pretraining corpus, or only a policy saying it respects copyright?
What To Watch Next
Watch whether the court separates the acquisition claim from the training claim at summary judgment. If it does — allowing the download claim to a jury while resolving fair use on training for the defense — the case becomes a damages negotiation against the Anthropic per-work benchmark, and settlement follows fast. If the claims stay bundled, this goes to trial, and every enterprise contract signed on the assumption of cheap indemnity gets repriced.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
