A Microsoft scientist called it the largest theft of labor in history
Unsealed filings, a narrow Ninth Circuit ruling and a new antitrust suit reveal three different legal defenses that all end the same way.
The quote came from inside the house
In a 92-page combined brief filed by news plaintiffs in the New York Times copyright case against OpenAI and Microsoft, a Microsoft Director of Applied Science, Dr. Brent Hecht, is quoted describing the case as “an astonishing theft of unprecedented proportions” and possibly the “largest theft of labor in human history.” That is not a campaigning journalist. That is a senior employee of one of the defendants, quoted in documents recently unsealed and reported on by 404 Media, as summarised in a Register opinion column published on 27 September.
These are the plaintiffs' arguments. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not ruled. OpenAI and Microsoft dispute the allegations, arguing fair use: that training on public articles and books advances public knowledge and does not act as an “unlawful economic market substitute” for the originals.
What the filings actually describe
Strip out the rhetoric and the documents describe a mechanism. Large language models are trained by ingesting enormous volumes of text; the model learns statistical patterns rather than storing copies, which is why “fair use” is the natural defence. The interesting part of fair use is the fourth factor — market effect. If the output substitutes for the original in the market, the defence weakens. That is precisely the ground the parties are fighting on.
Two details in the filings matter more than the insults. First, according to the brief, OpenAI's own corporate representative testified under oath that they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When cofounder Greg Brockman was told OpenAI could get past a paywall, the brief quotes him as replying, “ah, nice.” Second, an internal Microsoft policy document quoted in the papers concedes that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain” — what the same documents reportedly call a “doom loop,” leading to model collapse as the supply of fresh human work dries up.
There is a distribution problem underneath the legal one. As an OpenAI software engineer is quoted saying: “No matter how prominently we show the links, users won't click.” Citation as a remedy assumes traffic follows attribution. The engineer's own observation says it doesn't.
Why the GitHub win is narrower than it looks
The Ninth Circuit handed GitHub, Microsoft and OpenAI a win in Doe v. GitHub, but on one specific provision of the DMCA concerning copyright-management information — the author name, licence notice and terms attached to a work. Judge Eric Miller wrote: “One who creates a new work and fails to include CMI cannot be said to have ‘removed’ or ‘altered’ anything.”
That decides one claim, not the field. It did not hold that training on open source without restriction is permitted, that copying code into a training set is always fair use, that generated code cannot infringe, or that the GPL, Apache, BSD or MIT licences are unenforceable. An open source attorney quoted in the column flags a subtler risk: parties increasingly litigate open source licences as contracts rather than as the copyright licences they were drafted to be. Contract theories sidestep the questions copyright forces you to answer — do you own the work, is it protectable expression, was it actually copied? Win on contract and you may have enforced the terms while quietly converting the licence into something else.
The practical consequence is a provenance gap. A developer accepting a suggestion cannot tell which licence, if any, governed the code it was derived from, and in AI-assisted pipelines there is no link to check.
Questions You Should Be Asking
- If a court later finds that training data was unlawfully acquired, does your code assistant vendor's indemnity cover you — and does it cover contract claims over open source licences, or only copyright claims?
- Can your vendor produce any provenance record for generated code, or is “new work, no CMI” the entire answer?
- If your business model depends on being cited by an AI product, what is that worth given an OpenAI engineer's own statement that users won't click regardless of link placement?
- For the antitrust suit against Anthropic, OpenAI, SpaceXAI and Google: what evidence exists of an actual agreement and harm to competition, beyond executives making similar public statements about safety over one weekend?
- If Microsoft's internal documents concede LLMs destroy their own supply chain, what is the plan for sourcing fresh human work — and who pays for it?
What To Watch Next
Judge Stein's ruling in the Southern District of New York is the signal. If the fair use defence survives summary judgment, the “take it now, argue later” model is effectively ratified and licensing deals become optional goodwill. If it doesn't, every unsealed internal memo about doom loops and paywalls becomes evidence in the next twenty cases. Watch also whether any remedy has teeth — the column notes recent antitrust wins against Google's ad business and search were declawed at the punishment stage.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
