United States Department of Justice (DOJ) has told a federal court that training AI models on copyrighted text without a license does not violate copyright law, a position that could shape music industry lawsuits against Anthropic, Suno, and Udio.
The department set out that position on Tuesday (September 1) in a filing in the copyright lawsuit brought against OpenAI by The New York Times. It appears to be the first time the US government has intervened in any of the copyright cases now stacked up against AI companies.
DOJ filing in OpenAI case#
The filing is a Statement of Interest of the United States, advice rather than a ruling, and Judge Sidney Stein is free to ignore it in the OpenAI case. The 20-page document was signed by Stanley Woodward, the Associate Attorney General.
The document, which concerns large language models (LLMs), states: ‘The United States has a strong interest in this Court rejecting any argument that training LLMs on copyrighted texts violates copyright law.’
The filing rests on President Donald Trump‘s AI policy, citing two of the President’s executive orders, from January 2025 and June 2026. It also quotes his National Policy Framework for Artificial Intelligence, published in March, which states that the ‘training of AI models on copyrighted material,’ in and of itself, ‘does not violate copyright laws.’
The DOJ takes on the two questions that decide most fair use rulings: how far the new use transforms the original, and whether it damages the market for it.
On the first, the brief argues that copying text such as articles from The New York Times to train a model like ChatGPT is ‘a use of a different kind or character,’ and ‘extraordinarily transformative.’
On the second, the Justice Department argues that a training copy does not ‘serve as a substitute for the original,’ because training ‘does not reveal anything to the public at all.’
Large AI companies paying licensing fees to publishers of titles like the NYT ‘would disproportionately benefit legacy media outlets due to the sheer volume of their written publications,’ the US government adds.
It is not in the public’s interest, the DOJ argues, for the largest tech companies to hold ‘an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies.’
Written works, not songs#
The filing is limited to written works, not songs. It argues about ‘copyrighted texts,’ ‘written works,’ and ‘text articles.’ A footnote limits the DOJ’s reasoning to this case and, specifically, related suits brought by ‘book authors and publishers.’ Recordings and compositions go completely unmentioned across its 20 pages.
The fair use analysis could still be read by judges in music cases, including Suno’s defense.
Three stages of AI model building#
The DOJ divides AI model development into three stages:
- acquiring the material
- training the model on it
- generating outputs
‘Each stage may present distinct questions of copyright law,’ the DOJ says, and it defends only the middle one.
In the main music industry cases against AI companies, major labels and publishers are challenging all three stages.
The first stage is how the material was obtained, and it is an area where the AI industry has already lost ground. In the precedential book authors’ case against Anthropic, Judge William Alsup ruled in 2025 that downloading books from pirate libraries was not fair use, calling it ‘straightforward piracy but at massive scale.’ Anthropic settled with those authors for $1.5 billion in September 2025 over the same torrenting.
Two of the four counts in Sony Music Publishing and Warner Chappell Music‘s new suit against Anthropic, the fifth music copyright case against the Claude developer, concern torrenting. The DOJ’s filing says nothing about any of that.
Market dilution dispute#
The second stage at question in AI cases is the training itself, that is, models being fed information and content and learning from it. It is this stage the DOJ defends, and the one place it goes straight at music’s reasoning.
In 2025, book authors who had sued Meta over AI training lost on fair use. But the judge who decided it, Vince Chhabria, raised a theory that could help rightsholders in future cases. Chhabria suggested that AI outputs carry the ‘potential to flood the market with competing works’ and that developers should therefore ‘generally need to pay copyright holders for the right to use their materials’ even for training.
In other words, for Chhabria, what comes out is evidence that what went in should have been licensed. Lawyers call that market dilution, and it is the argument music has been building on ever since.
In a brief filed on March 30, the Recording Industry Association of America (RIAA), the National Music Publishers’ Association (NMPA), the American Association of Independent Music (A2IM), SoundExchange, and four other groups asked a court to reject Anthropic’s fair use defense in a legal fight with Universal Music Group (UMG), Concord, and ABKCO on similar market harm grounds.
However, the DOJ now calls Chhabria’s reasoning ‘deeply flawed,’ and says he ‘improperly collapsed LLM training and LLM outputs into a single continuous use.’ Training and outputs are two separate legal questions, the US government argues, and what a model produces has no bearing on whether training it was lawful.
