Trump Admin To Court: AI Training Is 'Fair Use'

The Trump administration is arguing to a federal judge that artificial intelligence (AI) companies make fair use of copyrighted material when they use it to train large language models.

"Consistent with the national interest in promoting innovation and free expression, the 'training of AI models on copyrighted material,' in and of itself, 'does not violate copyright laws,'" the Department of Justice said Tuesday in a court filing, quoting from the White House's March policy framework for artificial intelligence.

"LLMs are already helping researchers across fields achieve major breakthroughs," the administration writes, using an acronym for large language models.

"Constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility," the government adds.

advertisement

advertisement

The Justice Department submitted the statement to U.S. District Court Judge Sidney Stein, who is presiding over a copyright infringement lawsuit by The New York Times Co. and other publishers against OpenAI.

The publishers allege that OpenAI infringed copyright in at least two ways -- using news articles to train ChatGPT, and reproducing portions of articles in outputs that are generated in response to users' prompts.

The Trump administration argues to Stein that an artificial intelligence company's ingestion of articles for training is "transformative" -- which is one of the factors judges consider when evaluating fair use.

"The purpose of the copying (to build an intelligent, interactive model) differs in kind from the purpose of the copied work (to use language to directly entertain or educate a reading audience)," The Justice Department writes.

The administration is not necessarily arguing that the case should be dismissed -- The Justice Department also writes that displaying articles as outputs "may not be transformative" if a chatbot "reconstructs and disseminates an original copyrighted work."

In addition to arguing that harnessing articles to train chatbots is transformative, the government also contended that using copyrighted material for training does not in itself weaken a market for the original works.

"Training does not reveal anything to the public at all -- it simply creates a copy of a protected work in order to teach an LLM to recognize relationships between data and adapt to new information," the Justice Department writes. "The potential for future outputs that might cause market harm is simply not relevant to evaluating an LLM training use."

A spokesperson for the Times said the administration "is siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole."

"Both AI and creators can thrive -- AI companies simply need to pay fairly for the content that makes their products possible, as copyright law requires," the spokesperson said. "The Administration’s proposal to let companies take that content without permission or compensation would undermine the sustainability of the human-created content that a healthy society depends on, and which AI needs to function."

To date, two federal judges have weighed in on whether using copyrighted materials to train large language models infringes authors' rights.

In one matter, U.S. District Court Judge William Alsup, also in the Northern District of California, ruled that artificial intelligence company Anthropic did not infringe copyright by digitizing books it had purchased and then using the texts to train the chatbot Claude.

"The use of the books at issue to train Claude and its precursors was exceedingly transformative and was a fair use," Alsup wrote last year. He also sided against Anthropic with regard to allegations that it downloaded millions of pirated books.

Anthropic later agreed to a $1.5 billion settlement with authors.

In the other case, U.S. District Court Vince Chhabria in the Northern District of California said that copying material in order to train generative artificial intelligence would likely be illegal most of the time.

Chhabria said in that matter that using books to train large language models was transformative, but would also likely infringe copyright "in most cases" because generative artificial intelligence "has the potential to flood the market with endless amounts of images, songs, articles, books, and more."

The Justice Department criticized Chhabria's ruling, writing that it "misapplies copyright principles."

Next story loading loading..