The United States government has entered the legal dispute between OpenAI and a group of newspapers, filing a brief in Manhattan federal court that supports OpenAI's position that using copyrighted works to train large language models can qualify as fair use.
The brief, submitted on Tuesday, represents the first time the federal government has publicly weighed in on a wave of cases brought by copyright owners including authors, publishers, music labels and news outlets over the use of their material in AI training. The filing states that "The United States has a strong interest in this court rejecting any argument that training LLMs on copyrighted texts violates copyright law" and cites concerns such as scientific advancement and national security as reasons for that view.
Spokespeople for the White House, the New York Times and OpenAI did not immediately respond to requests for comment on the filing on Wednesday.
The origins of the dispute lie in a lawsuit first filed by the New York Times in 2023. In that complaint, the newspaper accused OpenAI and its largest financial backer, Microsoft, of using millions of newspaper articles without permission to train OpenAI's ChatGPT models. The Manhattan case is one among several legal actions initiated by copyright holders against technology firms, with plaintiffs contending that their material was misused to develop AI systems.
Those pending cases broadly turn on the legal question of whether AI systems make fair use of copyrighted material when they incorporate it into training datasets and produce new, transformative outputs. The courts' decisions so far have not been uniform. The first two judges to directly consider the issue issued diverging rulings last year, leaving a measure of legal uncertainty around the standards that will apply to LLM training.
In its brief, the government agreed with technology companies that AI training is "extraordinarily" transformative, and it emphasized potential benefits. The filing states, "Beyond the subject matter of this litigation, LLMs are already helping researchers across fields achieve major breakthroughs." It also warned that "Constraining LLM development under a misunderstanding of fair use doctrine would thwart such creative and scientific progress while hindering American prosperity and economic mobility."
Legal experts, rights holders and technology firms will be watching how the Manhattan court interprets those arguments. The government's intervention could have implications for a range of stakeholders, including news organizations that have pursued litigation and technology and AI firms facing related claims. The outcome of this case, and others like it, may shape the permissible scope of data used for model training and influence future litigation strategies on both sides.