bytevyte
bytevyte
Language
ai-beats

Seattle Times and Newsday lawsuit escalates the AI copyright fight

Seattle Times and Newsday lawsuit

The Seattle Times and Newsday are asking a federal court to order the destruction of OpenAI and Microsoft AI models trained on their journalism, along with monetary damages, in a complaint that combines copyright claims with trademark claims. The Seattle Times and Newsday lawsuit, filed September 4 in the U.S. District Court for the Southern District of New York, alleges that the two companies copied the newspapers' reporting without a license in order to build and operate products that include ChatGPT and Copilot.

The filing is the newest in a wave of copyright actions that media companies have brought against generative AI developers. It also shows how far the campaign has spread: the plaintiffs are regional outlets that sell digital subscriptions, not the national titles that led the first round of cases.

What the Seattle Times and Newsday lawsuit alleges

The complaint describes a systematic data pipeline. The newspapers contend that OpenAI and Microsoft scraped articles from the publishers' websites, worked around paywalls to reach subscriber-only material, and stripped copyright management information from the text. That content was then folded into large language models that now answer users' questions in their own words.

The paywall allegations carry particular weight for these two plaintiffs. Both papers built their digital operations around subscriptions to local and investigative reporting that appears nowhere else, and both argue that content only paying readers can see was harvested through deliberate circumvention and converted into free training material for a competing service. In their telling, the defendants took the product that readers fund and gave it away as a feature of their own tools.

The economic theory goes beyond the act of copying. AI-generated summaries substitute for original articles instead of pointing readers toward them, the publishers argue, so a person who gets a synthesized answer has less reason to visit the papers' sites, pay for a subscription, or generate the advertising revenue that keeps a newsroom running. The complaint describes a parasitic dynamic in which the assistants feed on reporting they never helped finance, and it portrays the technology as a self-consuming cycle that could leave the journalism industry unable to recover.

The publishers also separate two stages of harm: the copying that happened while the models were built, and the behavior of the finished products, which continue to draw on the papers' reporting every time they answer a question. Training data determines what the models know, they argue, while the outputs determine what readers are told, and both stages now feed off their work without payment.

That framing targets the heart of fair use. Courts weigh whether the copying harms the market for the original work, and an argument that chatbot answers displace reporting speaks directly to that factor. If the responses merely send readers back to the articles, the harm looks modest and the copying resembles permissible quotation; if they deliver the substance of the reporting in one place, the publishers' substitution theory holds together.

Removal of copyright management information adds a separate legal thread. Altering or deleting the metadata that identifies a work and its owner is independently actionable under U.S. law and can carry its own statutory damages. Even if a court rejects the core infringement theory, this claim gives the papers a route to recovery.

Trademark claims rest on what the models produce rather than on what they absorbed. Fabricated or distorted answers that appear to come from a named outlet mislead readers about the source of the information and damage the outlet's standing with its audience, the newspapers argue. Because attribution errors are routine in generative systems, the brand harm to the outlets is a recurring part of how these products behave, they say, not an isolated accident.

Model destruction: a remedy aimed at training economics

The request to destroy trained models is the most consequential demand in the filing. A damages award can be absorbed into a licensing negotiation, while an order that wipes out AI systems treats the model weights as carriers of the publishers' copyrighted expression and declares that products built on the disputed material cannot lawfully keep running. Compliance would require OpenAI and Microsoft to trace which capabilities depend on the contested content and to rebuild or remove them.

The remedy itself is familiar in copyright law, where courts may order the destruction of infringing copies. What is untested is whether a neural network counts as a copy, or as a container of copies, for that purpose. The publishers are asking a judge to stretch a tool designed for physical goods such as counterfeit books to cover systems whose internal weights no one can inspect by hand.

The structural risk reaches beyond the two defendants. If a court accepts that trained systems can be seized because they embody copied text, every developer that trained on scraped news content faces comparable exposure, and that possibility becomes part of the cost of every future training run. For the newspapers, the destruction remedy is the point: it converts a dispute over past copying into a judgment about whether unlicensed training can remain economically viable.

Enforcement would be a separate contest. A trained system blends licensed material, public-domain text, and the disputed articles into the same weights, so destroying a model would also erase capabilities that have nothing to do with the newspapers' work. Because expression is not stored as discrete files inside a network, verifying that particular articles were removed would itself be a technical and legal dispute that could outlast the liability phase of the case.

Filed alongside the pending New York Times case

The Seattle Times and Newsday lawsuit echoes the copyright action the New York Times brought in the same courthouse in late 2023, which has yet to be resolved. The new complaint extends that playbook with theories the earlier case does not press: trademark claims, harms defined by what the models output instead of only by what they absorbed, and the argument that synthesized answers replace reporting in the market.

The legal environment around the dispute is also shifting. In the New York Times proceeding, the U.S. Justice Department has told the court that the capacity to train AI systems on copyrighted works touches national-security priorities, and it has cautioned against a ruling that would end fair use for model development. That government position cuts against the publishers and all but ensures the two cases will litigate the same doctrinal question with the executive branch aligned with OpenAI and Microsoft.

The defendant lineup carries its own local tension. The Seattle Times Co. publishes the daily newspaper of Microsoft's home city and is suing a company based in nearby Redmond, Washington, while Newsday brings the dispute from Long Island. Microsoft is also OpenAI's principal corporate backer and distributes the underlying ChatGPT technology through its Copilot products, which is why the complaint names both companies. Two regional papers built on subscriptions joining a national litigation campaign shows how exposed paywalled journalism has become in the age of generative AI.

Why this matters

For publishers, the Seattle Times and Newsday lawsuit will test whether copyright protection stops at the boundary of the training data or extends into the model itself. For developers of large language models, a ruling that opens the door to destroying trained systems would raise the risk and cost of every future training run, and that uncertainty would reshape how content licensing is priced across the industry. The case will help determine whether original reporting is an input AI companies must pay for or a resource they are free to consume.

Photo by Jon Matthews on Unsplash

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.