Seattle Times and Newsday Join Lawsuits Against OpenAI and Microsoft

4 min read
Seattle Times and Newsday Join Lawsuits Against OpenAI and Microsoft

Background to the New Lawsuits

The Seattle Times and Newsday have recently filed legal actions that target two of the most prominent technology companies in the field of generative language models. Both newspapers claim that their articles were incorporated into training data without consent, a practice they argue violates copyright law.

These filings follow a series of similar cases that began with major outlets such as The New York Times and The Associated Press. The earlier suits set a precedent for how courts might address the tension between modern machine learning techniques and traditional publishing rights.

Key Claims Made by the Plaintiffs

According to the complaints, the defendants collected large amounts of publicly available news content, including articles from the Seattle Times and Newsday, and used that material to improve the performance of their language models. The plaintiffs argue that this process amounts to an unauthorized reproduction of copyrighted works.

The complaints emphasize several points:

  • Direct copying of verbatim excerpts from articles.
  • Systematic scraping of news sites without explicit licensing agreements.
  • Commercial benefit derived from the models that rely on the scraped content.

Each of these allegations is presented as a violation of the exclusive rights granted under the U.S. Copyright Act.

Copyright and Fair Use Analysis

The courts will likely apply the four‑factor test used to evaluate fair use. The first factor examines the purpose and character of the use. The defendants argue that the transformation of raw text into a predictive model constitutes a new purpose. However, the plaintiffs counter that the commercial nature of the services outweighs any transformative claim.

The second factor looks at the nature of the copyrighted work. News articles are factual in nature, but they also contain original expression, which courts have protected in past decisions.

The third factor measures the amount of material used. The complaints allege that entire articles were copied, which would weigh heavily against a fair use defense.

The fourth factor considers the effect on the market. The newspapers assert that the language models reduce the need for readers to visit original news sites, thereby harming advertising revenue.

Legal Landscape and Precedents

Recent rulings provide a mixed picture. In Reuters coverage of the New York Times case, a district court found that the use of news articles for training was not automatically exempt from copyright protection. The decision highlighted the need for clear licensing frameworks.

Another notable reference is a memorandum from the U.S. Copyright Office that acknowledges the growing complexity of digital content reuse. The office has called for legislative updates to address the specific challenges posed by machine learning.

Legal scholars at the Electronic Frontier Foundation have warned that overly broad restrictions could stifle innovation, while also emphasizing the rights of creators to control the distribution of their work.

Potential Outcomes

If the courts rule in favor of the Seattle Times and Newsday, the decision could establish a clear requirement for technology firms to secure licenses before incorporating news content into training datasets. Such a precedent would likely trigger a wave of negotiations between media companies and AI developers.

Conversely, a ruling that favors the defendants could reinforce the notion that publicly available text is fair game for training, provided that the use does not directly substitute the original work.

Industry Reactions

Both lawsuits have drawn attention from industry groups and advocacy organizations. The News Media Alliance issued a statement urging policymakers to create a balanced framework that protects journalistic labor while allowing responsible technological advancement.

Microsoft and OpenAI have responded with a joint statement emphasizing their commitment to respecting intellectual property. They noted that they have implemented “responsible data practices” and are open to discussing licensing solutions with publishers.

What This Means for Readers

For the average news consumer, the immediate impact may be subtle. However, the outcome of these cases could affect how news is delivered online, how paywalls are enforced, and whether free access to certain articles remains viable.

In the longer term, a shift toward formal licensing could lead to new revenue streams for news organizations, potentially supporting investigative reporting and local journalism.

Broader Implications for Technology and Media

The lawsuits underscore a broader debate about the balance between data-driven innovation and the protection of creative labor. As more companies rely on large text corpora to train sophisticated models, the legal standards governing data collection will become increasingly central to the industry’s evolution.

Stakeholders from both sides are calling for clearer guidelines. Lawmakers are being asked to consider amendments that address the unique nature of machine learning while preserving the economic interests of content creators.

Regardless of the final judicial decisions, the cases highlight the need for ongoing dialogue between technology firms, media outlets, and regulators.

Comments

No comments yet. Be first.

More from this author