What a Newspaper's Lawsuit Against OpenAI Says About How AI Systems Should Be Built
The Seattle Times and Newsday sued OpenAI and Microsoft on September 5, 2026 over AI training on their journalism. Underneath the copyright claim is a missing architectural feature: a link from every AI answer back to the source it came from.

On September 5, 2026, the Seattle Times and Newsday sued OpenAI and Microsoft in the U.S. District Court for the Southern District of New York. The complaint alleges the companies scraped the newspapers' websites, including content behind paywalls, and folded the articles into the datasets that train and run ChatGPT, Microsoft Copilot, and Bing's AI features.
What are the newspapers actually claiming?
The papers' central claim is specific: the resulting AI products can reproduce passages from their reporting, closely paraphrase full articles, and answer a reader's question using that reporting, all without sending the reader to the original story or the subscription that funded it. They are asking for damages and a court order to impound or destroy any copies of their work, training datasets, or AI models built on it. Microsoft's public response was surprise, paired with a stated openness to discuss solutions. TechCrunch also notes that Microsoft and OpenAI had previously funded Seattle Times journalism projects and fellowships, which makes the dispute sharper: a funder and a subject of the same reporting, on opposite sides of a courtroom.
Where does this filing sit in the wider docket?
This filing joins a widening docket. The New York Times sued the same two companies over the same underlying question in December 2023, and several other publishers have joined since. That earlier case is back in the news this week too: the Department of Justice filed a statement on September 3, 2026, siding with OpenAI's argument that training on copyrighted material counts as fair use, a position that drew public pushback even from within the administration's own political coalition.
What is the complaint really describing?
Strip away the legal argument for a moment and look at the underlying complaint as a systems problem. The papers are describing more than unauthorized use of their work. Their claim is that the output has no traceable link back to where it came from. A reader gets an answer built from real reporting, and nothing about that answer tells them whose reporting it was, whether it is current, or where to verify it. The information moved from source to output. The pointer back to the source stayed behind.
Provenance, in an AI system, is the link from a claim in the output back to the specific, checkable source it came from. It is the answer to a single question a reader should be able to ask of any sentence: where did this come from?
Is provenance a design choice?
Yes. It is possible to build an intelligence system the other way: every claim in the output carries a citation back to a specific, checkable source, and a claim that cannot be traced to a real source stays out of the output entirely. That is the core architectural bet behind ForIntel, Foragentis's knowledge-products line. Every report we produce is built so a claim without a verifiable source row cannot publish, a structural requirement of how the report gets assembled from the start, rather than a footnote added after the fact.
The same missing layer shows up wherever fluent output outruns verification. Apple rationed its own bug bounty after unverifiable security reports flooded its reviewers. In that case too, the output read as real. What was missing was the pointer back to a source someone could check.
What can be said with confidence right now?
How this lawsuit should be decided is a question for the court, and for the harder policy debate about who should own the value created when a machine learns from a body of human work. What can be said with confidence is narrower and more useful to anyone building or buying AI systems right now: provenance is buildable. A system can be designed so that every output carries a real answer to "where did this come from," and that design choice is available today, independent of how any particular lawsuit resolves.
If you are evaluating an AI vendor and want to know whether their outputs trace back to real, checkable sources, that is exactly the question ForIntel was built to answer: foragentis.com.
FAQ
Why are the Seattle Times and Newsday suing OpenAI and Microsoft?
On September 5, 2026, the two newspapers filed suit in the U.S. District Court for the Southern District of New York. The complaint alleges that OpenAI and Microsoft scraped the papers' websites, including content behind paywalls, and folded the articles into the datasets that train and run ChatGPT, Microsoft Copilot, and Bing's AI features. The papers say those products can reproduce passages from their reporting, closely paraphrase full articles, and answer a reader's question using that reporting, all without sending the reader to the original story or the subscription that funded it.
What are the newspapers asking the court for?
They are seeking damages and a court order to impound or destroy any copies of their work, the training datasets built from it, and the AI models trained on it. Microsoft's public response was surprise, paired with a stated openness to discuss solutions.
How does this case relate to the New York Times lawsuit against OpenAI?
The New York Times sued OpenAI and Microsoft over the same underlying question in December 2023, and several other publishers have joined since. On September 3, 2026, two days before the Seattle Times and Newsday filed, the Department of Justice submitted a statement in that earlier case siding with OpenAI's argument that training on copyrighted material counts as fair use. That position drew public pushback, including from within the administration's own political coalition.
What is AI provenance?
Provenance is the link from a claim in an AI system's output back to the specific, checkable source it came from. A system with provenance can answer, for every claim it makes, the question "where did this come from?" A system without it delivers answers built from real sources while leaving the pointer back to those sources behind, so a reader cannot tell whose work it was, whether it is current, or where to verify it.
How does ForIntel handle source provenance?
Every ForIntel report is assembled so that a claim without a verifiable source row cannot be published. The citation is a structural requirement of how the report is put together from the start, rather than a footnote added afterward. A claim that cannot be traced to a real source stays out of the report entirely.
ForIntel is a research-first business-intelligence product from Foragentis. Every quantitative claim in a ForIntel brief traces to source data, is cross-corroborated across independent measures, and passes a counter-signal verification step before it is published. To see a sample brief or commission a read, reach the ForIntel desk at forintel@foragentis.com.


