AI · 18 September 2026
Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.
What happened
Newly unsealed filings in the New York Times' copyright litigation against Microsoft and OpenAI reveal that a senior Microsoft executive privately described AI data-scraping practices as "the largest theft of labor in human history." The documents, previously redacted, show internal communications in which both companies acknowledged scraping paywalled Times content to build training datasets, while separately flagging concerns that such practices could severely damage news publishers.
According to reporting from TechCrunch and Ars Technica, the filings indicate that Microsoft and OpenAI personnel were aware of the scale and potential consequences of using copyrighted journalism to train large language models, even as the companies continued the practice. The unsealing adds new detail to the Times' ongoing lawsuit, which centres on whether AI firms unlawfully used copyrighted articles without licensing or compensation.
Why it matters
The case sits at the centre of one of the most consequential unresolved questions in AI development: how training data is sourced, and who bears the cost when that data is proprietary or paywalled. Internal admissions of this kind — from executives at companies building some of the most widely deployed AI systems — sharpen scrutiny on data provenance practices across the industry, not just at Microsoft and OpenAI.
For organisations deploying or building on large language models, the filings underscore that data-sourcing risk is no longer theoretical. Enterprises embedding third-party AI models into customer-facing or operational workflows should expect continued legal and reputational uncertainty around training data, and may need to factor licensing transparency into vendor due diligence going forward.
The Renascence take
Much of the coverage will focus on the litigation mechanics. The more durable story is about trust infrastructure: AI's usefulness has outpaced the industry's willingness to resolve how the underlying content was obtained, and that gap is now surfacing in discovery documents rather than policy statements.
What's notable isn't that scraping happened — that's long been suspected — it's that the people building these systems described it internally in terms that suggest they understood the harm and proceeded anyway. That's a governance failure as much as a legal one. Any organisation building AI-enabled experiences should be asking vendors not just "what can this model do," but "what did it cost someone else to make it possible," because that question is about to become a lot harder to avoid answering publicly.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.