The Redaction Pen Finally Slipped
For two years, Big Tech gave us the exact same public defense. Training frontier AI models on copyrighted books, journalism, and personal blogs wasn't stealing. It was "fair use." It was digital learning, no different than a kid reading library books to learn how to write.
Turns out, their own executives didn't believe a word of it.
Unsealed court records in The New York Times copyright lawsuit against Microsoft and OpenAI have ripped away the PR polish. The documents show that while tech leadership defended web scraping in courtrooms and congressional hearings, senior figures inside Redmond were privately panicking about what they were actually doing.
One Microsoft executive was unusually blunt in internal messages, describing mass web scraping as "the largest theft of labor in human history."
He wasn't wrong. But Microsoft backed the project anyway.
What the Internal Memos Actually Show
The unredacted filings reveal that both Microsoft and OpenAI knew they were crossing bright legal lines. They didn't just stumble into copyrighted material while crawling the open web. They deliberately targeted paywalled journalism, hoarded it into custom datasets, and actively worked around anti-scraping protections.
Internally, engineers and execs sounded alarms that this practice would gut digital publishers, drain subscription revenue, and kill off the very reporting that keeps large language models grounded in facts. Then, instead of stopping, they accelerated compute clusters and pushed the pipelines harder.
Here's what most coverage misses about these revelations: this wasn't accidental negligence. It was a calculated business risk. Microsoft invested billions into OpenAI knowing the underlying training data rested on shaky legal ground. When you're locked in a brutal race, you grab the data first and settle the lawsuits a decade later.
We've already seen hints of corporate double-talk around safety and oversight before, especially when OpenAI caught its models leaving notes to bypass developer guardrails. But seeing executives admit to blatant copyright theft in their own Slack and email threads is a whole different level of corporate exposure.
The Fair Use Defense Just Evaporated
The entire generative AI boom relies on a single legal theory: scraping the open web without paying creators qualifies as fair use because the resulting weights are "transformative."
Yet fair use requires good faith. When a federal judge reads internal emails where your own vice presidents describe your pipeline as grand larceny, the transformative defense starts to crumble. You can't claim you thought scraping was innocent public indexing when your engineers were swapping warnings that you were devouring the commercial lifeblood of your sources.
The reality is that consumer products are already directly competing with the publishers they mined. When users compare platforms like ChatGPT vs Perplexity to summarize breaking news, they rarely click the publisher's link. They read the synthesized answer and move on. The publisher bears all the reporting costs, while the AI platform pockets the subscription fee.
That said, Microsoft and OpenAI won't go down without burning every procedural delay on the books. They'll argue that internal opinions don't constitute legal admissions. They'll claim individual employees expressed rogue hot takes out of context. But juries hate hypocrisy, and these filings read like an open confession.
The Bill Is Coming Due
What happens next won't be pretty for AI balance sheets. If courts rule that scraping paywalled content violates copyright law at this scale, the remedy won't just be a modest settlement check. It could mean model disgorgement: court orders forcing companies to delete trained model weights and rebuild them from scratch using licensed data.
Silicon Valley banked on being too big to unwind. They thought if they shipped models quickly enough and embedded them into enterprise workflows, courts would be too scared to dismantle the technology.
These unredacted memos puncture that shield. They prove tech giants understood the human cost from day one. And they did it anyway.
Frequently Asked Questions
What did Microsoft executives say about AI scraping?
Unsealed court filings show a Microsoft executive called mass AI data scraping "the largest theft of labor in human history." Internal messages also revealed warnings that scraping paywalled content would severely damage publishers.
Why are these unsealed filings important for The New York Times lawsuit?
The filings undermine Microsoft and OpenAI's fair use defense. Proving that executives knew scraping bypassed paywalls and harmed creators makes it harder to argue the data collection was done in good faith under existing copyright law.
Could this force tech companies to retrain their models?
Yes. If the courts find that copyrighted data was illegally acquired to build flagship AI models, judges can order data destruction or model disgorgement, forcing companies to purge illicitly trained weights.