Actualité

A Microsoft Exec Once Called AI Scraping ‘Theft’ — Now His Own Words Are Evidence in Court

A 2023 internal memo in which a Microsoft science director branded AI data scraping an unprecedented form of theft has resurfaced in the copyright lawsuit pitting news publishers against OpenAI and Microsoft.

A memo that Microsoft would probably rather forget has landed at the center of one of the most closely watched legal fights over artificial intelligence and copyright. Written in January 2023 by Brent Hecht, who was then Microsoft’s director of Applied Science, the document didn’t mince words: he described the mass harvesting of online content to train AI models as “theft at an unprecedented scale” and suggested it might amount to the largest appropriation of human labor in history.

Those lines are now surfacing in newly unsealed court filings tied to the lawsuit brought by several news organizations, including the New York Times, against OpenAI and Microsoft. The publishers accuse both companies of scraping enormous volumes of copyrighted material to build their AI systems without ever asking permission or paying for it.

What makes the new filings sting a bit more is that they suggest some people inside these companies saw the problem coming. Plaintiffs are pointing directly to Hecht’s memo as proof that concerns about where the training data actually came from were circulating internally as far back as 2023.

A caveat worth keeping in mind: much of what’s now public comes filtered through the plaintiffs’ own legal brief, and many of the underlying documents remain sealed. That makes it hard to know the full context surrounding some of these quotes.

The filings go beyond training data. A separate internal Microsoft presentation reportedly found that Copilot’s AI-generated answers could sharply cut the number of users clicking through to news websites afterward. In tests involving New York Times content specifically, click-through rates dropped by as much as 93% compared to a standard Bing search, according to the court documents.

Hecht apparently framed this as something of a vicious cycle: AI systems lean on publishers’ content to generate better answers, but those same answers can starve the publishers of the traffic they need to survive — even as their work keeps feeding the machine.

The filings also reference a deposition from Microsoft CEO Satya Nadella, who reportedly said that content sitting behind a paywall should require a license before being used to train or power an AI system. He also indicated, according to the filings, that if he’d learned OpenAI was pulling in paywalled material without authorization, Microsoft might have pushed for the models involved to be retrained. That’s a notable admission given how deeply Microsoft is invested in OpenAI, both financially and technologically.

Plaintiffs further allege that OpenAI developed workarounds to grab paywalled articles, including a technique one researcher is said to have discussed internally for getting past the New York Times’ paywall. Court documents reportedly cite training datasets containing tens of thousands of copies of content from American news organizations, and a Common Crawl dataset said to hold more than two million documents from the nytimes.com domain alone.

None of this amounts to a ruling. These remain allegations laid out in an ongoing case, not a finding of legal responsibility against OpenAI or Microsoft.

But the underlying question the case raises is a big one: can a company train an AI model on copyrighted material without asking the people who made it? AI firms typically lean on the U.S. legal doctrine of fair use to justify these practices. Publishers counter that AI-generated answers can effectively replace their content outright, siphoning off both readers and revenue.

These documents won’t settle that argument. What they do show is that the unease publishers have been voicing publicly for years was apparently shared, quietly, by people working inside the AI companies themselves.

Leave a comment

Your email address will not be published.