OpenAI and Microsoft Saw Generative AI’s Web Backlash Coming
Introduction
Newly unsealed documents from The New York Times’ lawsuit against OpenAI and Microsoft bring an old contradiction in the generative AI business into sharper focus. AI companies depend on the open web for the material used to train and improve their models. Yet their products can also summarize that material so effectively that users no longer need to visit the sites where it was published.
The filings do more than present arguments about copyright doctrine. They show that people inside the companies had already identified the commercial tension: the same systems that need a healthy web may weaken the publishers, forums, and specialist sites that produce the web’s future training data.
What the documents indicate
- A “doom loop” was identified internally. A Microsoft document warned that its AI content strategy could hurt both model performance and the wider web. If a chatbot answers a question directly, the user has less reason to visit the original page. Lower traffic can mean less advertising, subscription, and referral revenue, which may eventually reduce the supply of new, high-quality material.
- The data-collection criticism was unusually blunt. Microsoft applied science leader Brent Hecht reportedly described the large-scale harvesting of online work for AI training as an extraordinary appropriation of labor and argued that a broad fair-use defense made a mockery of the concept. Microsoft later said these were his own divergent, academic, and forward-looking views rather than the company’s position.
- Paywalls were not clearly handled. The filings contrast Satya Nadella’s stated principle that paywalled material should be licensed with an OpenAI representative’s admission that he was unaware of an effort to systematically identify and remove paywalled content from training data.
- Memorization remained a known problem. Internal material acknowledged that GPT-4 had memorized a large amount of data and could therefore be unusually capable of reproducing it. The filing cites examples in which ChatGPT allegedly returned long passages associated with several publications, underscoring the gap between preventing memorization and avoiding copyright violations.
- Answer products can reduce referrals. OpenAI’s media and economic analysis linked declining referrals to AI summaries and similar search features. One internal view was that, after receiving a sufficient answer, a user may have little reason to click through to the source.
The larger conflict
The dispute is not limited to whether training data qualifies as fair use. It also concerns how value is distributed across the information economy. AI developers absorb material created by others, then package it into a response that may compete with the original publisher’s product. If the publisher loses traffic and revenue, it may have fewer resources to report, research, edit, and publish new work.
That creates a supply-chain paradox. The more successful an AI answer engine becomes at replacing parts of search and information consumption, the more pressure it places on the content producers and distribution platforms that its models depend on. In that sense, the product can weaken the ecosystem supplying its essential inputs.
What happens next
Microsoft has sought to distance itself from the most forceful statements, emphasizing that they do not constitute legal analysis or represent the company as a whole. That response does not erase the business problem documented in the filings. Courts still have to distinguish among training-data use, verbatim model output, and the effect of AI answers on publisher traffic.
The industry also needs workable answers about licensing, paywall detection, attribution, and compensation. The web is unlikely to vanish because of chatbots, but its economic foundation can change gradually if users stop reaching original sources. Generative AI therefore faces a challenge beyond model quality: it must scale without hollowing out the information supply on which its progress depends.
Source: The Verge AI
Comments
Checking sign-in status...
Loading comments...