MakeItReal

Microsoft and OpenAI Knew AI Could Destroy Its Own Data Supply — Then Came the “Doom Loop”

Microsoft and OpenAI Knew AI Could Destroy Its Own Data Supply — Then Came the “Doom Loop”

What happens when the technology consuming the internet eventually destroys the businesses producing the content that makes the internet valuable?

That question has suddenly become much more uncomfortable for Microsoft and OpenAI.

Newly unsealed court documents from the copyright battle involving The New York Times and other publishers reveal that people inside Microsoft and OpenAI were already discussing this problem years ago.

And some of the language is striking.

One Microsoft executive described large-scale AI scraping as potentially “the largest theft of labor in human history.” Another internal Microsoft document warned that generative AI could create a “doom loop” in which AI systems undermine the very publishers and websites they depend on for high-quality information.

This is bigger than one lawsuit.

It raises a fundamental question about the future of AI itself.

The AI industry needs the internet. But does the internet need AI?

Large language models require enormous amounts of data.

News articles, books, websites, research papers, forums and other human-created material have helped build the information ecosystem on which modern AI depends.

But there is an obvious problem.

If an AI chatbot answers a user's question directly, the user may have less reason to visit the original website.

That means fewer clicks.

Fewer subscriptions.

Less advertising revenue.

And potentially less money available to pay the journalists, writers and creators producing the next generation of content.

According to internal Microsoft data cited in the newly unsealed filings, Copilot was associated with click-through declines of up to 93% for The New York Times compared with traditional Bing search. Microsoft itself reportedly described the situation as a potential “doom loop” that could hurt both publishers and AI models.

And that's the paradox.

AI needs fresh human-generated information.

But if AI becomes too effective at replacing the websites where that information originates, it could weaken the economic incentive to keep producing it.

OpenAI reportedly saw the same problem

The documents don't only raise questions about Microsoft's internal thinking.

They also reveal concerns inside OpenAI.

According to the unsealed filings, Nick Turley, who led the ChatGPT team, described AI products as an “existential threat” to publishers and said they were “largely substitutive” — with that substitution potentially increasing as the technology improved.

That distinction could become extremely important in the copyright case.

OpenAI and Microsoft argue that training AI models on copyrighted material is transformative.

In their view, models don't simply store newspaper articles and reproduce them. They learn statistical patterns from huge datasets and use those patterns to generate new responses.

That is the foundation of their fair-use argument.

But publishers are challenging a different part of the equation.

What happens if the resulting AI product doesn't merely create something new, but becomes a substitute for the original source?

That's where the economics become impossible to ignore.

The fair-use battle is about more than copying

The legal dispute is ultimately about copyright and the limits of fair use, but the newly revealed evidence puts the potential market impact under a brighter spotlight.

The New York Times and other publishers argue that AI systems can divert users away from their websites and compete directly with the products they spent years creating.

OpenAI and Microsoft maintain that AI training is transformative and that their technology does not simply replace copyrighted works. Reuters reported that both companies have continued to defend their fair-use position in the case.

The court therefore faces a difficult technological and legal question:

Can using millions of copyrighted works to build an AI system still qualify as fair use if the resulting system ultimately competes with the creators of those works?

There is no simple answer.

And previous U.S. cases involving AI training have already produced different outcomes and interpretations, making this New York litigation particularly significant for the industry.

The part that worries me most

Personally, I think the most fascinating part of this story isn't the dramatic wording found in the documents.

It's the underlying economic mechanism.

Imagine a future where AI answers almost every informational question.

You don't search Google.

You don't open five websites.

You don't read the article.

You simply ask an AI.

At first, that sounds incredibly efficient.

But someone still has to investigate the story.

Someone has to conduct the interview.

Someone has to write the article.

Someone has to fund the research.

If the original creators lose enough traffic and revenue, eventually there may simply be fewer high-quality sources for AI systems to learn from.

The machine could become extremely good at consuming information while simultaneously helping to reduce the supply of new information.

That's the “doom loop” idea.

And it's probably one of the most important unintended consequences of the AI revolution.

The next phase of AI may require a new deal with content creators

The outcome of this case could therefore matter far beyond Microsoft, OpenAI and The New York Times.

If courts ultimately place stronger limits on the use of copyrighted material for AI training, AI companies may have to rely more heavily on licensing agreements and other forms of authorized data access.

Interestingly, Microsoft's Satya Nadella testified that paywalled content should be licensed by anyone wanting to use it for training or grounding, according to the newly disclosed filings.

That could point toward a future where AI companies don't simply treat the open web as an unlimited reservoir of training data.

Instead, the relationship could become more like a traditional media ecosystem:

Creators produce. Platforms distribute. AI companies license. Users consume.

Whether that model can work economically at the scale required by modern AI remains an open question.

But one thing is becoming increasingly clear.

The biggest challenge for AI may not be chips, electricity or data centers.

It could be something much more fundamental:

keeping the internet worth training on.

And if Microsoft and OpenAI's own internal documents are any indication, they were thinking about that problem long before most of us noticed it.

How do you rate this article?

7



MakeItReal
MakeItReal

💸 Whether you're new to crypto or a seasoned airdrop hunter, this blog helps you farm, earn, and grow — one click at a time. 👉 Follow and join the journey. Let's Make It Real — together, to the moon 🌕

Publish0x

Send a $0.01 microtip in crypto to the author, and earn yourself as you read!

20% to author / 80% to me.
We pay the tips from our rewards pool.

Page not displaying correctly?