Recently unsealed court documents in the New York Times' case against OpenAI and Microsoft reveal that the companies knew their AI models would damage the web. The documents show that Microsoft's Director of Applied Science, Brent Hecht, described the scraping of data to train their models as the 'largest theft of labor in human history' and said it made a 'complete mockery of the idea of fair use.'
Microsoft has tried to distance itself from Hecht's assertions, with spokesperson Alex Haurek saying that 'These comments reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views.' However, it's clear that the situation has come to pass, with Google Zero and AI eating into the web.
The court filing includes quotes from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. Nadella admitted that chatbots have replaced search and removed the need to go straight to the source for information. An internal Microsoft document described the situation as a 'doom loop' that would hurt both their models and the web.
OpenAI cofounder Greg Brockman was more interested in the potential profits, saying the company could make 'gazillions' of dollars through commercial AI. Despite claims of altruistic intent, it seems that OpenAI was aware of ChatGPT's tendency to reproduce copyrighted material 'verbatim.' The company acknowledged that preventing memorization was important to minimize copyright violations, but employees admitted that GPT-4 'memorized a ton of data and therefore will be insanely good at regurgitation.'
The filing includes examples of ChatGPT outputting long strings of copy straight from articles in the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to queries. Microsoft knew how its wholesale scraping of the internet would be perceived and admitted that 'almost no one intended for they [sic] content they created to be used in this fashion, nor are they compensated for its use.' OpenAI's Nick Turley said that once you get an answer from its chatbot, there is 'no good reason to click' on a link to the source.
Microsoft is quoted as admitting that 'LLMs are a product that destroys its own supply chain' because it's a substitute for its own training data in many cases. OpenAI's own media and economic experts attributed the drop in referral traffic for sites like the Times directly to AI summaries like Google's AI Overviews. They speculated that search referrals may be down as much as 60 percent.
It seems clear that both Microsoft and OpenAI knew they were going to irreparably harm the publishing industry and damage their own product, but carried forward anyway in pursuit of profits. The situation has significant implications for the future of the web and the publishing industry, with the rise of zero-click search and the potential for answer engine optimization to further disrupt traditional search models.
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.
