Microsoft is pushing back against copyright claims from publishers and authors, saying its Copilot chatbot rarely reproduces copyrighted content. As part of a lawsuit, the company provided 8.2 million chat logs to an expert, who found that only 59,545 contained at least 16 words in common with news content used to train the AI model.

The analysis also showed that only 51 instances of "substantial overlap" with content from the Center for Investigative Reporting were found, and just 24 responses contained at least 30 matching words with authors' work. Microsoft argues that these numbers bolster its case for fair use, saying the resulting systems are used for significantly different purposes than the original content.

The New York Times disagreed with Microsoft's conclusions, saying the documents and testimony lead to only one conclusion: Microsoft and OpenAI stole from the Times to make commercial products that substitute for its journalism. The Times' lead counsel, Ian Crosby, said the company looks forward to Microsoft and OpenAI being held accountable for their theft.

Microsoft submitted its filing as it argues for a summary judgement, which would end the case at an early stage. The company's filing comes as part of a larger legal battle between publishers, authors, and tech companies over the use of copyrighted content in AI training datasets.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.