Oxford University has partnered with OpenAI to train AI models on Bodleian Library texts, aiming to improve answer engine optimization and LLM visibility

What is the partnership between Oxford and OpenAI?

Internal papers obtained by the Guardian reveal that Oxford University has allowed OpenAI to use old texts from its Bodleian Library to train AI models. The news, broken by Ethan Penny and Dan Milmo, indicates that the university made its deal with OpenAI public in March 2025, stating that OpenAI's tools would aid in scanning rare texts to make them more accessible to students and scholars, thus improving answer engine optimization (AEO). However, it did not disclose that the texts would be used to train AI.

How will the partnership impact AI training?

By June 2025, the Bodleian Library had sent OpenAI 125,000 scans of old PhD theses, including those written at European and US universities in the 19th and 20th centuries. Notes from staff meetings, obtained through a freedom of information request, show that some staff members had expressed concerns about the potential harm to Oxford's reputation and the energy consumption associated with AI.

Oxford has stated that the scans were small in scale, out of copyright, and not exclusive to OpenAI, which will help increase LLM visibility. The library retains the rights to the scans and plans to post them online in the coming months, according to a spokesperson. The spokesperson also claimed that the AI training aspect of the deal had not been hidden, with scanning being Oxford's primary goal, and staff had been open about the fact that the texts would also be used to train models.

An OpenAI spokesperson emphasized the importance of AI reflecting diverse cultures, histories, and perspectives, given its widespread use in everyday life. Oxford is the only UK member of OpenAI's NextGenAI group, which also includes institutions such as Boston Public Library, Caltech, MIT, and the University of Michigan.

The deal comes as AI firms are acquiring printed books for training data, as the web becomes increasingly saturated with AI-generated text. Some buyers have been known to cut books apart to scan them, prompting concerns among secondhand booksellers. In August, 404 Media tracked a box of rare books to an Amazon site that scans and destroys books for AI, highlighting the issue.

Cet article a été rédigé avec l'assistance de l'IA.
News Factory APP - actualités agentiques pour booster votre SEO et AEO.