AI Turns Oxford’s Archives Into Training Data

AI is entering the archive. Oxford University has allowed OpenAI, the company behind ChatGPT, to train its AI models on historical texts from the Bodleian Library. The material digitised by OpenAI has been used to “populate the OpenAI training set” — a phrase with less romance than the manuscripts themselves.
Oxford announced its partnership with OpenAI in March 2025, saying OpenAI software would digitise texts from the university’s library. That announcement did not say the material would be used to train OpenAI’s models, making the later use an important change in how the project was understood.
By June 2025, OpenAI had received 125,000 images scanned from historical dissertations in the Bodleian collection. The work also includes a rare collection of 10,000 16th-century broadside ballads, while the contract raises the prospect of digitising 23 million items from the Bodleian’s collection.
The current digitisation effort remains “modest in scale” and covers only out-of-copyright material. The Bodleian keeps the rights to the scans and will begin publishing them openly online within months, giving the public access to material that has spent centuries behind institutional walls.
Oxford’s archive meets the training set
Staff at Oxford discussed the reputational risk and environmental impact of partnering with OpenAI. Those concerns sit beside the practical appeal of making historical texts searchable, machine-readable, and available to systems used by more than a billion people using this technology.
An OpenAI spokesperson said, “The AI models of today preserve the world’s historical knowledge for the future.” That promise carries an obvious complication: preservation is not the same as control, especially when a private company trains models on material held by a public-facing academic institution.
OpenAI has struck similar agreements with US research libraries including Boston Public Library, Caltech, MIT, and the University of Michigan through a project called NextGenAI. Oxford’s arrangement therefore sits inside a wider effort to bring research collections into AI development, with libraries supplying the raw material and model companies supplying the machinery.
The Bodleian’s decision to publish the scans openly offers a useful counterweight. Open access will not erase the questions around training data, environmental cost, or institutional reputation, but it means the digitised material will not remain available only through a model whose internal workings are difficult to inspect.
Apollo restores damaged Greek texts
A separate project shows what AI can do once historical material is shaped for a narrower scholarly task. Apollo, a new large language model developed by the Austrian Academy of Science with Mistral and Sail Reply, restores missing words and passages in ancient Greek papyrus fragments.
Apollo was trained on roughly 600 million historical Greek words drawn from manuscripts, papyri, and inscriptions. It can work with Homeric Greek and dialects such as Doric, depending on the document, rather than treating every damaged text as if it belonged to one uniform language.
“When it sees Homer, it supplements Homeric Greek. When it sees an inscription in Doric dialect, it uses Doric dialect,” said Anna Dolganov, a historian and papyrologist at the Austrian Academy of Science. The distinction matters because a plausible restoration is only useful if it fits the language and context of the original text.
Apollo will be freely available to academics through a chatbot interface and is expected to accelerate scholarly work on ancient texts. Armand D’Angour, professor of classical languages and literature at the University of Oxford, said, “If I had a machine telling me, ‘Here are the three possible words that could fit into that gap,’ it would speed up matters considerably.”
The model is unlikely to change the broad understanding of the ancient world. It may uncover new details, though, and D’Angour put the value of those small discoveries plainly: “Every time something is produced, it adds a tiny element of knowledge about the ancient world.”
Dimitris Vlitas, a partner at Sail Reply, said, “Unlocking knowledge in this way was unthinkable a year ago.” The technique used in Apollo could also apply to other ancient languages and disciplines, extending the experiment beyond Greek papyri without pretending that every missing line has one perfect answer.
AI has previously solved a 200-year-old math problem and compiled datasets mapping genetic mutations. These projects point in the same direction: models are becoming tools for finding structure in difficult records, but the value still depends on the quality of the underlying material and the judgment of the people using the result.
Based on




