Conference on zoom at 5:30 PM by Rachael Griffiths, Post-doctoral researcher at the EPHE – PSL University, and CRCAO.
The PaganTibet project examines the religious traditions preserved in a largely unexplored corpus of over 100,000 manuscript pages, most of which currently exist only as digitised images. Few people have worked on these owing to difficulties of script, language, and the concepts conveyed.
This talk presents the computational pipeline the PaganTibet project is developing to open up this corpus for analysis: from handwritten text recognition (HTR), which transcribes manuscript images into machine-readable text, to natural language processing (NLP) methods that annotate and analyse these texts. Recent advances in digital humanities allow deep-learning methods to automatically detect topics, patterns, and other content features, helping us organise and classify texts whose contents are, for the most part, still unknown. One of the project’s key outputs will be a catalogue of the collection to display and better understand the content of the corpus, helping lay the groundwork to reconstruct the nature of Tibet’s pagan religion.