Text Recognition techniques for premodern Italian and Devanāgarī manuscripts Graduate fellows Priyamvada Nambrath and Eleanor Webb used the open-source Handwritten Text Recognition platform eScriptorium, which interfaces with the Kraken automatic text recognition system, to transcribe premodern manuscripts at the University of Pennsylvania's Schoenberg Institute of Manuscript Studies during their 2025-2026 fellowship. Nambrath worked on three eighteenth- and nineteenth-century Devanāgarī Indic manuscripts (Ms. Coll. 390 Items 1914, 2478, and 1167), while Webb transcribed UPenn Oversize Ms. Codex 1663, a late seventeenth-century Italian manual of mathematics, astrology, and chiromancy. The fellows tested whether HTR could ease labor-intensive manuscript transcription and what tradeoffs the shortcut might carry. Written by Priyamvada Nambrath PhD Candidate, South Asia Studies and Eleanor Webb PhD Candidate, History During our time as Graduate Fellows at the Schoenberg Institute of Manuscript Studies SIMS 2025-2026 , we explored how Handwritten Text Recognition HTR could be used to transcribe premodern manuscripts from Penn’s collections. HTR uses machine learning techniques to build models that can transcribe handwritten manuscripts. We came to the project with different disciplinary backgrounds and geographical specializations. Priya worked on UPenn Ms. Coll. 390 Item 1914 https://find.library.upenn.edu/catalog/9954196113503681?hld id=resource link 0 , Item 2478 https://colenda.library.upenn.edu/catalog/81431-p3sq8qm1n , and Item 1167 https://find.library.upenn.edu/catalog/9963566933503681?hld id=resource link 0 , all Indic manuscripts written in Devanāgarī script from the eighteenth and nineteenth centuries on omenology, devotional poetry and philosophy respectively, while Ellie worked on transcribing UPenn Oversize Ms. Codex 1663 https://find.library.upenn.edu/catalog/9962934883503681?hld id=resource link 0 , a manual of mathematics, astrology, and chiromancy that was produced in Italy in the late seventeenth century. We nevertheless shared an interest as researchers in HTR and in exploring how using it might change the process of researching premodern manuscripts. We worked alongside one another, and under the guidance of Lynn Ransom Curator of SIMS Programs and Schoenberg Database Manager and Jessie Dummer Digitization Project Coordinator , with technical support from Doug Emery Special Collections Digital Content Programmer . Dot Porter SIMS Curator of Digital Humanities , and Matt Hunter Head of Digital Scholarship, Research Data and Digital Scholarship https://www.library.upenn.edu/rdds stepped in as needed with timely technical guidance and support. We used the platform eScriptorium https://escriptorium.rich.ru.nl/ , an open source software that serves as an interface to Kraken https://kraken.re/main/index.html , an automatic text recognition ATR system that is described on its website as “a universal text recognizer for the humanities.” eScriptorium allows users to segment manuscript pages into their individual lines and regions title text, main text, marginalia, and diagrams and to transcribe them. This data is used to create eScriptorium’s ground truth, or manually verified data. Kraken uses the ground truth to learn which characters on the written page correspond to which characters in the alphabet. In eScriptorium, segmentation divides the page into text regions and individual lines, allowing Kraken to learn from line-level images while the separation of different regions provides a basic analysis of page layout. The ground truth, which includes accurate page segmentation and transcriptions, can then be used to create a new transcription model from scratch, or to help refine an existing one. This makes it especially valuable for materials poorly served by conventional optical character recognition OCR , such as handwritten texts. The advantage of eScriptorium https://escriptorium.rich.ru.nl/ ’s open source status is that it is free for users and fosters a culture of open data between researchers and institutions. The processing capacity required for training models, however, required that the software be hosted on the university’s servers. Could HTR alleviate the labor-intensive work of transcribing manuscripts? If so, what might be the downsides of taking this “shortcut”? Could HTR facilitate different kinds of engagement with the premodern manuscripts in Penn collections? Might the availability of HTR technology prompt or enable researchers to ask different questions about these manuscripts? As we began transcribing our manuscripts and testing our respective HTR models, these questions motivated many rich discussions in our weekly meetings. Conducting our transcription projects alongside each other also highlighted HTR’s different capabilities when applied to digitally well-represented scripts like seventeenth-century Italian cursive and to digitally under-represented scripts like Devanāgarī . We spent the first few weeks getting familiar with eScriptorium https://escriptorium.rich.ru.nl/ , learning a lot of new technical language, and beginning to build “ground truth” for our respective manuscripts basically, lots of hours of old fashioned transcribing We first imported IIIF manifests for our respective manuscripts into eScriptorium https://escriptorium.rich.ru.nl/ IIIF, the International Image Interoperability Framework https://iiif.io/ , is a well-established standard for image exchange among institutions . We then asked the platform to “segment” each page. This is a process by which the software recognizes the lines, diagrams, and different regions of the page. We imported an ontology that organized the regions into recognizable categories, such as the main text, marginal text, diagrams, and others. When projects are exported from eScriptorium https://escriptorium.rich.ru.nl/ in XML format, the XML preserves the ontology and records the information about where on the page the handwriting occurs, in which region it occurs, and on which lines – so the output is interoperable with other systems in which researchers may want to display their work. Ellie’s manuscript, Ms. Codex 1663, has a lot of mathematical diagrams and, in the section on chiromancy, several annotated diagrams of hands Figure 1 . For this manuscript, eScriptorium https://escriptorium.rich.ru.nl/ did a fairly good job of identifying the different regions of the page, although it had some difficulty in accurately numbering the lines. Ellie spent some time working through the manuscript and making minor corrections to this segmentation. Priya began her work on Bulbul बुलबुल , UPenn Ms. Coll. 390 Item 1914 https://find.library.upenn.edu/catalog/9954196113503681?hld id=resource link 0 . This manuscript contains a text that is, as far as we know, not present elsewhere. It is an omenological work of performative sortilege, or the practice of fortune-telling by the random drawing of a card from within a collection. Both sides of the eight folios are identical in layout: on the left side is a painting of a bird with its name given below, and on the right side are ten numbered lines of text in the Hindustani language. The first page has a diagram of a bird titled ‘Bulbul,’ and this has been used as the title of the entire manuscript in the catalog record. Figure 2 shows the segmented version of a sample folio side of this manuscript, with the lines and regions indicated. The next step was to begin transcribing the text to create the ground truth. This was the first step at which our projects diverged. Creating ground truth can proceed from scratch – by producing your own transcription – or by fine tuning a model built by other researchers and made available on platforms like HuggingFace by running the model over the manuscript and then correcting the transcription that model produces. The suggested amount of ground truth to build a reliable HTR model from scratch is considered by most estimates to be upwards of 1000 lines, while less is needed for fine-tuning a model. While many models are available for medieval and early modern Latin and other European languages including Italian , we could find no models for Devanāgarī that worked in eScriptorium https://escriptorium.rich.ru.nl/ . Ellie was therefore able to apply a model to Ms. Codex 1663 and then refine it, while Priya was required to build ground truth from scratch for the Bulbul बुलबुल manuscript. She transcribed fifteen pages of the manuscript manually, and then used her transcription to read the sixteenth and final page, which was then used as the validation set for the model. The process of creating ground truth required us to make many decisions that will be familiar to scholars who have produced a transcription or critical edition from a premodern manuscript. Ellie’s manuscript contained several common abbreviations in Italian see Figure 3 . What was distinct about producing a transcription using eScriptorium https://escriptorium.rich.ru.nl/ was the obligation to consider how the model would interpret the abbreviation. Unlike the normal transcription process, which is concerned with rendering the text to be comprehensible to a human reader, we needed to ensure our transcription would enable the computer to learn to recognize the characters depicted on the manuscript and connect them to known letters. This meant taking a different approach to abbreviations, errors, and scribal corrections. For instance, while one might usually transcribe the abbreviation for the word “per” in Italian diplomatically p er , or silently per , doing either in eScriptorium https://escriptorium.rich.ru.nl/ risked confusing the model by including letters that were not in the manuscript. It was therefore necessary to reproduce the text as accurately as possible – by, for instance, simply transcribing “p” for per and “dl” for d e l . In the case of errors and scribal corrections, we tried to reproduce the visual appearance of the manuscript as best as possible in the transcription so that the model could learn to recognize the characters depicted. Ellie used unicode symbols to represent the zodiac and other astrological signs in her manuscript.Where errors or areas of damage rendered the text illegible, we excluded those regions from the computer’s scope of vision to avoid confusion. Priya’s initial work on the Bulbul बुलबुल manuscript https://find.library.upenn.edu/catalog/9954196113503681?hld id=resource link 0 was an experimental effort, as this manuscript was smaller and more manageable than the others she was working with. The eight folios of this document together comprise around 160 lines, which seemed like an achievable amount of input data that needed to be manually input into the model. On transcribing the text over the next couple of weeks, however, she realized that all sixteen sides of the folios essentially had the same textual content even though their individual bird images were different: the same ten fortune outcomes were listed on each face in varying sequences. For purposes of creating the ground truth, this meant that the seven folios had essentially collapsed to just one folio, and the test results reflected this, returning an accuracy rate of just over 5%. As seen in the transcription output in the right of Figure 4 below, the model was unsuccessful in deciphering all the lines on the page, returning the same single, unreadable character for all ten lines. Given the highly restricted character set of the input data supplied to the training model, this quality of output was not really surprising. So this manuscript turned out to be a useful learning exercise in how to use eScriptorium https://escriptorium.rich.ru.nl/ , even though it was unsuccessful in generating a model. In Ellie’s case, she applied a model that she found on Zenodo called catmus-medieval-1.6.0 https://zenodo.org/records/15030337 to her manuscript. This is a Kraken HTR model that was trained on Old and Middle French, Latin, Spanish, and Italian. Interestingly, while this model returned a 94.9% accuracy rate on eScriptorium https://escriptorium.rich.ru.nl/ , a cursory examination of the transcription it produced for Ms. Codex 1663 indicated a far lower accuracy. Ellie transcribed over 50 folios around 2800 lines of the manuscript taken from different sections of the volume in order to produce a representative sample while the scribe is consistent across the manuscript, the handwriting is rather messier in later sections, which also include more complex diagrams . She then trained the catmus model using the transcription. The first subsequent test of the trained model returned an accuracy rate of 93.3% – which was less than before any training While this prompted discussion of how precisely eScriptorium https://escriptorium.rich.ru.nl/ produces its accuracy ratings, it also sent Ellie back on the hunt for models that might be better suited to her seventeenth-century manuscript. The task of searching for appropriate models is a learning experience in itself. While many models are available online, for the uninitiated researcher, it is not always immediately obvious where to find them or which one might be best suited for a particular manuscript. There are also many large and small gaps in the digital representation of premodern scripts, which is reflected in the available models. After her failed experiment with catmus, Ellie found another model – called Tridis v2 Medieval EarlyModern https://zenodo.org/records/13862096 – which had been trained on early modern manuscripts. This model was markedly more successful at reading Ms. Codex 1663, and after several rounds of training produced an eScriptorium https://escriptorium.rich.ru.nl/ accuracy rating of 96.8% and an Ellie accuracy rating of around 80% . For Priya’s second attempt, she decided to select a more well-known Sanskrit manuscript, as this would help to improve the accuracy of her manual transcription. She chose a manuscript copy of the Saundaryalaharī Stotra सौन्दर्यलहरी स्तोत्र , UPenn Ms. Coll. 390 Item 2478 https://colenda.library.upenn.edu/catalog/81431-p3sq8qm1n , a very popular tāntric work of contemplative poetry composed by the 8