Qoruyo OCR & HTR
Qoruyo: Models for Automatic Transcription of Manuscripts
The Beth Mardutho Qoruyo project seeks to develop tools and resources for successful optical character recognition (OCR) and handwritten-text recognition (HTR) of printed and handwritten Syriac texts. They take their name from the Syriac word for "reader," a reference to liturgical lectors.
As of 2026, new and updated models have been developed with Kraken for automatic text recognition of both printed texts (Qoruyo) and manuscripts (Sophro Mhiro).
Models
| Qoruyo (Printed Texts) | Sophro Mhiro (Manuscripts) |
|---|---|
| Qoruyo: Syriac model for the segmentation of printed texts: https://doi.org/10.5281/zenodo.17406625 | Sophro Mhiro: Syriac ATR model for the segmentation of manuscripts with text in one column: https://doi.org/10.5281/zenodo.17406717 |
| Qoruyo: Syriac model for the recognition of printed text in Serto: https://doi.org/10.5281/zenodo.17406676 | Sophro Mhiro: Syriac ATR model for the segmentation of manuscripts with text in two columns: https://doi.org/10.5281/zenodo.17406754 |
| Qoruyo: Syriac model for the recognition of printed text in Estrangela: https://doi.org/10.5281/zenodo.17406702 | Sophro Mhiro: Syriac ATR model for the segmentation of manuscripts with text in four columns: https://doi.org/10.5281/zenodo.17406766 |
| Qoruyo: Syriac model for the recognition of printed text in Eastern Syriac script: https://doi.org/10.5281/zenodo.17406689 | Sophro Mhiro: Syriac ATR model for the recognition of manuscript text in Serto, Estrangela and Eastern Syriac scripts: https://doi.org/10.5281/zenodo.17406773 |
Please note: these ATR models are made freely available, but Beth Mardutho is unable to provide technical support to individual users. Users may consult the documentation page on GitHub created by Jimmy Issa.
Models Overview
Qoruyo models work on printed texts. The respective segmentation models recognize the differing structures of a page layout.
Sophro Mhiro models work on handwritten manuscripts. Models have been developed for pages with one, two, and four columns. Users should select the model that best aligns with their manuscript layout.
How to Use
- Users should first select the appropriate ATR model for their document: Qoruyo for printed texts, and Sophro Mhiro for manuscript pages.
- With Qoruyo models, users should select the model that aligns with the font/script of their document (Estrangela, Serto, or Eastern Syriac).
- Users should select the Sophro Mhiro model according to column layout on the manuscript page (single, double, or quadruple columns). Sophro Mhiro models work on all scripts.
- Model specifications are noted in their names.
- Upload images to eScriptorium.
- Select segmenting model to segment the page into lines.
- Select reading model to read the text.
History of the Project
Beth Mardutho has been working on the problem of Syriac Automatic Text Recognition (ATR) for almost a decade. Back in 2016, George Kiraz and Benjamin Kiessling wrote a proof-of-concept with the Kraken engine. In 2018, Digital Humanities Fellows Emily Chesley, Jillian Marcantonio, and Abigail Pearson tested and evaluated Tesseract 4.0 for Syriac OCR.
The following year, Abigail Pearson, Kyle Brunner, and Christine Roughan turned to developing the Transkribus engine out for Syriac manuscripts. Beth Mardutho announced the Qoruyo HTR models for automatic recognition of printed texts in the major Syriac scripts in 2019. These models transcribed handwritten Syriac manuscripts in all three major scripts with up to 98% accuracy. Unfortunately, due to updates in Transkribus, we are currently unable to offer these models to new users. Attention then returned to Kraken as an engine.
The current Qoruyo and Sophro Mhiro models were developed with Kraken, using eScriptorium as an interface. Development of HTR (handwritten text recognition) on manuscripts grew exponentially during a 2024 Transcribathon at Princeton University's Center for Digital Humanities, organized by Christine Roughan, Daniel Stökl Ben Ezra, and George Kiraz. Through the global collaboration of fellows and volunteers, the Sophro Mhiro models for the recognition of manuscript texts were built. These allow automatic text recognition of manuscripts in multiple column formats.
Project Publications
- Stökl Ben Ezra, Daniel, Luigi Bambaci, George Kiraz, Christine Roughan, and Matthieu Freyder. "Steps Towards Mining Manuscript Images for Untranscribed Texts: A Case Study From the Syriac Collection at the Vatican Library." In CHR 2024: Computational Humanities Research Conference 2024, Aarhus, Denmark. (December 4, 2024): 48–64. https://doi.org/10.5281/zenodo.15396841
- Chesley, Emily, Jillian Marcantonio, and Abigail Pearson. "Towards Digital Syriac Corpora: Evaluation of Tesseract 4.0 for Syriac OCR." Hugoye: Journal of Syriac Studies 22, no. 1 (2019): 109–192. https://doi.org/10.31826/hug-2019-220105