Turns of the centuries. The Transkribus automated tool for recognition, transcription and translation of handwritten historical documents

Babel ◽  
2020 ◽  
Vol 66 (2) ◽  
pp. 294-310
Author(s):  
Miodrag M. Vukčević

Abstract The translation of handwritten historical documents faces many challenges due to variation in the writing style, local language, and an inevitable language change. Even the transliteration from Cyrillic to Latin characters is standardized by the bijective transliteration standard ISO 9. This presentation introduces a number of tools offered by Transkribus for the automated processing of documents, such as Handwritten Text Recognition (HTR) and Document Understanding, which are needed for the translation of historical documents. Next to the problem of decoding handwritten documents, written for example in Kurrentschrift using ancient terminology, changed meanings and different spelling have additionally to be considered during the translation of texts from earlier centuries. Resolution strategies on a case study show different methods for ensuring quality translations.

2020 ◽  
Vol 6 (5) ◽  
pp. 32 ◽  
Author(s):  
Yekta Said Can ◽  
M. Erdem Kabadayı

Historical document analysis systems gain importance with the increasing efforts in the digitalization of archives. Page segmentation and layout analysis are crucial steps for such systems. Errors in these steps will affect the outcome of handwritten text recognition and Optical Character Recognition (OCR) methods, which increase the importance of the page segmentation and layout analysis. Degradation of documents, digitization errors, and varying layout styles are the issues that complicate the segmentation of historical documents. The properties of Arabic scripts such as connected letters, ligatures, diacritics, and different writing styles make it even more challenging to process Arabic script historical documents. In this study, we developed an automatic system for counting registered individuals and assigning them to populated places by using a CNN-based architecture. To evaluate the performance of our system, we created a labeled dataset of registers obtained from the first wave of population registers of the Ottoman Empire held between the 1840s and 1860s. We achieved promising results for classifying different types of objects and counting the individuals and assigning them to populated places.


2019 ◽  
Vol 94 ◽  
pp. 122-134 ◽  
Author(s):  
Joan Andreu Sánchez ◽  
Verónica Romero ◽  
Alejandro H. Toselli ◽  
Mauricio Villegas ◽  
Enrique Vidal

2021 ◽  
Vol 7 (12) ◽  
pp. 260
Author(s):  
Lazaros Tsochatzidis ◽  
Symeon Symeonidis ◽  
Alexandros Papazoglou ◽  
Ioannis Pratikakis

Offline handwritten text recognition (HTR) for historical documents aims for effective transcription by addressing challenges that originate from the low quality of manuscripts under study as well as from several particularities which are related to the historical period of writing. In this paper, the challenge in HTR is related to a focused goal of the transcription of Greek historical manuscripts that contain several particularities. To this end, in this paper, a convolutional recurrent neural network architecture is proposed that comprises octave convolution and recurrent units which use effective gated mechanisms. The proposed architecture has been evaluated on three newly created collections from Greek historical handwritten documents that will be made publicly available for research purposes as well as on standard datasets like IAM and RIMES. For evaluation we perform a concise study which shows that compared to state of the art architectures, the proposed one deals effectively with the challenging Greek historical manuscripts.


2020 ◽  
Vol 21 (4) ◽  
pp. 40-44
Author(s):  
Dawn Behrend

Sex & Sexuality, Module I: Research Collections from the Kinsey Institute Library & Special Collections published by Adam Matthew Digital is a collection of digitized primary sources obtained exclusively from the Kinsey Institute Library & Special Collections dedicated to the study of human sexuality throughout the twentieth century. The collection makes use of the artificial intelligence capabilities of Handwritten Text Recognition (HTR) to enable keyword searching of handwritten documents. The documents and images in the collection have been meticulously digitized by Adam Matthew Digital making them discoverable, visually appealing, and adjustable. The proprietary interface is intuitive to navigate with the product being compatible with a range of browsers and electronic devices. Contract provisions are standard to the product and permit for use across locations and interlibrary loan sharing. As pricing is primarily determined by size and enrollment, the collection may be affordable for libraries of varying sizes. Users seeking more current research on gender and women’s studies may find ProQuest’s GenderWatch a more suitable choice, while those seeking information on sexuality from the sixteenth to mid-twentieth centuries may prefer Part III of Gale’s Archives of Sexuality & Gender with both resources providing access to a range of sources beyond that of the Kinsey Institute.


Sign in / Sign up

Export Citation Format

Share Document