Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study

Boros, Emanuela; Ehrmann, Maud; Matteo Romanello; Najem-Meyer, Sven; Kaplan, Frédéric

Boros, Emanuela; Ehrmann, Maud; Matteo Romanello; Najem-Meyer, Sven; Kaplan, Frédéric

2024

Formats

Format
BibTeX
MARCXML
TextMARC
MARC
DublinCore
EndNote
NLM
RefWorks
RIS

Files

Abstract

The quality of automatic transcription of heritage documents, whether from printed, manuscripts or audio sources, has a decisive impact on the ability to search and process historical texts. Although significant progress has been made in text recognition (OCR, HTR, ASR), textual materials derived from library and archive collections remain largely erroneous and noisy. Effective post-transcription correction methods are therefore necessary and have been intensively researched for many years. As large language models (LLMs) have recently shown exceptional performances in a variety of text-related tasks, we investigate their ability to amend poor historical transcriptions. We evaluate fourteen foundation language models against various post-correction benchmarks comprising different languages, time periods and document types, as well as different transcription quality and origins. We compare the performance of different model sizes and different prompts of increasing complexity in zero and few-shot settings. Our evaluation shows that LLMs are anything but efficient at this task. Quantitative and qualitative analyses of results allow us to share valuable insights for future work on post-correcting historical texts with LLMs.

Details

Title Post-correction of Historical Text Transcripts with Large Language Models: An Exploratory Study

Author(s) Boros, Emanuela ; Ehrmann, Maud ; Matteo Romanello ; Najem-Meyer, Sven ; Kaplan, Frédéric

Published in Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2024)

Pages 133-159

Conference The 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature, St Julian's, Malta, March 22, 2024

Date 2024-02-18

Publisher Association for Computational Linguistics

ISBN 979-8-89176-069-1

Keywords

large language models; OCR post-correction; historical texts; evaluation

Additional link Link ACL Anthology

Laboratories DHLAB

Record Appears in Scientific production and competences > CDH - College of Humanities and social sciences > Digital Humanities Institute > DHLAB - Digital Humanities Laboratory
Peer-reviewed publications
Conference Papers
Work produced at EPFL

Record creation date 2024-02-18

Files

Abstract

Details

PDF