The Algorithmic Challenges and Opportunities in Tracing Meaning at Scale in Large Historical Document Collections
Filip Ginter, Mikko Tolonen, Anna Plassart, Jenna Kanerva, Kira Hinderks, Yu Wu
DOI: http://dx.doi.org/10.15439/2026F8031
Citation: Filip Ginter, Mikko Tolonen, Anna Plassart, Jenna Kanerva, Kira Hinderks, Yu Wu (2026). The Algorithmic Challenges and Opportunities in Tracing Meaning at Scale in Large Historical Document Collections. In M. Bolanowski, M. Ganzha, M. Grzegorowski, L. Maciaszek, M. Paprzycki, A. Paszkiewicz, D. Ślęzak (eds), Proceedings of the 21st Conference on Computer Science and Intelligence Systems (FedCSIS). ACSIS, Vol. 47, pages 21–24.
Abstract. The capabilities of modern language models are expanding the questions that computational history can pose. In this position paper we argue for a research program organized as a progression of increasingly semantic forms of text reuse, moving from surface-level lexical matching, through cross-lingual translation mining, toward the indirect forms of engagement through which ideas are actually transformed. We briefly review the state of the art at each step of this progression, and identify the open frontier, the least direct kinds of engagement. We argue that recognizing such indirect text reuse requires encoder models that move beyond paraphrase, and that historians should contribute to the development of such models. The research program is a case study in fruitful two-directional collaboration: AI enables new history research, whose demands in turn can shape new AI.
References
- S. F. Altschul, W. Gish, W. Miller, E. W. Myers, and D. J. Lipman, “Basic local alignment search tool,” Journal of Molecular Biology, vol. 215, no. 3, pp. 403–410, 1990. https://dx.doi.org/10.1016/S0022-2836(05)80360-2
- A. Vesanto, A. Nivala, H. Rantala, T. Salakoski, H. Salmi, and F. Ginter, “Applying BLAST to text reuse detection in Finnish newspapers and journals, 1771–1910,” in Proceedings of the NoDaLiDa 2017 Workshop on Processing Historical Language. Gothenburg: Linköping University Electronic Press, 2017, pp. 54–58.
- G. Roe, “Text reuse as cultural practice: Intertextuality in the 18thcentury digital archive,” Digital Enlightenment Studies, vol. 2, no. 1, pp. 1–30, 2024. https://dx.doi.org/10.61147/des.23
- D. A. Smith, R. Cordell, and A. Mullen, “Computational methods for uncovering reprinted texts in antebellum newspapers,” American Literary History, vol. 27, no. 3, pp. E1–E15, 2015. https://dx.doi.org/10.1093/alh/ajv029
- C. Gladstone, R. Horton, and M. Olsen, “TextPAIR (pairwise alignment for intertextual relations),” Software, ARTFL Project, University of Chicago, 2021, developed 2008–2021. https://github.com/ ARTFL-Project/text-pair.
- K. Hinderks, C. Ledins, F. Ginter, and M. Tolonen, “Translation mining: An AI-driven taxonomy of eighteenth-century Anglo-French translation practices,” Historical Methods: A Journal of Quantitative and Interdisciplinary History, pp. 1–25, 2026. https://dx.doi.org/10.1080/01615440.2026.2675558 Advance online publication.
- D. Rosson, E. Mäkelä, V. Vaara, A. Mahadevan, Y. Ryan, and M. Tolonen, “Reception reader: Exploring text reuse in early modern British publications,” Journal of Open Humanities Data, vol. 9, no. 5, pp. 1– 11, 2023. https://dx.doi.org/10.5334/johd.101
- Y. C. Ryan, A. Mahadevan, and M. Tolonen, “A comparative text similarity analysis of the works of Bernard Mandeville,” Digital Enlightenment Studies, vol. 1, no. 1, pp. 28–58, 2023. https://dx.doi.org/10.61147/des.6
- R. Péter, “Uncovering hidden influences: The reception reader as a tool for intellectual historians,” Global Intellectual History, pp. 1–13, 2025. https://dx.doi.org/10.1080/23801883.2025.2474486
- J. Kanerva, C. Ledins, S. Käpyaho, and F. Ginter, “OCR error postcorrection with LLMs in historical documents: No free lunches,” in Proceedings of the Third Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2025). Tallinn, Estonia: University of Tartu Library, Estonia, 2025, pp. 38–47.
- N. Reimers and I. Gurevych, “Making monolingual sentence embeddings multilingual using knowledge distillation,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Online: Association for Computational Linguistics, Nov. 2020. https://dx.doi.org/10.18653/v1/2020.emnlp-main.365 pp. 4512–4525.
- M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. Mazaré, M. Lomeli, L. Hosseini, and H. Jégou, “The Faiss library,” 2024.
- J. Matas, C. Galambos, and J. Kittler, “Robust detection of lines using the progressive probabilistic Hough transform,” Computer Vision and Image Understanding, vol. 78, no. 1, pp. 119–137, 2000. https://dx.doi.org/10.1006/cviu.1999.0831
- Y. Wu, A. Mahadevan, F. Ginter, M. Mathioudakis, and M. Tolonen, “Matching meaning at scale: Evaluating semantic search for 18thcentury intellectual history through the case of Locke,” 2026, to appear in Proceedings of the 6th International Conference on Natural Language Processing for the Digital Humanities (NLP4DH 2026), San Diego, CA.