Proceedings chapter
OA Policy
French

Données et modèles pour le traitement des documents en néolatin : le cas Lambert Daneau

Presented atActes du colloque de l'association francophone des humanités numériques, Paris, 20-22 mai 2026
Published inCrespi, S., Gabay, S., Grandjean, M., Pinche, A., Puren, M. & Saint-Raymond, L. (Ed.), Humanistica 2026, p. 120-133
PublisherParis : Association francophone des humanités numériques
Collection
  • Anthology of Computers and the Humanities; 4
First online date2026-05-21
Abstract

This article presents the construction of a corpus of sixteenth-century commentaries on the Epistles of Paul, based on the digitization of numerous printed works in Neo-Latin. As this subtype of Latin is still underrepresented in existing datasets, it required the development of specific resources for training suitable models. The prepared data and models for ATR post-correction and lemmatization are described here to enable systematic digital exploitation of the historical material.

Keywords
  • Latin philology
  • Neo-latin
  • Layout analysis
  • Automatic text recognition
  • Linguistic normalisation
  • Lemmatisation
Citation (ISO format)
GOY, Floriane et al. Données et modèles pour le traitement des documents en néolatin : le cas Lambert Daneau. In: Humanistica 2026. Crespi, S., Gabay, S., Grandjean, M., Pinche, A., Puren, M. & Saint-Raymond, L. (Ed.). Paris. Paris : Association francophone des humanités numériques, 2026. p. 120–133. (Anthology of Computers and the Humanities) doi: 10.63744/5TcizCXUUTmJ
Main files (1)
Proceedings chapter (Published version)
Identifiers
2views
4downloads

Technical informations

Creation10/06/2026 07:39:34
First validation15/06/2026 07:26:03
Update15/06/2026 07:26:03
Status update15/06/2026 07:26:03
Last indexation15/06/2026 07:26:04
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack