Scientific article
OA Policy
English

Generative AI for automatic topic labelling

Published inCanadian journal of information and library science, vol. 49, no. 2, p. 26-34
First online date2026-06-18
Abstract

Topic modelling has become a prominent tool for the study of scientific fields, as they allow for a large-scale interpretation of research trends. Nevertheless, the output of these models is structured as a list of keywords, which requires a manual interpretation for the labelling. This paper proposes to assess the reliability of three LLMs, namely flan, GPT-4o, and GPT-4 mini for topic labelling. Drawing on previous research leveraging BERTopic, we generate topics from a dataset of all the scientific articles (n=34,797) authored by all biology professors in Switzerland between 2008 and 2020, as recorded in the Web of Science database. We assess the output of the three models both quantitatively and qualitatively and measure the effect of the temperature parameter in GPT models and find that, first, both GPT models are capable of correctly and precisely labelling topics from the models' output keywords at the default temperature. Second, 3-word labels are preferable to grasp the complexity of research topics.

Citation (ISO format)
KOZLOWSKI, Diego, PRADIER, Carolina, BENZ, Pierre. Generative AI for automatic topic labelling. In: Canadian journal of information and library science, 2026, vol. 49, n° 2, p. 26–34. doi: 10.5206/cjils-rcsib.v49i2.24099
Main files (1)
Article (Published version)
Identifiers
Journal ISSN1195-096X
2views
2downloads

Technical informations

Creation20/06/2026 00:33:35
First validation07/07/2026 08:10:01
Update07/07/2026 08:10:01
Status update07/07/2026 08:10:01
Last indexation07/07/2026 08:10:02
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack