Scientific article
OA Policy
English

Is tokenization needed for masked particle modeling?

Published inMachine learning: science and technology, vol. 6, no. 2, 025075
Publication date2025-06-30
First online date2025-06-27
Abstract

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

Keywords
  • Jet
  • Self-supervised learning
  • High-energy physics
  • Conditional generative models
  • Jet physics
Funding
Citation (ISO format)
LEIGH, Matthew et al. Is tokenization needed for masked particle modeling? In: Machine learning: science and technology, 2025, vol. 6, n° 2, p. 025075. doi: 10.1088/2632-2153/addb98
Main files (1)
Article (Published version)
Identifiers
Journal ISSN2632-2153
4views
167downloads

Technical informations

Creation12/03/2026 14:23:02
First validation24/03/2026 13:37:35
Update28/04/2026 12:01:01
Status update28/04/2026 12:01:01
Last indexation28/04/2026 12:01:04
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack