Doctoral thesis
OA Policy
English

Efficient and Scalable Language Models for Training, Inference and Test Time Compute

ContributorsPaliotta, Danièle
Imprimatur date2026-05-08
Defense date2026-02-26
Abstract

Large Language Models (LLMs) have emerged as the central paradigm of modern AI, but their rapid growth in size and capabilities has created pressing challenges in training efficiency, inference speed, and deployment scalability.

This dissertation develops a systematic study of architectural, algorithmic, and system-level techniques to make LLMs more efficient without sacrificing performance.

We investigate methods that tackle bottlenecks across three critical stages.

For pretraining, we introduce fast attention algorithms and novel solutions to mitigate outlier activations, a key obstacle to effective model quantization. For inference, we accelerate generation through advances in speculative decoding, cross-architecture distillation, and the development of subquadratic and hybrid models.

Finally, at test time, we scale reasoning capabilities by designing distilled models that trade additional inference compute for improved accuracy.

We provide new insights into the emergence and mitigation of outlier features in Transformers, novel dynamic sparse attention mechanisms with efficient GPU kernels, and a principled framework for distilling hybrid architectures that bridge Transformers and structured state space models such as Mamba.

We further demonstrate that scaling inference-time compute with distilled reasoning students can rival or surpass strong Transformers at fixed compute budgets for complex reasoning tasks, pointing toward a more cost-effective trajectory for LLM development.

Together, these results provide theoretical and empirical evidence that LLM efficiency can be substantially improved across the training–inference spectrum.

Overall, this work advances the design of scalable, more sustainable foundation models that retain the capabilities of current state-of-the-art systems while reducing their computational bottlenecks.

Citation (ISO format)
PALIOTTA, Danièle. Efficient and Scalable Language Models for Training, Inference and Test Time Compute. Thèse, 2026. doi: 10.13097/archive-ouverte/unige:195581
Main files (1)
Thesis
accessLevelPublic
Secondary files (1)
Imprimatur
accessLevelPublic
Identifiers
6views
76downloads

Technical informations

Creation31/08/2026 10:23:22
First validation01/09/2026 09:21:31
Update29/09/2026 07:26:31
Status update29/09/2026 07:26:31
Last indexation29/09/2026 07:27:29
All rights reserved by Archive ouverte UNIGE and the University of GenevaunigeBlack