Speech perception is a complex cognitive process that transforms continuous acoustic signals into meaningful linguistic representations, engaging distributed cortical networks across multiple temporal and spatial scales. In particular, the \ac{STG} plays a central role in this transformation, encoding acoustic-phonetic features that form the building blocks of higher-order linguistic structures. Recent advances in high-density intracranial recordings have enabled us to directly record the spiking activity of hundreds of neurons simultaneously, allowing us to study how speech is processed through the dynamical spiking activity of neuronal networks. However, the way in which neuronal spiking activity conveys the conversion of acoustic signals into linguistic components, like phonemes and their sequences that blend into words, remains elusive.
To answer this question, we developed two computational frameworks that enable the understanding of the dynamical interplay of spiking neuronal networks during cognitive functions. First, we developed and applied a novel modeling framework based on diffusion approximation of neuronal point-process network dynamics (\acs{DIFUSNP}), which is a continuous approximation of nonlinear Hawkes piecewise deterministic Markov models (\acs{NH-PDMP}). This formulation preserves the statistical flexibility of \acs{NH-PDMP}, allowing direct fitting of the model to spiking recordings while enabling analytical insights using tools from dynamical systems theory.
Second, we introduced a computational framework for cognitive encoding based on a single, input-driven point attractor in an \ac{RNN}. Using the network's near-critical conditions at rest, this framework can explain a range of cognitive functions, including path integration, working memory, and sequential coding, all of which are essential for a deeper understanding of speech processing.
We demonstrated that our frameworks can effectively replicate canonical spiking behaviors, including tonic/phasic spiking/bursting neurons, prominent network motifs such as winner-take-all competition and input-driven attractors, and sustained spiking behavior. Additionally, we validated them on single-unit recordings from non-human subjects during a motor task, showing that they accurately reflect the task's geometry and sustained activity, thereby demonstrating their utility for examining the dynamics of spiking networks. Ultimately, we applied our frameworks to human \ac{STG} micro-electrode array recordings collected during speech perception tasks. Our findings uncovered a single, phonetically driven point attractor mechanism that preserves the structural properties of vowels, such as F1 vowel formants. Importantly, we found that spiking neurons displayed slowing dynamics during specific phonemes and bigrams, shedding light on the dynamcial mechanisms that help integrate phonetic information into more complex linguistic units.
Overall, these results advance our understanding of how structured spiking population dynamics in the human \ac{STG} support speech processing.