While it may seem easy to understand a written sentence, how does our brain parse information when it’s spoken? A new model helps explain this phenomenon.
When we read the sentence “He left for work,” we can clearly distinguish the different words that make it up, because they are separated by a space. But if, instead of reading, we hear the same sentence spoken by someone, the different parts—which are called “discrete linguistic units,” such as words or syllables—are not as directly and easily discernible.
In fact, what reaches the listener’s ear—the “speech signal”—is not organized into discrete, clearly distinct units, but rather as a continuous, uninterrupted stream. So how do we transform this continuous signal into distinct linguistic units? It is this question—which has driven decades of research on speech perception—that we address in an original mathematical model, recently published in the journal Frontiers in Systems Neuroscience.
Different Models of Speech Perception
In the literature, there are two main classes of speech perception models. Models in the first category, such as TRACE—the classic model in the field—assume that speech segmentation occurs naturally as part of the decoding of the acoustic content of speech: the listener can directly decode the continuous speech stream from the acoustic information contained in the signal, using their knowledge of words and sounds. Segmentation would then be a simple byproduct of decoding.
On the contrary, for the second class of models, there would indeed be a segmentation process (involving the detection of the boundaries of linguistic units) distinct from another process that associates the segments thus obtained with lexical units. This segmentation would rely on the detection of events that mark the boundaries between segments. These two distinct processes would work together in an integrated manner to facilitate the understanding and processing of the continuous flow of speech.
Such mechanisms can be observed in infants who, although they have not yet developed a vocabulary in their language, are nonetheless able, to a certain extent, to segment speech into distinct units.
In line with this second conception of segmentation, developments in neuroscience over the past 15 years have led to new proposals regarding the processes of speech stream segmentation, in connection with synchronization and neural oscillation processes. These processes refer to the coordinated brain activities that occur at different frequencies in our brain. When we listen to speech, our brain must synchronize and organize the various acoustic signals reaching our ears to form a coherent perception of language. Neurons in the brain’s auditory areas oscillate at specific frequencies, and this rhythmic oscillation facilitates the segmentation of the speech stream into discrete units.
A leading model in this field is the neurobiological TEMPO model. TEMPO focuses on the temporal detection of amplitude peaks in the speech signal to determine the boundaries between segments.
This approach is based on neurophysiological data showing that neurons in the auditory cortex are sensitive to the temporal structure of speech, and more specifically on the fact that there are synchronization processes between neural oscillations and syllabic rhythm.
How to Understand a Sentence Amid the Commotion
However, although these models provide a more nuanced and precise perspective on how our brain analyzes and processes the complex acoustic signals of speech, they do not yet explain all the mechanisms involved in speech perception. One open question concerns the role of higher-level knowledge—such as lexical knowledge, that is, knowledge of the words we know—in the process of speech segmentation. More specifically, researchers are still investigating how this knowledge is transmitted and combined with cues extracted from the speech signal to achieve the most robust speech segmentation possible.
Suppose, for example, that a speaker named Bob says the sentence “He left for work” to Alice. If there isn’t too much background noise, if Bob articulates clearly, and if he doesn’t speak too quickly, Alice will have no difficulty understanding the message conveyed by her conversation partner. With no apparent effort, she will have understood that Bob pronounced the individual words il, E, paRti, o, tRavaj (the phonetic transcription of the spoken words in the SAMPA transcription system). In such an “ideal” situation, a model based solely on fluctuations in the signal’s amplitude—without relying on additional knowledge—would suffice for segmentation.
However, in everyday life, the acoustic signal is “noised,” for example by the sounds of car engines, birdsong, or music coming from a neighbor next door. Under these conditions, Alice will have a harder time understanding Bob when he says the same sentence. In this case, it’s likely that Alice would use her knowledge of the language to get an idea of what Bob is likely to say—or not say. This knowledge would allow her to supplement the information provided by the acoustic cues for more effective segmentation.
In fact, Alice knows a great deal about language. She knows that words are strung together in syntactically and semantically acceptable sequences, and that words are made up of syllables, which are themselves made up of smaller linguistic units. Since she speaks the same language as Bob, she even knows very precisely the “standard” durations for generating and producing the speech signal herself. She therefore knows the expected durations of syllables and can rely on this information to aid her segmentation process, especially when she encounters a challenging situation, such as background noise. If ambient noise “suggests” syllable boundaries that do not match her expectations, she can ignore them; conversely, if a noise masks a boundary actually produced by Bob, she can recover it if her predictions suggest one at that moment.
In our article published in the scientific journal *Frontiers in Systems Neuroscience*, we explore these various theories of speech perception. The model we developed includes a module for decoding the spectral content of the speech signal and a temporal control module that guides the segmentation of the continuous speech signal stream. This temporal control module combines, in an innovative way, information sources derived from the signal itself (in accordance with the principles of neural oscillations) and those derived from the listener’s lexical knowledge of syllable durations—regardless of whether the speech signal is disrupted by an extra event or a missing event. We have thus developed various fusion models that allow us either to eliminate irrelevant events caused by acoustic noise—if they do not correspond to consistent prior knowledge—or to recover missing events using linguistic predictions. Simulations using the model confirm that incorporating lexical predictions of syllable durations results in a more robust perceptual system. A variant of the model also explains behavioral observations from a recent experiment in which syllable durations in sentences were manipulated to either match or deviate from naturally expected durations.
In conclusion, in a real-world communication situation, when we find ourselves in an environment where the spoken signal is not distorted, relying on the signal alone is likely sufficient to identify the syllables, as well as the words that make up the signal. On the other hand, when this signal is degraded, our modeling work explains how the brain might draw on additional knowledge—such as what we know about the typical durations of the syllables we produce—to aid in speech perception.
This article is republished from The Conversation under a Creative Commons license. Readthe original article.