Seeing sounds
Your brain constructs a single, unified perception of the world that rarely feels like a patchwork of different senses. Speech perception shows this seamless integration. You don't just hear words; you see them, too. This phenomenon is vividly demonstrated by the McGurk effect, an illusion first described by Harry McGurk and John MacDonald in 1976. If you watch a video of a person mouthing the syllable "ga" while the audio plays the syllable "ba," you will most likely perceive a completely different sound: "da". Your brain fuses the conflicting visual and auditory information to create a new, coherent perception.
This integration happens in a part of the brain called the superior temporal sulcus (STS). The STS acts as a convergence zone, positioned between the visual and auditory cortices. It receives inputs from both areas and combines the sight of mouth movements with the sound of a voice. Research at institutions like Haskins Laboratories in New Haven, which moved to the city in 1970 and has long-standing affiliations with Yale University, has deeply explored these mechanisms. Since the 1950s, Haskins researchers have been at the forefront, developing early speech synthesis technology and the influential "Motor Theory of Speech Perception," which posits that we perceive speech by referencing our own systems for producing it.
A repurposed cortex
The brain's flexibility in processing speech is apparent in individuals who are deaf. Functional magnetic resonance imaging (fMRI) studies show that when a congenitally deaf person lip-reads, their auditory cortex—the part of the brain that processes sound—becomes active. Specifically, areas like the posterior superior temporal gyrus (pSTS) and even parts of the secondary auditory cortex (BA42) show greater activation in deaf participants compared to hearing ones during silent speechreading. It seems the brain, deprived of auditory input, repurposes its "hearing" circuits to process visual speech information. This neural reorganization demonstrates that the function of a brain area is not rigidly fixed.
This cross-modal activity is not limited to those who cannot hear. In hearing individuals, lip-reading silent speech also activates the auditory cortex, though often to a lesser degree. The brain essentially "fills in" the missing sound, synthesizing auditory features from visual cues alone. Magnetoencephalographic (MEG) recordings show that when someone is lip-reading, their auditory cortex activity synchronizes with the rhythm of the unseen, silent speech. This suggests a predictive mechanism where visual information helps the brain anticipate and decode the auditory signal faster and more accurately, an advantage in noisy environments. Studies using sEEG (stereoelectroencephalography) have found that the response in the STS to audiovisual speech is 40% faster and 18% larger than to sound alone.
