Audiovisual speech perception: Moving beyond McGurk

Although listeners use both auditory and visual cues during speech perception, the cognitive and neural bases for their integration remain a matter of debate. One common approach to measuring multisensory integration is to use McGurk tasks, in which discrepant auditory and visual cues produce auditory percepts that differ from those based solely on unimodal input. Not all listeners show the same degree of susceptibility to the McGurk illusion, and these individual differences in susceptibility are frequently used as a measure of audiovisual integration ability. However, despite their popularity, we argue that McGurk tasks are ill-suited for studying the kind of multisensory speech perception that occurs in real life: McGurk stimuli are often based on isolated syllables (which are rare in conversations) and necessarily rely on audiovisual incongruence that does not occur naturally. Furthermore, recent data show that susceptibility on McGurk tasks does not correlate with performance during natural audiovisual speech perception. Although the McGurk effect is a fascinating illusion, truly understanding the combined use of auditory and visual information during speech perception requires tasks that more closely resemble everyday communication.

Download Full-text

Audiovisual Speech Perception

Perception ◽

10.1068/v970029 ◽

1997 ◽

Vol 26 (1_suppl) ◽

pp. 347-347

Author(s):

M Sams

Keyword(s):

Speech Perception ◽

Visual Information ◽

Word Meaning ◽

Source Area ◽

Mcgurk Effect ◽

Audiovisual Speech ◽

Audiovisual Speech Perception ◽

Processing Level ◽

Early Processing ◽

Audiovisual Fusion

Persons with hearing loss use visual information from articulation to improve their speech perception. Even persons with normal hearing utilise visual information, especially when the stimulus-to-noise ratio is poor. A dramatic demonstration of the role of vision in speech perception is the audiovisual fusion called the ‘McGurk effect’. When the auditory syllable /pa/ is presented in synchrony with the face articulating the syllable /ka/, the subject usually perceives /ta/ or /ka/. The illusory perception is clearly auditory in nature. We recently studied the audiovisual fusion (acoustical /p/, visual /k/) for Finnish (1) syllables, and (2) words. Only 3% of the subjects perceived the syllables according to the acoustical input, ie in 97% of the subjects the perception was influenced by the visual information. For words the percentage of acoustical identifications was 10%. The results demonstrate a very strong influence of visual information of articulation in face-to-face speech perception. Word meaning and sentence context have a negligible influence on the fusion. We have also recorded neuromagnetic responses of the human cortex when the subjects both heard and saw speech. Some subjects showed a distinct response to a ‘McGurk’ stimulus. The response was rather late, emerging about 200 ms from the onset of the auditory stimulus. We suggest that the perisylvian cortex, close to the source area for the auditory 100 ms response (M100), may be activated by the discordant stimuli. The behavioural and neuromagnetic results suggest a precognitive audiovisual speech integration occurring at a relatively early processing level.

Download Full-text

Neural Correlates of Modality-Sensitive Deviance Detection in the Audiovisual Oddball Paradigm

Brain Sciences ◽

10.3390/brainsci10060328 ◽

2020 ◽

Vol 10 (6) ◽

pp. 328

Author(s):

Melissa Randazzo ◽

Ryan Priefer ◽

Paul J. Smith ◽

Amanda Nagler ◽

Trey Avery ◽

...

Keyword(s):

Speech Perception ◽

Visual Information ◽

Time Window ◽

Mcgurk Effect ◽

Predictive Processing ◽

Visual Signal ◽

Oddball Paradigm ◽

Audiovisual Speech ◽

Late Time ◽

Audiovisual Speech Perception

The McGurk effect, an incongruent pairing of visual /ga/–acoustic /ba/, creates a fusion illusion /da/ and is the cornerstone of research in audiovisual speech perception. Combination illusions occur given reversal of the input modalities—auditory /ga/-visual /ba/, and percept /bga/. A robust literature shows that fusion illusions in an oddball paradigm evoke a mismatch negativity (MMN) in the auditory cortex, in absence of changes to acoustic stimuli. We compared fusion and combination illusions in a passive oddball paradigm to further examine the influence of visual and auditory aspects of incongruent speech stimuli on the audiovisual MMN. Participants viewed videos under two audiovisual illusion conditions: fusion with visual aspect of the stimulus changing, and combination with auditory aspect of the stimulus changing, as well as two unimodal auditory- and visual-only conditions. Fusion and combination deviants exerted similar influence in generating congruency predictions with significant differences between standards and deviants in the N100 time window. Presence of the MMN in early and late time windows differentiated fusion from combination deviants. When the visual signal changes, a new percept is created, but when the visual is held constant and the auditory changes, the response is suppressed, evoking a later MMN. In alignment with models of predictive processing in audiovisual speech perception, we interpreted our results to indicate that visual information can both predict and suppress auditory speech perception.

Download Full-text

A value-driven McGurk effect: Value-associated faces enhance the influence of visual information on audiovisual speech perception and its eye movement pattern

Attention Perception & Psychophysics ◽

10.3758/s13414-019-01918-x ◽

2020 ◽

Vol 82 (4) ◽

pp. 1928-1941

Author(s):

Xiaoxiao Luo ◽

Guanlan Kang ◽

Yu Guo ◽

Xingcheng Yu ◽

Xiaolin Zhou

Keyword(s):

Speech Perception ◽

Eye Movement ◽

Visual Information ◽

Movement Pattern ◽

Mcgurk Effect ◽

Audiovisual Speech ◽

Audiovisual Speech Perception ◽

A Value

Download Full-text

Weaker McGurk Effect for Rubin’s Vase-Type Speech in People With High Autistic Traits

Multisensory Research ◽

10.1163/22134808-bja10047 ◽

2021 ◽

pp. 1-17

Author(s):

Yuta Ujiie ◽

Kohske Takahashi

Keyword(s):

Speech Recognition ◽

Speech Perception ◽

Visual Information ◽

Autistic Traits ◽

Autism Spectrum ◽

Mcgurk Effect ◽

Visual Speech ◽

Audiovisual Speech ◽

Autism Spectrum Quotient ◽

Audiovisual Speech Perception

Abstract While visual information from facial speech modulates auditory speech perception, it is less influential on audiovisual speech perception among autistic individuals than among typically developed individuals. In this study, we investigated the relationship between autistic traits (Autism-Spectrum Quotient; AQ) and the influence of visual speech on the recognition of Rubin’s vase-type speech stimuli with degraded facial speech information. Participants were 31 university students (13 males and 18 females; mean age: 19.2, SD: 1.13 years) who reported normal (or corrected-to-normal) hearing and vision. All participants completed three speech recognition tasks (visual, auditory, and audiovisual stimuli) and the AQ–Japanese version. The results showed that accuracies of speech recognition for visual (i.e., lip-reading) and auditory stimuli were not significantly related to participants’ AQ. In contrast, audiovisual speech perception was less susceptible to facial speech perception among individuals with high rather than low autistic traits. The weaker influence of visual information on audiovisual speech perception in autism spectrum disorder (ASD) was robust regardless of the clarity of the visual information, suggesting a difficulty in the process of audiovisual integration rather than in the visual processing of facial speech.

Download Full-text

Audiovisual Speech Perception: Acoustic and Visual Phonetic Features Contributing to the McGurk Effect

i-Perception ◽

10.1068/ic768 ◽

2011 ◽

Vol 2 (8) ◽

pp. 768-768

Author(s):

Kaisa Tiippana ◽

Martti Vainio ◽

Mikko Tiainen

Keyword(s):

Speech Perception ◽

Mcgurk Effect ◽

Audiovisual Speech ◽

Audiovisual Speech Perception ◽

Phonetic Features

Download Full-text

Merging auditory and visual phonetic information: A critical test for feedback?

Behavioral and Brain Sciences ◽

10.1017/s0140525x00243240 ◽

2000 ◽

Vol 23 (3) ◽

pp. 327-328 ◽

Cited By ~ 1

Author(s):

Lawrence Brancazio ◽

Carol A. Fowler

Keyword(s):

Speech Perception ◽

Visual Information ◽

Audiovisual Speech ◽

Critical Test ◽

Audiovisual Speech Perception ◽

Phonetic Information ◽

Present Description

The present description of the Merge model addresses only auditory, not audiovisual, speech perception. However, recent findings in the audiovisual domain are relevant to the model. We outline a test that we are conducting of the adequacy of Merge, modified to accept visual information about articulation.

Download Full-text

Brain activity during audiovisual speech perception: An fMRI study of the McGurk effect

Neuroreport ◽

10.1097/00001756-200306110-00006 ◽

2003 ◽

Vol 14 (8) ◽

pp. 1129-1133 ◽

Cited By ~ 97

Author(s):

Jeffery A. Jones ◽

Daniel E. Callan

Keyword(s):

Speech Perception ◽

Brain Activity ◽

Mcgurk Effect ◽

Audiovisual Speech ◽

Audiovisual Speech Perception ◽

Fmri Study

Download Full-text

Sound Location Can Influence Audiovisual Speech Perception When Spatial Attention Is Manipulated

Seeing and Perceiving ◽

10.1163/187847511x557308 ◽

2011 ◽

Vol 24 (1) ◽

pp. 67-90 ◽

Cited By ~ 11

Author(s):

Riikka Möttönen ◽

Kaisa Tiippana ◽

Mikko Sams ◽

Hanna Puharinen

Keyword(s):

Speech Perception ◽

Spatial Attention ◽

Reaction Times ◽

Mcgurk Effect ◽

Visual Speech ◽

Audiovisual Speech ◽

Sound Location ◽

Audiovisual Speech Perception ◽

The Right ◽

Talking Face

AbstractAudiovisual speech perception has been considered to operate independent of sound location, since the McGurk effect (altered auditory speech perception caused by conflicting visual speech) has been shown to be unaffected by whether speech sounds are presented in the same or different location as a talking face. Here we show that sound location effects arise with manipulation of spatial attention. Sounds were presented from loudspeakers in five locations: the centre (location of the talking face) and 45°/90° to the left/right. Auditory spatial attention was focused on a location by presenting the majority (90%) of sounds from this location. In Experiment 1, the majority of sounds emanated from the centre, and the McGurk effect was enhanced there. In Experiment 2, the major location was 90° to the left, causing the McGurk effect to be stronger on the left and centre than on the right. Under control conditions, when sounds were presented with equal probability from all locations, the McGurk effect tended to be stronger for sounds emanating from the centre, but this tendency was not reliable. Additionally, reaction times were the shortest for a congruent audiovisual stimulus, and this was the case independent of location. Our main finding is that sound location can modulate audiovisual speech perception, and that spatial attention plays a role in this modulation.

Download Full-text

Reducing Playback Rate of Audiovisual Speech Leads to a Surprising Decrease in the McGurk Effect

Multisensory Research ◽

10.1163/22134808-00002586 ◽

2018 ◽

Vol 31 (1-2) ◽

pp. 19-38 ◽

Cited By ~ 5

Author(s):

John F. Magnotti ◽

Debshila Basu Mallick ◽

Michael S. Beauchamp

Keyword(s):

Visual Information ◽

Mcgurk Effect ◽

Visual Speech ◽

Large Individual ◽

Audiovisual Speech ◽

Unexpected Finding ◽

Video Playback ◽

Natural Rate ◽

Audiovisual Speech Perception ◽

Bayesian Integration

We report the unexpected finding that slowing video playback decreases perception of the McGurk effect. This reduction is counter-intuitive because the illusion depends on visual speech influencing the perception of auditory speech, and slowing speech should increase the amount of visual information available to observers. We recorded perceptual data from 110 subjects viewing audiovisual syllables (either McGurk or congruent control stimuli) played back at one of three rates: the rate used by the talker during recording (the natural rate), a slow rate (50% of natural), or a fast rate (200% of natural). We replicated previous studies showing dramatic variability in McGurk susceptibility at the natural rate, ranging from 0–100% across subjects and from 26–76% across the eight McGurk stimuli tested. Relative to the natural rate, slowed playback reduced the frequency of McGurk responses by 11% (79% of subjects showed a reduction) and reduced congruent accuracy by 3% (25% of subjects showed a reduction). Fast playback rate had little effect on McGurk responses or congruent accuracy. To determine whether our results are consistent with Bayesian integration, we constructed a Bayes-optimal model that incorporated two assumptions: individuals combine auditory and visual information according to their reliability, and changing playback rate affects sensory reliability. The model reproduced both our findings of large individual differences and the playback rate effect. This work illustrates that surprises remain in the McGurk effect and that Bayesian integration provides a useful framework for understanding audiovisual speech perception.

Download Full-text

Own-race faces promote integrated audiovisual speech information

Quarterly Journal of Experimental Psychology ◽

10.1177/17470218211044480 ◽

2021 ◽

pp. 174702182110444

Author(s):

Yuta Ujiie ◽

Kohske Takahashi

Keyword(s):

Speech Perception ◽

Mcgurk Effect ◽

The Other ◽

Emotional Expressions ◽

Audiovisual Speech ◽

Audiovisual Speech Perception ◽

Facial Identity ◽

Race Effect ◽

Speech Information ◽

Effect Experiment

The other-race effect indicates a perceptual advantage when processing own-race faces. This effect has been demonstrated in individuals’ recognition of facial identity and emotional expressions. However, it remains unclear whether the other-race effect also exists in multisensory domains. We conducted two experiments to provide evidence for the other-race effect in facial speech recognition, using the McGurk effect. Experiment 1 tested this issue among East Asian adults, examining the magnitude of the McGurk effect during stimuli using speakers from two different races (own-race vs. other-race). We found that own-race faces induced a stronger McGurk effect than other-race faces. Experiment 2 indicated that the other-race effect was not simply due to different levels of attention being paid to the mouths of own- and other-race speakers. Our findings demonstrated that own-race faces enhance the weight of visual input during audiovisual speech perception, and they provide evidence of the own-race effect in the audiovisual interaction for speech perception in adults.

Download Full-text