Right Brain · Reading
Why What You Hear Changes What You See
You already know that a bad audio sync makes a film unwatchable, and that you follow a conversation better in a loud bar when you can see the speaker's mouth. What you probably do not know is how far this goes. Under the right conditions, watching a face can make you hear a syllable that nobody said — and being told exactly how the trick works does not switch it off. Take in the whole. Before the parts.
Hearing lips
In 1976 Harry McGurk and John MacDonald reported an accident. They were dubbing speech onto video for a study on infant perception, mismatched the track, and found that adults watching the result heard something that was in neither channel.1 Strong
The canonical version: a face articulating ga, a voice saying ba. Most adult observers report hearing da — a compromise between the two, produced by neither. Close your eyes and you hear ba, correctly. Open them and the third syllable comes back.
The part worth sitting with is not that vision won. It is that there was never a moment when you had access to the losing side. You are not choosing between two reports. You are handed a single conclusion, already reconciled, with the reconciliation hidden.
And it is not a lab curiosity built on a fragile stimulus. Twenty-two years earlier, W. H. Sumby and Irwin Pollack had shown the useful version of the same machinery: being able to see a talker's face improves speech intelligibility in noise, and improves it most when the noise is worst.2 Strong The bar-conversation effect is real, it is large, and it is the reason this cross-wiring exists at all.
It runs the other way too
Vision does not simply dominate. Ladan Shams, Yukiyasu Kamitani and Shinsuke Shimojo flashed a single disc on a screen while playing two brief beeps. Observers saw two flashes.3 Moderate Sound rewrote a visual event — one of the cleanest cases anywhere of hearing overruling seeing.
Which raises the obvious question: what decides who wins? David Alais and David Burr answered it elegantly using the ventriloquist effect — the everyday illusion that a voice comes from the puppet's mouth, or from the actor on screen rather than the speaker beside it. They degraded the visual signal progressively and found that the perceived location shifted from vision-dominated to sound-dominated in the way you would expect if the brain were weighting each sense by its reliability.4 Moderate
So there is no fixed hierarchy of the senses. Vision usually wins at locating things because vision is usually more precise at locating things. Hearing wins at timing, which is why two beeps can split one flash. The system is not deferring to an authority; it is combining estimates, and the more reliable estimate carries more weight.
Two conditions, and one big caveat
Integration is not indiscriminate. It requires the two signals to arrive close enough in time — there is a tolerance window rather than a strict simultaneity requirement, and audiovisual speech is notably more forgiving of sound arriving late than early, which is the asymmetry you would expect from a system built for a world where light beats sound to the observer.5 Moderate
It also needs attentional room. Agnès Alsius and colleagues loaded observers with a demanding second task and found the McGurk effect diminished.6 Moderate Not abolished — reduced. So this is not a purely automatic reflex sealed off from the rest of cognition, which is a useful correction to how the effect is usually described.
Now the caveat, and it is a real one. The McGurk effect is far more variable between people than popular accounts admit. Debshila Basu Mallick, John Magnotti and Michael Beauchamp measured it carefully and found susceptibility ranging from almost never to almost always, depending heavily on the individual and on which particular stimulus was used — while being fairly stable within a person over time.7 Moderate
If you try a McGurk demo and hear nothing unusual, you are not broken and you are not immune to multisensory integration. You are somewhere on a wide distribution — one that the "everybody hears da" framing quietly erases.
Where it happens
Audiovisual speech converges in a specific piece of cortex, and the individual differences point at it. Audrey Nath and Michael Beauchamp found that susceptibility to the illusion tracked activity in the superior temporal sulcus, with stronger STS responses in the people who experienced it more.8 Moderate Beauchamp and colleagues went further and disrupted the region with TMS, which reduced the effect — evidence for a causal role rather than a correlation.9 Moderate
That name should look familiar. Posterior STS is the same neighbourhood implicated in reading a person from a dozen moving dots. A region that specialises in binding signals across time and across channels is doing both jobs, which is a tidier story than perception usually offers.
Where the evidence stands
- Seeing a face changes which syllable you hear. Strong — replicated for nearly fifty years.1
- Seeing the talker improves speech understanding in noise. Strong — large, practical, well replicated.2
- Sound can change how many flashes you see. Moderate — clean demonstration, small samples.3
- The senses are weighted by reliability rather than ranked. Moderate — supported for spatial localisation.4
- Integration needs rough temporal proximity and some attention. Moderate5,6
- Susceptibility varies enormously between people. Moderate — and the variation is often ignored.7
- Superior temporal sulcus is causally involved. Moderate — imaging plus a TMS result.8,9
- Knowing about the illusion lets you switch it off. No evidence — the effect is standardly reported in observers who know exactly what is happening.
- You can train yourself to keep the senses separate. No evidence — nothing here supports it, and we are not going to imply it.
What this means for Right Brain
Right Brain is named for a metaphor, not a hemisphere — the open, whole-first way of looking, as against the narrow, part-by-part one. Multisensory integration is the version of that argument that is hardest to argue with, because it happens below the level where arguing is possible.
A part-by-part system would keep the channels clean: here is what the eyes report, here is what the ears report, now compare them and decide. That system would never hear da. It would also be much worse at understanding you in a loud bar, because it would have no way to let the mouth fill in what the noise took out. The illusion and the ability are not two phenomena. They are the same mechanism, seen from the front and the back.
This is the same lesson as why optical illusions fool everyone, arriving through a different door: the output of perception is an inference the system has already committed to, and your knowing better arrives too late to be consulted. Pair it with Attention vs. Perception for the part attention does and does not control.
The perception games in the app are that idea shrunk to a few quiet minutes: no score, no verdict about your abilities, just your own inference machinery caught briefly in the act. The standing limit applies as always — there is no good evidence that playing perception games improves your everyday seeing or hearing. Broad transfer is the weak link in this whole field, and we covered it properly in Do Brain-Training Games Actually Work? Contested
Watch your visual system build a world
Right Brain is 30 quick perception minigames — illusions, gist-catching, spot-the-change, find-the-target-in-the-noise. It's a calm wellness app, not brain training and not medical advice, and every game carries a clear evidence grade, from Strong to Contested.
It's live on the App Store for iPhone. Browse the full game catalogue or see how we grade the science on the evidence section. Take in the whole. Before the parts.
Get Right Brain on the App StoreReferences
- McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746–748. doi:10.1038/264746a0
- Sumby, W. H., & Pollack, I. (1954). Visual contribution to speech intelligibility in noise. The Journal of the Acoustical Society of America, 26(2), 212–215. doi:10.1121/1.1907309
- Shams, L., Kamitani, Y., & Shimojo, S. (2000). What you see is what you hear. Nature, 408(6814), 788. doi:10.1038/35048669
- Alais, D., & Burr, D. (2004). The ventriloquist effect results from near-optimal bimodal integration. Current Biology, 14(3), 257–262. doi:10.1016/j.cub.2004.01.029
- van Wassenhove, V., Grant, K. W., & Poeppel, D. (2007). Temporal window of integration in auditory-visual speech perception. Neuropsychologia, 45(3), 598–607. doi:10.1016/j.neuropsychologia.2006.01.001
- Alsius, A., Navarra, J., Campbell, R., & Soto-Faraco, S. (2005). Audiovisual integration of speech falters under high attention demands. Current Biology, 15(9), 839–843. doi:10.1016/j.cub.2005.03.046
- Basu Mallick, D., Magnotti, J. F., & Beauchamp, M. S. (2015). Variability and stability in the McGurk effect: contributions of participants, stimuli, time, and response type. Psychonomic Bulletin & Review, 22(5), 1299–1307. doi:10.3758/s13423-015-0817-4
- Nath, A. R., & Beauchamp, M. S. (2012). A neural basis for interindividual differences in the McGurk effect, a multisensory speech illusion. NeuroImage, 59(1), 781–787. doi:10.1016/j.neuroimage.2011.07.024
- Beauchamp, M. S., Nath, A. R., & Pasalar, S. (2010). fMRI-guided transcranial magnetic stimulation reveals that the superior temporal sulcus is a cortical locus of the McGurk effect. The Journal of Neuroscience, 30(7), 2414–2417. doi:10.1523/JNEUROSCI.4865-09.2010
Right Brain is a general wellness app for relaxation and play. It is not a medical device and does not diagnose, treat, or prevent any condition, and it is not brain training. The name is a metaphor for a mode of looking, not a claim about brain hemispheres. Evidence grades reflect our reading of the research; whether perception games transfer to everyday seeing remains unproven as a general claim. Difficulties with hearing, speech, or vision are matters for a qualified clinician, not a game.