Sensation, Perception & Memory · Unit 13

Perception

You don't see reality. You see your brain's best guess about reality — usually right, and revealingly wrong when it isn't.

~15 min · pairs with the Perception lecture

Your slides covered the mechanics — how raw sensory signals get organized into a coherent scene. This page fills in the "so what": why your brain bothers to interpret rather than just record, what happens when the interpretation runs ahead of the data, and why that's a feature, not a bug.

Two directions of processing

Sensation gives the brain raw material — light hitting the retina, pressure waves hitting the eardrum. Perception is what happens next: organizing and interpreting that raw material into something meaningful. Two different processing directions do this work together, and your slides likely named them without dwelling on why the distinction matters.

Bottom-up processing

Starts with the raw sensory data and builds upward — edges, colors, and contrasts get assembled into shapes, then objects. No prior knowledge required. This is how you can recognize something completely novel that you've never seen or thought about before.

Top-down processing

Starts with your knowledge, context, and expectations, and uses them to interpret incoming sensory data faster and with less information. This is how you can read a smudged word in a sentence instantly, because context fills in the gaps.

Top-down processing has a well-known cousin worth naming precisely: perceptual set, a mental predisposition to perceive something in a particular way based on expectation. What you expect to see, hear, or find changes what you actually perceive — not just how you interpret it afterward, but the raw experience itself. Show two groups an ambiguous image after priming them with different words, and they'll report seeing genuinely different things.

Bodhi says

Don't treat bottom-up and top-down as two competing systems fighting for control. They run together, constantly, on almost everything you perceive. The question is never "which one is happening" — it's "how much weight is each one carrying right now."

Check yourself
You instantly read a text message even though half the words are typo'd and missing letters. Which processing direction is doing the heavy lifting?

The modern reframe: your brain as a prediction machine

◆ The update your textbook may be missing

The older way of describing perception treats it almost like a camera with some editing software attached: light comes in, gets processed, and a picture comes out. The current framing in cognitive science is more radical. Under predictive processing (Clark, 2013), the brain is constantly generating a hypothesis about what's out there in the world, and incoming sensory signals are used mainly to correct that hypothesis when it's wrong. Perception, on this view, isn't built from the bottom up out of raw data first — it's the brain's best current guess, continuously checked against the senses. Most of the time the guess and the input line up so well you never notice the guessing is happening at all. Illusions are the moments when they don't.

This reframing folds bottom-up and top-down processing into a single ongoing loop rather than two separate paths: top-down expectation generates the hypothesis, bottom-up sensory input tests it, and perception is whatever comes out of that constant back-and-forth.

Attention as the gate

Recall from the Consciousness unit that attention decides what reaches awareness at all. Perception depends on that gate being open. Inattentional blindness is the failure to notice a fully visible stimulus because attention was allocated elsewhere. The best-known demonstration is the invisible-gorilla study (Simons & Chabris, 1999): participants asked to count basketball passes in a video routinely failed to notice a person in a gorilla suit walk through the middle of the scene, even though it was in plain view the entire time. A close relative, change blindness, is the failure to notice that something in a scene has changed between two views, again because attention wasn't on the changing detail.

Both findings make the same point from different angles: perception is not a passive recording of everything hitting your senses. Without attention pointed at something, it can be right in front of you and still never become part of your conscious experience.

Organizing fragments into wholes

Even before top-down knowledge gets involved, your visual system automatically groups raw fragments into organized wholes. The Gestalt psychologists captured this with a phrase your slides may have quoted: the whole differs from the sum of its parts (Wertheimer, 1923). A handful of dots isn't perceived as a handful of dots — it's perceived as a line, a cluster, or a shape, automatically and without effort.

The Gestalt psychologists identified several specific grouping principles the visual system applies automatically, and they're worth naming individually. By proximity, elements near one another are grouped together. By similarity, elements that look alike — same color or shape — are grouped even when they're spread apart. By closure, we mentally fill in gaps to perceive a complete figure, seeing a whole circle in a ring of dashes. By continuity, we perceive smooth, continuous paths rather than abrupt breaks where lines cross. And the most basic of all is figure-ground organization: we automatically split any scene into a figure that stands out and a ground that recedes behind it — a division that can flip back and forth in deliberately ambiguous images, like the classic vase-versus-two-faces silhouette. None of these require conscious effort; they are the visual system's default settings for imposing order on raw input.

A little history helps these principles hang together. Gestalt psychology began with a train ride: in 1910 Max Wertheimer, watching lights blink on and off, realized that two stationary lights flashed in quick succession are seen as one light moving — the phi phenomenon, the illusion of motion that every film, flipbook, and LED sign depends on. No individual frame contains motion, yet you see it plainly; the perception is something the whole delivers that the parts do not. Wertheimer and his colleagues Wolfgang Köhler and Kurt Koffka built a school around that insight. Their master rule, sitting above all the grouping principles, is the law of Prägnanz (roughly, "good form"): faced with an ambiguous input, the visual system settles on the simplest, most stable, most regular organization available. Five overlapping rings become the Olympic logo, not a tangle of arcs. And the figure–ground demonstration in your slides has a name too — Rubin's vase, the Danish psychologist Edgar Rubin's silhouette that reads as either a white goblet or two black faces in profile, but never both at once, because the visual system assigns the shared contour to only one of them at a time.

Slide titled Gestalt Laws of Grouping with definitions and examples. Proximity: rows of dashes grouped by spacing. Similarity: a grid of dots grouped by color. Continuity: a smooth curve seen as one line rather than segments. Closure: a horse outlined by broken lines that the mind completes.
Four classic Gestalt principles — proximity, similarity, continuity, and closure. The mind organizes fragments into wholes automatically. From the PSY 100 lecture slides (Magee).
Check yourself
What did the invisible-gorilla study (Simons & Chabris, 1999) demonstrate?

Depth: turning a flat retina into a 3D world

Your retina is flat, but your experience of the world is not. Depth perception reconstructs three dimensions from two-dimensional input, using cues that fall into two categories depending on whether they need one eye or two.

Binocular cues (need two eyes)

Retinal disparity — each eye gets a slightly different image, and the size of that difference signals distance. Convergence — the inward turning of the eyes to focus on something close produces a muscular signal the brain uses as a distance cue.

Monocular cues (work with one eye)

Linear perspective (parallel lines appear to converge with distance), interposition or overlap (a closer object blocks part of a farther one), relative size (smaller objects are perceived as farther away), texture gradient (detail gets denser and finer with distance), and light and shadow (shading implies three-dimensional form).

Close one eye and the world doesn't go flat — that's a useful demonstration of just how much of depth perception the monocular cues alone can carry.

Three more monocular cues round out the list, and they're the ones students most often leave off. Aerial perspective: distant hills look bluer, hazier, and lower in contrast than near ones, because you're looking at them through more atmosphere — painters have faked distance this way for five hundred years. Motion parallax: as you move, nearby objects sweep across your view quickly while distant ones barely budge, which is why the fence posts blur past a car window while the mountains sit still. And accommodation: the same lens-shaping muscles you met in the vision unit send the brain a signal about how hard they're working, and that effort is itself a rough cue to distance for anything within arm's reach. On the binocular side, retinal disparity also goes by the name binocular disparity — you'll see both on exams — and you can feel it directly: hold up a finger at arm's length and alternate closing each eye, and the finger jumps against the background. That jump is the raw material of 3-D movies, which simply feed each eye its own slightly offset image. The cells that read disparity need to be exercised early in life to survive, which is why a childhood eye that never worked with its partner can leave someone stereoblind — living, quite competently, on monocular cues alone.

Perceptual constancies: correcting for a changing image

An opening door casts a rapidly changing trapezoid-shaped image on your retina, yet you perceive a rectangular door swinging on a hinge. A friend walking away across a field casts a shrinking image on your retina, yet you don't perceive them shrinking. This is perceptual constancy — the tendency to perceive objects as unchanging (in size, shape, color, or brightness) even as the raw retinal image changes. The brain isn't reporting the retinal image faithfully; it's correcting for distance, angle, and lighting to preserve a stable, useful interpretation of the object itself.

Size constancy in particular works by combining an object's retinal image with its perceived distance — the farther away you judge something to be, the larger the brain infers it must actually be to cast an image of that size. This normally works beautifully, but it can be exploited. The moon illusion — the moon looking strikingly larger near the horizon than high overhead, despite casting an identical-sized retinal image in both positions — is widely explained by exactly this mechanism: horizon cues like trees, buildings, and terrain make the low moon seem farther away, so the brain "corrects" its perceived size upward. Change the perceived distance and you change the perceived size, even though the raw image never changed at all.

When the rules misfire: illusions

Illusions are not evidence that perception is broken or unreliable in general. They're the opposite: they're what happens when the same rules that normally produce accurate perception get applied to a situation engineered to trip them up. Gregory (1997) framed illusions as a window into the assumptions perception is built on — you can learn the rules precisely by studying the cases where they fail.

The Müller-Lyer illusion — two equal-length lines that appear different in length because of the direction their arrow-like fins point — is a classic case. The Ponzo illusion uses converging lines (like linear perspective cues) to make equal-sized objects look different in size. The Ames room is a distorted room, built to look rectangular from one specific viewing point, that makes people standing in different corners appear dramatically different in size.

Illusions aren't confined to vision, either — perception constantly fuses information across the senses, and that fusion can be fooled too. The McGurk effect is the striking demonstration: when people watch a video of a mouth clearly articulating one syllable (say, "ga") while the soundtrack actually plays a different syllable ("ba"), they frequently hear a third syllable entirely ("da") that blends the two (McGurk & MacDonald, 1976). What you consciously hear is reshaped by what you see the lips doing — direct evidence that perception integrates the senses into a single best guess rather than keeping each one in its own lane. Tellingly, knowing the trick doesn't dispel it; the integration happens well below the level you can consciously override.

◆ Not just a lab curiosity — a cross-cultural finding

The Müller-Lyer illusion is weaker in people raised in environments without much "carpentered," rectangular architecture — flat walls, square corners, straight hallways. That pattern suggests the illusion isn't a hardwired universal glitch; it's partly a byproduct of a visual system tuned by a lifetime of experience with rectangular environments. Perception's built-in assumptions are shaped, at least in part, by what kind of world you grew up perceiving.

◆ Where perception meets prejudice

If experience and expectation shape how you see a line, they also shape how you see a person — and that's where this unit stops being an academic exercise. Payne (2001) flashed a Black or White face for a fraction of a second, followed by an object that was either a handgun or a hand tool, and asked participants to classify the object as fast as they could. After a Black face, people identified guns faster and, when rushed, misidentified tools as guns more often. Correll and colleagues (2002) built a video-game version in which armed and unarmed targets appear and participants must decide, in under a second, whether to shoot: participants shot armed Black targets faster than armed White ones, and were more likely to shoot an unarmed Black target by mistake. Nothing about the images changed; what changed was the perceptual set the face activated. These are the same top-down machinery you've been reading about — a prior expectation filling in ambiguous input under time pressure — applied to a decision where the false alarm is a human life. It's the strongest argument this unit can make that "I just saw what was there" is never the whole story.

Seeing faces that aren't there

Pareidolia is the tendency to perceive a specific, familiar pattern — very often a face — in a vague or random stimulus: a face in the clouds, in an electrical outlet, in the front grille of a car. Human vision has processing dedicated to detecting faces quickly and reliably, and that system is tuned to be trigger-happy rather than cautious. A false alarm (seeing a face that isn't there) costs almost nothing. A miss (failing to notice a real face — or, deeper in evolutionary history, a real predator) could cost a great deal. A system built to err on the side of over-detecting patterns is exactly what you'd expect that trade-off to produce.

The myth

Seeing is believing — perception is a faithful, accurate recording of what's actually out there in the world.

What's actually true

Perception is an active, constructive process — a best guess, continuously built and corrected, not a passive recording. That construction is usually so accurate you never notice it happening. But it runs on assumptions and shortcuts that can be systematically and predictably fooled, from a two-line illusion to a whole gorilla walking straight through your field of view unnoticed.

Check yourself
Which pair correctly matches a depth cue to its category?

Check yourself

The one thing to carry out of this unit

Perception feels like direct, effortless access to the world as it really is. It isn't. It's your brain's continuously updated best hypothesis, built from a blend of raw sensory data and everything you already know or expect, filtered through an attentional gate that determines what even gets a chance to become conscious. Most of the time that system is remarkably accurate. The illusions, the missed gorillas, and the faces in the clouds aren't failures of that system — they're the clearest evidence of how it actually works.

References

Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3), 181–204.

Correll, J., Park, B., Judd, C. M., & Wittenbrink, B. (2002). The police officer's dilemma: Using ethnicity to disambiguate potentially threatening individuals. Journal of Personality and Social Psychology, 83(6), 1314–1329. https://doi.org/10.1037/0022-3514.83.6.1314

Gregory, R. L. (1997). Knowledge in perception and illusion. Philosophical Transactions of the Royal Society of London B, 352(1358), 1121–1127.

McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature, 264(5588), 746–748. https://doi.org/10.1038/264746a0

Payne, B. K. (2001). Prejudice and perception: The role of automatic and controlled processes in misperceiving a weapon. Journal of Personality and Social Psychology, 81(2), 181–192. https://doi.org/10.1037/0022-3514.81.2.181

Simons, D. J., & Chabris, C. F. (1999). Gorillas in our midst: Sustained inattentional blindness for dynamic events. Perception, 28(9), 1059–1074.

Wertheimer, M. (1923). Laws of organization in perceptual forms. Psychologische Forschung, 4, 301–350.