Learning, Sleep & Consciousness · Unit 08

Learning

How experience rewires behavior — three engines of learning, and why two of the textbook classics have been rewritten.

~13 min · pairs with the Learning lecture

The lecture walks you through Pavlov's dogs, Skinner's boxes, and Bandura's Bobo doll. This page slows down on the pieces students most often blur together — what actually counts as learning, why "negative" doesn't mean "bad", and why one of psychology's most famous stories about helplessness just got flipped after fifty years — plus the ethical asterisk hanging over its most famous human demonstration.

What "learning" even means

Psychologists define learning narrowly: a relatively permanent change in behavior that results from experience. Both halves of that definition are doing work. "Relatively permanent" rules out temporary states like fatigue or drug intoxication — you behave differently when you're exhausted, but you haven't learned anything. "Due to experience" rules out behavior that's simply wired in from birth.

That second clause is the important contrast. Some behavior doesn't need to be learned at all. A fixed action pattern is a stereotyped, species-wide behavior triggered automatically by a specific cue, no practice required. Konrad Lorenz's goslings offer the classic case: a newly hatched gosling will follow the first moving object it sees during a brief critical window and treat it as its mother — a process called imprinting. That's biology handing an organism a ready-made behavior. Learning, by contrast, is biology leaving room for experience to do the shaping.

Classical conditioning: Pavlov's discovery by accident

Ivan Pavlov wasn't studying learning at all — he was studying digestion, measuring how much dogs salivated to food. He noticed his dogs started salivating before the food arrived, at the sound of the assistant's footsteps. That "problem" became one of psychology's foundational discoveries.

The vocabulary matters because you'll use it all semester. Food is an unconditioned stimulus (US) — it triggers salivation, the unconditioned response (UR), automatically, no learning needed. A bell, on its own, means nothing to the dog; it's a neutral stimulus. Pair the bell with food repeatedly, and the dog starts salivating to the bell alone. The bell has become a conditioned stimulus (CS), and the salivation it now triggers is a conditioned response (CR). Same behavior (salivating), different cause — that's the whole trick of classical conditioning.

A handful of terms describe how that learned association behaves over time. Acquisition is the initial learning phase, while the CS and US are being paired. Extinction is what happens when you keep presenting the CS without the US — the CR fades. But extinction isn't erasure: after a rest period, the CR can reappear weakly on its own, a phenomenon called spontaneous recovery, which tells us the original learning is suppressed, not deleted. Generalization is responding to stimuli similar to the CS as if they were the CS (a dog salivating to a buzzer that sounds like the bell); discrimination is the opposite — learning to respond only to the specific CS and not to similar stimuli.

Bodhi says

A trick for keeping acquisition, extinction, and spontaneous recovery straight: think of it as learning, unlearning, and a ghost. The CR shows up, the CR fades when you stop reinforcing it, and then — days later, unprompted — a faint version of it pops back up anyway. That "ghost" of the response is exactly why extinction is described as suppression, not deletion.

Before During (pairing) After Bell (neutral) → nothing Food (US) → Salivate (UR) Bell + Food → Salivate Bell (CS) alone → Salivate (CR) A neutral stimulus (bell) becomes a conditioned stimulus (CS) only by being paired with something that already causes the response.
Classical conditioning in three snapshots: before pairing, during pairing, and after learning has occurred.

Little Albert — a famous demo with a serious asterisk

The textbook's go-to human example is Watson and Rayner's (1920) "Little Albert" study, in which an infant who initially showed no fear of a white rat was conditioned to cry at the sight of it after the rat was repeatedly paired with a loud, frightening noise. It's presented as proof that even complex emotional responses — fear — can be classically conditioned in humans.

Black-and-white photo of the infant 'Little Albert' sitting calmly between two researchers.
Before conditioning: "Little Albert" shows no fear of the white rat. Public domain, via Wikimedia Commons.
Black-and-white photo of Little Albert crying, with a furry animal in front of him and the researchers nearby.
After the rat was paired with a loud, frightening clang, Albert cried at the sight of it — fear, classically conditioned. Public domain, via Wikimedia Commons.

◆ The update your textbook may be soft-pedaling

Treat Little Albert as a cautionary classic, not settled science. It was a single-subject demonstration that has never been cleanly replicated, and it would be considered ethically indefensible by modern standards — deliberately conditioning fear in an infant with no plan to extinguish it afterward. It gets worse: Fridlund et al. (2012) argue, based on medical records and film analysis, that "Albert" was very likely a neurologically impaired infant, which would undercut the study's implicit claim that this was a healthy, typically developing child. The takeaway isn't "ignore classical conditioning in humans" — it's "this particular founding demonstration is far shakier than the two-sentence textbook version suggests."

Biological preparedness: not all associations are created equal

Classical conditioning can make it sound like any stimulus can be linked to any response, given enough pairings. Garcia and Koelling's (1966) taste-aversion research showed that's false, and the exception reveals something deep about how learning is shaped by evolution.

The myth

You can condition any response to any stimulus equally well, as long as you pair them enough times — conditioning is a blank, general-purpose learning process.

What's actually true

Garcia and Koelling paired flavored water with either radiation-induced nausea or with a light-and-noise combination paired with shock. Rats readily learned to avoid the flavor when it preceded nausea — even with just one pairing and a delay of several hours between taste and sickness — but rarely learned to associate the flavor with shock, or the light/noise with nausea. Taste links to sickness; external cues link to pain. Evolution appears to have "prepared" certain associations (an animal that gets sick after eating something is far better off learning to avoid that food than to avoid whatever sound happened to be playing) and made others hard to form. This is biological preparedness — learning is powerful, but it isn't a blank slate.

Notice how unusual taste aversion is compared with the slow, many-trials picture of classical conditioning you just read about: it can form in a single trial, across a delay of hours (not seconds), and it's selective about which stimuli it will link. That combination — one trial, long delay, narrow stimulus range — is the signature of a system evolution built for a specific job, not a general-purpose association machine.

Operant conditioning: consequences shape behavior

Classical conditioning is about associating two stimuli. Operant conditioning is about associating a behavior with its consequence. Edward Thorndike's law of effect — behaviors followed by satisfying consequences become more likely, behaviors followed by unpleasant consequences become less likely — set up the idea; B. F. Skinner (1938) built the systematic experimental program around it, using the "Skinner box" to study how reinforcement and punishment shape behavior in controlled, measurable ways.

Here's the four-way scheme that trips up more students than anything else in this unit. Two independent questions get asked about any consequence: is something added or removed, and does the behavior increase or decrease. Cross those, and you get four cells — and the word "negative" in this scheme has nothing to do with something being bad.

Reinforcement (behavior increases)

Positive reinforcement — add something desirable to increase behavior (a dog sits, gets a treat). Negative reinforcement — remove something aversive to increase behavior (you fasten your seatbelt, the annoying chime stops). Both strengthen the behavior; they just differ in whether something pleasant showed up or something unpleasant went away.

Punishment (behavior decreases)

Positive punishment — add something aversive to decrease behavior (touch a hot stove, get burned, stop touching stoves). Negative punishment — remove something desirable to decrease behavior (a teenager loses phone privileges after breaking curfew). Both weaken the behavior; again, the "positive/negative" label only describes adding versus removing.

◆ The one line worth memorizing

Positive and negative tell you whether a stimulus was added or removed. Reinforcement and punishment tell you whether the behavior increased or decreased. "Negative reinforcement" is not a punishment, and it is not "bad" — it's still strengthening a behavior, just by taking something unpleasant away.

Check yourself
A driver starts wearing their seatbelt because it makes an irritating chime stop. This is an example of…

Shaping, reinforcers, and schedules

Shaping is how you get a behavior that never occurs on its own: reinforce successive approximations that get closer and closer to the target behavior, rather than waiting for the whole thing to happen spontaneously. It's how animal trainers get a dolphin to jump through a hoop — no dolphin does that by accident.

Reinforcers themselves split into two kinds. Primary reinforcers satisfy a biological need directly — food, water, warmth — and don't need to be learned. Secondary (conditioned) reinforcers get their power through association with primary reinforcers — money buys nothing you can eat directly, but it reliably gets exchanged for things that do, so it becomes reinforcing in its own right.

Reinforcement doesn't have to happen every single time a behavior occurs, and in fact it rarely does outside the lab. Schedules of reinforcement describe the pattern. A fixed-ratio (FR) schedule reinforces after a set number of responses (a punch card: every 10th coffee free). A variable-ratio (VR) schedule reinforces after an unpredictable number of responses, averaging around some value. A fixed-interval (FI) schedule reinforces the first response after a set amount of time has passed. A variable-interval (VI) schedule reinforces the first response after an unpredictable amount of time.

Cumulative-response graph comparing four reinforcement schedules — variable ratio, fixed ratio, variable interval, and fixed interval — with ratio schedules rising steepest and fixed interval showing a scalloped curve.
The four schedules on a cumulative-response graph: ratio schedules (steeper) drive faster responding than interval schedules, and fixed-interval produces the tell-tale post-payoff "scallop." Diagram; rights with the original creator.

◆ Why this shows up in your pocket every day

Variable-ratio schedules produce the highest, steadiest rate of responding of any schedule, and behavior reinforced on a VR schedule is the most resistant to extinction — because the organism never knows whether the very next response is the one that pays off, so it keeps going. Slot machines run on VR schedules by design. So do the pull-to-refresh gesture and the notification badge on your social media apps: you don't know which check will surface something rewarding, so you keep checking. This isn't a metaphor borrowed from psychology — it's the literal mechanism, engineered on purpose.

Check yourself
Which reinforcement schedule produces behavior that is the most resistant to extinction?

Why punishment is a weaker tool than it looks

Punishment can suppress a behavior quickly, which is exactly why it's so tempting to reach for. But it has three well-documented downsides that reinforcement-based approaches avoid. First, punishment tells the organism what not to do without teaching what to do instead — it suppresses without instructing. Second, especially when physical, punishment models aggression as a way of solving problems, which the punished individual may go on to imitate. Third, the suppression is often tied tightly to the punisher and the context — a kid who gets punished for misbehaving by one parent may simply learn "don't do this in front of Dad," rather than "don't do this."

Latent learning: a crack in strict behaviorism

Early behaviorists argued that learning only happens with reinforcement — no reward, no learning. Edward Tolman's (1948) research on rats in mazes challenged that directly. Rats that explored a maze with no reward at all still built up a mental representation of its layout — a cognitive map — and when reinforcement (food) was introduced later, those rats suddenly navigated the maze as efficiently as rats that had been rewarded from the start. The learning had been there all along, just hidden until there was a reason to show it. This is latent learning: learning that occurs without reinforcement and isn't demonstrated until it's needed. It was an early piece of evidence that something like internal, cognitive representation was going on — not just observable stimulus-response bonds — well before the "cognitive revolution" the lecture covers later in the course.

The big update: helplessness isn't what's learned

The classic story, going back to Seligman's original research, goes like this: expose an animal to a stressor it cannot escape or control, and it will later fail to escape even when escape becomes possible — it has supposedly learned to be helpless. For decades, "learned helplessness" was taught as a straightforward case of an animal acquiring passivity through experience, and it became an influential model for understanding depression in humans.

◆ The headline update

Maier and Seligman (2016) — yes, that Seligman, revisiting his own fifty-year-old theory in light of decades of neuroscience — reversed the causal story. Passivity in the face of prolonged, uncontrollable stress turns out to be the default mammalian response, built in rather than acquired. What is actually learned is the detection of control: a circuit involving the ventromedial prefrontal cortex that, when an animal discovers it has control over a stressor, becomes active and inhibits a brainstem region (the dorsal raphe nucleus) that would otherwise drive the passive, "helpless" response. In plain terms: organisms don't learn to give up. They start out prone to giving up under prolonged stress, and what experience teaches them — when it teaches them anything — is that they have some control, which switches the shutdown response off. Helplessness isn't learned. Control is.

This isn't just semantic hair-splitting. It flips which side of the equation needs explaining: instead of asking "why did this organism learn to be helpless," the newer model asks "why did this organism fail to learn that it had control." That's a different question with different implications for intervention, and it's a good example of a well-established finding being reinterpreted, not overturned, as the underlying mechanism became clearer.

Observational learning: watching is enough

Not all learning requires direct experience with reinforcement or punishment at all. Albert Bandura's classic Bobo doll studies (Bandura, Ross, & Ross, 1961) showed that children who simply watched an adult model act aggressively toward an inflatable doll — hitting it, kicking it, using specific novel actions — would later imitate those exact behaviors themselves, without ever being reinforced for doing so. This is observational learning: acquiring a behavior by watching a model perform it, no direct reinforcement of the observer required. It's a reminder that classical and operant conditioning, powerful as they are, don't cover everything — a huge amount of human behavior is picked up secondhand, just by paying attention to what other people do.

A nine-panel photo sequence of a young child imitating an adult's aggressive actions toward an inflatable Bobo doll — hitting, kicking, and striking it.
Bandura's Bobo doll study: children who watched an adult attack the doll reproduced the same specific aggressive acts themselves — no reinforcement required. From Bandura, Ross, & Ross (1961); rights with the original source.
Check yourself
According to Maier and Seligman's (2016) reversal of the classic learned-helplessness story, what does an organism actually learn?

The one thing to carry out of this unit

Learning is not one thing — it's at least three different mechanisms (associating stimuli, associating behaviors with consequences, and watching others) layered on top of a brain that was never a blank slate to begin with. Evolution constrains what associates easily with what; consequences shape behavior in ways that hinge on whether something was added or removed, not on whether it "feels negative"; and even the most famous, decades-old findings — Little Albert, learned helplessness — keep getting re-examined and, sometimes, rewritten. The mechanisms are real and replicable. The stories textbooks tell about them are a work in progress.

References

Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.

Fridlund, A. J., Beck, H. P., Goldie, W. D., & Irons, G. (2012). Little Albert: A neurologically impaired child. History of Psychology, 15(4), 302–327.

Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123–124.

Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty: Insights from neuroscience. Psychological Review, 123(4), 349–367.

Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.

Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208.

Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. Journal of Experimental Psychology, 3(1), 1–14.