Learning, Sleep & Consciousness · Unit 08

Learning

How experience rewires behavior — three engines of learning, and why two of the textbook classics have been rewritten.

~16 min · pairs with the Learning lecture

The lecture walks you through Pavlov's dogs, Skinner's boxes, and Bandura's Bobo doll. This page slows down on the pieces students most often blur together — what actually counts as learning, why "negative" doesn't mean "bad", and why one of psychology's most famous stories about helplessness just got flipped after fifty years — plus the ethical asterisk hanging over its most famous human demonstration.

What "learning" even means

Psychologists define learning narrowly: a relatively permanent change in behavior that results from experience. Both halves of that definition are doing work. "Relatively permanent" rules out temporary states like fatigue or drug intoxication — you behave differently when you're exhausted, but you haven't learned anything. "Due to experience" rules out behavior that's simply wired in from birth.

That second clause is the important contrast. Some behavior doesn't need to be learned at all. A fixed action pattern is a stereotyped, species-wide behavior triggered automatically by a specific cue, no practice required. Konrad Lorenz's goslings offer the classic case: a newly hatched gosling will follow the first moving object it sees during a brief critical window and treat it as its mother — a process called imprinting. That's biology handing an organism a ready-made behavior. Learning, by contrast, is biology leaving room for experience to do the shaping.

Classical conditioning: Pavlov's discovery by accident

Ivan Pavlov wasn't studying learning at all — he was studying digestion, measuring how much dogs salivated to food. He noticed his dogs started salivating before the food arrived, at the sound of the assistant's footsteps. That "problem" became one of psychology's foundational discoveries.

The vocabulary matters because you'll use it all semester. Food is an unconditioned stimulus (US) — it triggers salivation, the unconditioned response (UR), automatically, no learning needed. A bell, on its own, means nothing to the dog; it's a neutral stimulus. Pair the bell with food repeatedly, and the dog starts salivating to the bell alone. The bell has become a conditioned stimulus (CS), and the salivation it now triggers is a conditioned response (CR). Same behavior (salivating), different cause — that's the whole trick of classical conditioning.

A handful of terms describe how that learned association behaves over time. Acquisition is the initial learning phase, while the CS and US are being paired. Extinction is what happens when you keep presenting the CS without the US — the CR fades. But extinction isn't erasure: after a rest period, the CR can reappear weakly on its own, a phenomenon called spontaneous recovery, which tells us the original learning is suppressed, not deleted. Generalization is responding to stimuli similar to the CS as if they were the CS (a dog salivating to a buzzer that sounds like the bell); discrimination is the opposite — learning to respond only to the specific CS and not to similar stimuli.

Bodhi says

A trick for keeping acquisition, extinction, and spontaneous recovery straight: think of it as learning, unlearning, and a ghost. The CR shows up, the CR fades when you stop reinforcing it, and then — days later, unprompted — a faint version of it pops back up anyway. That "ghost" of the response is exactly why extinction is described as suppression, not deletion.

Infographic of classical conditioning in three panels. Before conditioning: food (unconditioned stimulus) makes the dog salivate (unconditioned response); a bell (neutral stimulus) produces no response. During conditioning: the bell is rung just before the food, and the dog salivates. After conditioning: the bell alone (conditioned stimulus) makes the dog salivate (conditioned response).
Classical conditioning in three snapshots: before pairing, during pairing, and after learning has occurred. From the PSY 100 lecture slides (Magee).

◆ Update: it's about prediction, not just pairing

The Pavlov story makes conditioning sound like simple co-occurrence: ring the bell alongside the food often enough and the association forms. Robert Rescorla (1988) showed the real story is subtler and more interesting — what matters is not how often the CS and US are paired, but how well the CS actually predicts the US. If a bell rings 10 times with food, but food also appears just as often without the bell, the animal learns very little, because the bell carries no genuine predictive information. Conditioning, on this view, is the brain detecting real contingency in the environment — working out what reliably signals what — rather than mechanically stamping in whatever happens to co-occur. That reframing (captured in the title of Rescorla's paper, "Pavlovian conditioning: it's not what you think it is") helped push learning theory toward the information-processing, cognitive view of even this most seemingly automatic kind of learning.

Two more wrinkles round out the classical-conditioning toolkit. First, order matters. Conditioning works best with forward pairing — the CS comes just before the US, so the bell announces the food. Present them at exactly the same moment (simultaneous pairing) and learning is weak; present the food first and the bell after (backward pairing) and almost nothing is learned, because a bell that arrives after the food predicts nothing. That's Rescorla's point again: the CS has to carry information about what's coming. Second, a conditioned stimulus can itself do the conditioning. Once the bell reliably triggers salivation, pair a new neutral stimulus — a light — with the bell (no food at all), and the dog will eventually salivate to the light. This is higher-order conditioning (also called second-order conditioning; Pavlov, 1927). It's how a doctor's office can come to trigger nausea in a chemotherapy patient, and then the sight of a syringe in that office can trigger it too. The chain doesn't go on forever — a third-order stimulus rarely picks up much of the response — but two links are enough to let a single reflex spread to a whole web of cues.

Little Albert — a famous demo with a serious asterisk

The textbook's go-to human example is Watson and Rayner's (1920) "Little Albert" study, in which an infant who initially showed no fear of a white rat was conditioned to cry at the sight of it after the rat was repeatedly paired with a loud, frightening noise. It's presented as proof that even complex emotional responses — fear — can be classically conditioned in humans.

Black-and-white photo of the infant 'Little Albert' sitting calmly between two researchers.
Before conditioning: "Little Albert" shows no fear of the white rat. Public domain, via Wikimedia Commons.
Black-and-white photo of Little Albert crying, with a furry animal in front of him and the researchers nearby.
After the rat was paired with a loud, frightening clang, Albert cried at the sight of it — fear, classically conditioned. Public domain, via Wikimedia Commons.

◆ The update your textbook may be soft-pedaling

Treat Little Albert as a cautionary classic, not settled science. It was a single-subject demonstration that has never been cleanly replicated, and it would be considered ethically indefensible by modern standards — deliberately conditioning fear in an infant with no plan to extinguish it afterward. It gets worse: Fridlund et al. (2012) argue, based on medical records and film analysis, that "Albert" was very likely a neurologically impaired infant, which would undercut the study's implicit claim that this was a healthy, typically developing child. The takeaway isn't "ignore classical conditioning in humans" — it's "this particular founding demonstration is far shakier than the two-sentence textbook version suggests."

Watson's bigger idea — and the ad on your feed

Little Albert wasn't a one-off stunt. John B. Watson is usually credited as the founder of behaviorism, the school of thought holding that psychology should study only observable behavior — stimulus in, response out — and leave unmeasurable "mental processes" alone. Where Pavlov had conditioned reflexes in dogs, Watson's bet was that the same machinery could condition human emotions. After leaving academia he took that bet to Madison Avenue, and advertising has run on it ever since. A car commercial pairs the car (neutral stimulus) with an attractive model, a soaring soundtrack, or a sun-drenched coastline (unconditioned stimuli that already produce good feelings). After enough pairings, the car alone produces a little of that glow — a conditioned response you never chose to learn. The same logic explains why a sponsor drops an athlete the moment a scandal breaks: the athlete has stopped being a US for positive feelings, so pairing them with the product no longer helps and may actively hurt. The next time an ad shows you almost nothing about the product and a great deal about a lifestyle, you're watching Pavlov's dogs in a thirty-second cut.

Biological preparedness: not all associations are created equal

Classical conditioning can make it sound like any stimulus can be linked to any response, given enough pairings. Garcia and Koelling's (1966) taste-aversion research showed that's false, and the exception reveals something deep about how learning is shaped by evolution.

The myth

You can condition any response to any stimulus equally well, as long as you pair them enough times — conditioning is a blank, general-purpose learning process.

What's actually true

Garcia and Koelling paired flavored water with either radiation-induced nausea or with a light-and-noise combination paired with shock. Rats readily learned to avoid the flavor when it preceded nausea — even with just one pairing and a delay of several hours between taste and sickness — but rarely learned to associate the flavor with shock, or the light/noise with nausea. Taste links to sickness; external cues link to pain. Evolution appears to have "prepared" certain associations (an animal that gets sick after eating something is far better off learning to avoid that food than to avoid whatever sound happened to be playing) and made others hard to form. This is biological preparedness — learning is powerful, but it isn't a blank slate.

Notice how unusual taste aversion is compared with the slow, many-trials picture of classical conditioning you just read about: it can form in a single trial, across a delay of hours (not seconds), and it's selective about which stimuli it will link. That combination — one trial, long delay, narrow stimulus range — is the signature of a system evolution built for a specific job, not a general-purpose association machine.

Operant conditioning: consequences shape behavior

Classical conditioning is about associating two stimuli. Operant conditioning is about associating a behavior with its consequence. Edward Thorndike's law of effect — behaviors followed by satisfying consequences become more likely, behaviors followed by unpleasant consequences become less likely — set up the idea; B. F. Skinner (1938) built the systematic experimental program around it, using the "Skinner box" to study how reinforcement and punishment shape behavior in controlled, measurable ways.

Here's the four-way scheme that trips up more students than anything else in this unit. Two independent questions get asked about any consequence: is something added or removed, and does the behavior increase or decrease. Cross those, and you get four cells — and the word "negative" in this scheme has nothing to do with something being bad.

Reinforcement (behavior increases)

Positive reinforcement — add something desirable to increase behavior (a dog sits, gets a treat). Negative reinforcement — remove something aversive to increase behavior (you fasten your seatbelt, the annoying chime stops). Both strengthen the behavior; they just differ in whether something pleasant showed up or something unpleasant went away.

Punishment (behavior decreases)

Positive punishment — add something aversive to decrease behavior (touch a hot stove, get burned, stop touching stoves). Negative punishment — remove something desirable to decrease behavior (a teenager loses phone privileges after breaking curfew). Both weaken the behavior; again, the "positive/negative" label only describes adding versus removing.

◆ The one line worth memorizing

Positive and negative tell you whether a stimulus was added or removed. Reinforcement and punishment tell you whether the behavior increased or decreased. "Negative reinforcement" is not a punishment, and it is not "bad" — it's still strengthening a behavior, just by taking something unpleasant away.

Check yourself
A driver starts wearing their seatbelt because it makes an irritating chime stop. This is an example of…

Shaping, reinforcers, and schedules

Shaping is how you get a behavior that never occurs on its own: reinforce successive approximations that get closer and closer to the target behavior, rather than waiting for the whole thing to happen spontaneously. It's how animal trainers get a dolphin to jump through a hoop — no dolphin does that by accident.

Reinforcers themselves split into two kinds. Primary reinforcers satisfy a biological need directly — food, water, warmth — and don't need to be learned. Secondary (conditioned) reinforcers get their power through association with primary reinforcers — money buys nothing you can eat directly, but it reliably gets exchanged for things that do, so it becomes reinforcing in its own right.

Reinforcement doesn't have to happen every single time a behavior occurs, and in fact it rarely does outside the lab. Schedules of reinforcement describe the pattern. A fixed-ratio (FR) schedule reinforces after a set number of responses (a punch card: every 10th coffee free). A variable-ratio (VR) schedule reinforces after an unpredictable number of responses, averaging around some value. A fixed-interval (FI) schedule reinforces the first response after a set amount of time has passed. A variable-interval (VI) schedule reinforces the first response after an unpredictable amount of time.

Before any of those four schedules, there's a simpler one. Continuous reinforcement — a reward every single time the behavior occurs — is the fastest way to teach a new behavior, which is why a trainer treats a puppy for every sit at first. But behavior learned that way is also fragile: stop the treats and the sitting stops fast, because the change is obvious. Once the behavior is established, trainers deliberately switch to partial (intermittent) reinforcement, rewarding only some of the responses. The payoff is the partial reinforcement effect: behavior maintained on a partial schedule is far more resistant to extinction than behavior reinforced continuously, because the organism has already learned that unrewarded responses are normal and can't tell when the rewards have stopped for good. That's the same reason a toddler whose tantrums are "usually" ignored but occasionally rewarded is harder to extinguish than one whose tantrums were rewarded every time — and it's why the four partial schedules you just read about matter so much.

Cumulative-response graph comparing four reinforcement schedules — variable ratio, fixed ratio, variable interval, and fixed interval — with ratio schedules rising steepest and fixed interval showing a scalloped curve.
The four schedules on a cumulative-response graph: ratio schedules (steeper) drive faster responding than interval schedules, and fixed-interval produces the tell-tale post-payoff "scallop." Diagram; rights with the original creator.

◆ Why this shows up in your pocket every day

Variable-ratio schedules produce the highest, steadiest rate of responding of any schedule, and behavior reinforced on a VR schedule is the most resistant to extinction — because the organism never knows whether the very next response is the one that pays off, so it keeps going. Slot machines run on VR schedules by design. So do the pull-to-refresh gesture and the notification badge on your social media apps: you don't know which check will surface something rewarding, so you keep checking. This isn't a metaphor borrowed from psychology — it's the literal mechanism, engineered on purpose.

Check yourself
Which reinforcement schedule produces behavior that is the most resistant to extinction?

Operant conditioning has its own biological limits, mirroring the taste-aversion constraint on classical conditioning. Keller and Marian Breland — former students of Skinner who went into professional animal training — found that animals trained with reinforcement would sometimes drift away from the trained behavior and back toward instinctive, species-typical actions. Pigs trained to carry coins to a "piggy bank" for food would increasingly stop and root at the coins instead; raccoons trained to drop coins in a slot would compulsively rub them together, as though washing food. The Brelands called this instinctive drift (Breland & Breland, 1961): reinforcement is powerful, but it operates on an animal's evolved instincts rather than on a blank slate, and when the two conflict, instinct often wins. It's the operant echo of biological preparedness — a reminder that conditioning is always negotiating with biology, never simply overwriting it.

Why punishment is a weaker tool than it looks

Punishment can suppress a behavior quickly, which is exactly why it's so tempting to reach for. But it has three well-documented downsides that reinforcement-based approaches avoid. First, punishment tells the organism what not to do without teaching what to do instead — it suppresses without instructing. Second, especially when physical, punishment models aggression as a way of solving problems, which the punished individual may go on to imitate. Third, the suppression is often tied tightly to the punisher and the context — a kid who gets punished for misbehaving by one parent may simply learn "don't do this in front of Dad," rather than "don't do this."

Latent learning: a crack in strict behaviorism

Early behaviorists argued that learning only happens with reinforcement — no reward, no learning. Edward Tolman's (1948) research on rats in mazes challenged that directly. Rats that explored a maze with no reward at all still built up a mental representation of its layout — a cognitive map — and when reinforcement (food) was introduced later, those rats suddenly navigated the maze as efficiently as rats that had been rewarded from the start. The learning had been there all along, just hidden until there was a reason to show it. This is latent learning: learning that occurs without reinforcement and isn't demonstrated until it's needed. It was an early piece of evidence that something like internal, cognitive representation was going on — not just observable stimulus-response bonds — well before the "cognitive revolution" the lecture covers later in the course.

The big update: helplessness isn't what's learned

The classic story, going back to Seligman's original research, goes like this: expose an animal to a stressor it cannot escape or control, and it will later fail to escape even when escape becomes possible — it has supposedly learned to be helpless. For decades, "learned helplessness" was taught as a straightforward case of an animal acquiring passivity through experience, and it became an influential model for understanding depression in humans.

◆ The headline update

Maier and Seligman (2016) — yes, that Seligman, revisiting his own fifty-year-old theory in light of decades of neuroscience — reversed the causal story. Passivity in the face of prolonged, uncontrollable stress turns out to be the default mammalian response, built in rather than acquired. What is actually learned is the detection of control: a circuit involving the ventromedial prefrontal cortex that, when an animal discovers it has control over a stressor, becomes active and inhibits a brainstem region (the dorsal raphe nucleus) that would otherwise drive the passive, "helpless" response. In plain terms: organisms don't learn to give up. They start out prone to giving up under prolonged stress, and what experience teaches them — when it teaches them anything — is that they have some control, which switches the shutdown response off. Helplessness isn't learned. Control is.

This isn't just semantic hair-splitting. It flips which side of the equation needs explaining: instead of asking "why did this organism learn to be helpless," the newer model asks "why did this organism fail to learn that it had control." That's a different question with different implications for intervention, and it's a good example of a well-established finding being reinterpreted, not overturned, as the underlying mechanism became clearer.

Observational learning: watching is enough

Not all learning requires direct experience with reinforcement or punishment at all. Albert Bandura's classic Bobo doll studies (Bandura, Ross, & Ross, 1961) showed that children who simply watched an adult model act aggressively toward an inflatable doll — hitting it, kicking it, using specific novel actions — would later imitate those exact behaviors themselves, without ever being reinforced for doing so. This is observational learning: acquiring a behavior by watching a model perform it, no direct reinforcement of the observer required. It's a reminder that classical and operant conditioning, powerful as they are, don't cover everything — a huge amount of human behavior is picked up secondhand, just by paying attention to what other people do.

A nine-panel photo sequence of a young child imitating an adult's aggressive actions toward an inflatable Bobo doll — hitting, kicking, and striking it.
Bandura's Bobo doll study: children who watched an adult attack the doll reproduced the same specific aggressive acts themselves — no reinforcement required. From Bandura, Ross, & Ross (1961); rights with the original source.

Watching, though, isn't automatically learning — you've watched thousands of things you couldn't reproduce. Bandura (1977) laid out four steps that all have to happen for modeling to work. You have to be paying attention to the model (a distracted kid learns nothing from the demonstration). You have to retain what you saw — hold it in memory long enough to use it. You have to be physically capable of reproduction — a toddler can watch a slam dunk all day. And you have to have the motivation to actually do it. That last step is where consequences sneak back in, just secondhand. In a follow-up to the Bobo doll study, Bandura (1965) showed children a film of an adult attacking the doll and then varied what happened to the adult: rewarded, punished, or nothing. Children who saw the model punished imitated far less — until they were offered a prize for copying, at which point they reproduced the aggression just as well as everyone else. They had learned the behavior either way; what the punishment changed was whether they performed it. Seeing a model rewarded is vicarious reinforcement; seeing a model punished is vicarious punishment. The model doesn't even need to be in the room: Bandura distinguished live models (your coach demonstrating a kick), verbal models (your coach describing it), and symbolic models (a character doing it on a screen) — which is why the debate over violent media is, at bottom, a debate about observational learning.

Check yourself
According to Maier and Seligman's (2016) reversal of the classic learned-helplessness story, what does an organism actually learn?

Check yourself

The one thing to carry out of this unit

Learning is not one thing — it's at least three different mechanisms (associating stimuli, associating behaviors with consequences, and watching others) layered on top of a brain that was never a blank slate to begin with. Evolution constrains what associates easily with what; consequences shape behavior in ways that hinge on whether something was added or removed, not on whether it "feels negative"; and even the most famous, decades-old findings — Little Albert, learned helplessness — keep getting re-examined and, sometimes, rewritten. The mechanisms are real and replicable. The stories textbooks tell about them are a work in progress.

References

Bandura, A. (1965). Influence of models' reinforcement contingencies on the acquisition of imitative responses. Journal of Personality and Social Psychology, 1(6), 589–595. https://doi.org/10.1037/h0022070

Bandura, A. (1977). Social learning theory. Prentice Hall.

Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.

Breland, K., & Breland, M. (1961). The misbehavior of organisms. American Psychologist, 16(11), 681–684. https://doi.org/10.1037/h0040090

Fridlund, A. J., Beck, H. P., Goldie, W. D., & Irons, G. (2012). Little Albert: A neurologically impaired child. History of Psychology, 15(4), 302–327.

Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123–124.

Maier, S. F., & Seligman, M. E. P. (2016). Learned helplessness at fifty: Insights from neuroscience. Psychological Review, 123(4), 349–367.

Pavlov, I. P. (1927). Conditioned reflexes: An investigation of the physiological activity of the cerebral cortex (G. V. Anrep, Trans.). Oxford University Press.

Rescorla, R. A. (1988). Pavlovian conditioning: It's not what you think it is. American Psychologist, 43(3), 151–160. https://doi.org/10.1037/0003-066X.43.3.151

Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.

Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208.

Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. Journal of Experimental Psychology, 3(1), 1–14.