← All lessons
Lesson 11 · Experiments

The Logic of Experiments

This is the payoff of the whole course. The experiment is the only design that earns the word cause — and it does so with a deceptively simple recipe: manipulate one thing, measure another, and let a coin flip decide who gets what.

The problemWhy "cause" is such a hard word to earn

Here is a claim: writing about a painful memory with self-compassion causes people to feel less distressed. To test it, you might be tempted to compare people who naturally treat themselves kindly with people who don't. You already know from Lessons 9 and 10 why that fails — third variables everywhere, direction unknown. But it's worth seeing exactly why no amount of measuring can fix it, because that's what makes the experiment such a strange and beautiful invention.

What you really want to know is a counterfactual: how would this same person, at this same moment have felt if they had written self-compassionately instead of the way they actually wrote? Statisticians call the impossibility of ever observing that the fundamental problem of causal inference (Holland, 1986). Each person can be in only one condition at a time. The other outcome — the one that would tell you the causal effect directly — is forever missing. Every causal design in science is a workaround for this one missing data point.

The experiment's workaround is the cleverest one we have (Rubin, 1974). If you can't observe the same person under both conditions, create two groups that are interchangeable — identical, on average, in every respect: personality, history, mood, genes, everything, including all the things you never thought to measure. Then treat the groups differently in exactly one way. Whatever difference appears at the end can have only one source, because you built the groups to differ in only one way. The control group stands in for the treatment group's missing counterfactual.

And how do you build two groups identical on every variable, including the ones nobody has discovered yet? You can't do it by matching — you can only match on what you can name. The answer, when it finally arrived, was almost insulting in its simplicity: flip a coin.

The historyFisher, a cup of tea, and the coin flip that changed science

Random assignment feels obvious now, but it's young — younger than psychology itself. Working at an agricultural research station in the 1920s and 30s, Ronald Fisher faced the problem in the dirt: you can't give the same plot of land two fertilizers at once, and no two plots are truly identical. His solution, laid out in The Design of Experiments (Fisher, 1935), was to assign treatments to plots by a formal chance procedure. Randomization doesn't make any two groups identical — it makes them identical in expectation, and (this was Fisher's deeper insight) it makes the remaining chance differences quantifiable, so you can compute how likely your result would be if the treatment did nothing. Fisher opened the book with a now-famous story: a colleague claimed she could taste whether milk was poured before or after the tea. Rather than scoff, he designed a randomized tasting to find out. The lesson stuck: any claim, however odd, becomes testable the moment you control who gets what.

Psychology imported the idea and then systematized it. Campbell and Stanley (1963) catalogued the designs and named the threats that stalk them; Shadish, Cook, and Campbell (2002) built that into the modern framework of the four validities you've been using all semester. This lesson is where the framework pays off: the experiment is the design built, from the ground up, to win the internal-validity battle.

The recipeThree ingredients, one causal claim

A true experiment is not defined by lab coats or fancy equipment. It's defined by three moves, and it needs all three:

(1) Manipulate an independent variable. The researcher, not the world, sets who gets which level. The independent variable (IV) is the manipulated variable; its levels are the conditions (self-compassion prompts vs. control prompts; drug vs. placebo). If nobody manipulated anything — if people arrived already differing — it isn't an IV and it isn't an experiment. (2) Measure a dependent variable. The dependent variable (DV) is the outcome you hope "depends" on the IV, measured the same way for everyone. (3) Randomly assign participants to the IV's levels. Alongside the manipulated IV, good experiments also hold other variables constant — control variables: same room, same script, same time limits. Control variables aren't measured outcomes; they're the "everything else equal" that makes the comparison clean.

That third move is the quiet hero. It's what lets an experiment satisfy all three causal criteria at once, which no correlational design can do:

Causal criterionHow the experiment delivers it
CovarianceThe groups differ on the DV → the IV and outcome move together.
Temporal precedenceYou manipulate the IV first, then measure the DV. Order is guaranteed by design.
Internal validityRandom assignment scatters every third variable — measured or not — evenly across groups.

A classic from the prejudice literature shows the recipe turning a correlational observation into a causal claim. Word, Zanna, and Cooper (1974) first observed that white interviewers sat farther away, made more speech errors, and ended interviews sooner with Black applicants than with white ones. Observation alone couldn't say whether that cold treatment caused worse performance — so they ran the experiment. White applicants were randomly assigned to be interviewed by trained interviewers reproducing either the warm style or the cold, distant style. Applicants who received the cold style were rated as performing worse and appeared more nervous. Same people-pool, one manipulated difference, coin-flip assignment: the interviewer's behavior, not the applicant, was doing causal work. That is what discrimination research looks like when it earns causal language.

Myth · random assignment and random sampling are the same thing

This is the biggest mix-up in the entire course, and it's easy to breeze past. They are different words for different jobs. Random assignment takes the people you already have and splits them into conditions by chance — it buys you internal validity, the right to say the IV caused the effect. Random sampling pulls people from a population by chance — it buys you external validity, the right to generalize to that population. A study can have one without the other: a lab experiment on 40 psych majors can have flawless random assignment (so it's causally airtight) and terrible sampling (so it may not generalize). Almost every experiment in social psychology is exactly that combination, and it's a reasoned trade, not an accident: experiments prioritize internal validity and let replication across samples handle generalization. Assignment = cause. Sampling = generalize. Memorize the pairing.

Compared to what?Control groups, placebos, and comparison conditions

"Manipulate the IV" implies at least two levels, because a causal claim is always a comparison. The simplest baseline is a control group — a level that receives no treatment, or a neutral version of the task. But "no treatment" is often the wrong comparison. If the treatment group swallows a pill and the control group swallows nothing, the groups differ in two ways: the drug, and the experience of being treated. Expecting improvement can itself produce improvement, so drug trials use a placebo group — identical pill, no active ingredient — so that expectation is held constant and only the chemistry differs. (Lesson 12 takes placebo effects, and the double-blind designs that tame them, much further.)

Behavioral research has the same problem in different clothes. If people writing self-compassionate letters feel better than people who wrote nothing, is it the self-compassion — or just the act of writing about the event at all? The fix is a comparison condition that does everything the treatment does except the active ingredient: everyone writes, but only one group writes self-compassionately. Choosing the comparison condition is where experiments are won or lost, because the comparison defines exactly which causal claim you're entitled to. "Better than nothing" and "better than an equally engaging alternative" are very different findings.

Confound vs noiseThe difference between fatal and merely annoying

Students hear "variability" and panic. But not every stray variable sinks a study. The question is always the same: does it vary systematically with the IV?

Confound = systematic = fatal

A design confound is a second variable that tracks the IV — an accidental second difference between conditions. If every self-compassion session were run by a warm, encouraging experimenter and every control session by a brusque one, "experimenter warmth" moves in lockstep with condition. Now you can't tell whether the writing prompts or the warmth drove the effect. That's a fatal alternative explanation.

Unsystematic variability = noise

If warm and brusque experimenters were sprinkled randomly across both conditions, their style is just noise — haphazard scatter that doesn't favor either group. It makes a real effect harder to detect (more on power in Lesson 12) but can't manufacture a false one. Random assignment is exactly what turns would-be confounds into mere noise.

There's a third failure mode worth its own name. A selection effect is what happens when the people differ across conditions before the study even starts — usually because they chose their own condition or were assigned by convenience. Compare volunteers who signed up for an eight-week meditation course with people who didn't, and the groups differ in motivation, free time, openness, and a hundred unnamed things before minute one. A design confound is a flaw in what you did to the groups; a selection effect is a flaw in who ended up in the groups. Random assignment is the cure for selection effects — which is why "participants chose their condition" is the fastest way to spot a study that only looks like an experiment. Matched-group designs (match participants on a key variable, then randomly assign within pairs) can sharpen small samples, but note the fine print: matching supplements the coin flip, it never replaces it.

Go deeper · manipulation checks — did your IV actually work?

Here's an easily-missed step: how do you know your manipulation did what you intended? You add a manipulation check — a separate measure whose only job is to confirm the IV landed. If your prompts were supposed to induce a self-compassionate frame of mind, ask participants afterward how kindly they were treating themselves; the induction group should score higher. If it doesn't, your manipulation failed — and any null result on your real DV is uninterpretable: you can't distinguish "self-compassion doesn't help" from "we never actually induced self-compassion." Manipulation checks protect the construct validity of the IV, and they're the difference between "my IV didn't matter" and "my IV never happened."

The four designsBetween vs within, and their trade-offs

Two designs split the sample into separate groups (between-groups). In a posttest-only design, you randomly assign and measure the DV once, after the manipulation — the minimalist true experiment, and thanks to random assignment it's fully sufficient for a causal claim. In a pretest/posttest design, you measure the DV both before and after: the pretest lets you verify the groups really did start equivalent and lets you track change in each person, at the cost of possibly tipping participants off to what you're studying.

Two designs give every participant every level (within-groups). In a repeated-measures design, each person experiences the levels in sequence and is measured after each. Plassmann, O'Doherty, Shiv, and Rangel (2008) had every participant taste the same wine labeled with different prices; the identical wine tasted better — and produced more activity in the brain's pleasantness-coding regions — when labeled $90 than $10. Because each person is compared with themselves, individual differences in wine snobbery cancel out entirely. In a concurrent-measures design, all levels are presented at roughly the same time and a single preference is the DV. It's the workhorse of infant research, where babies can't fill out questionnaires: Fantz (1963) showed newborns a patterned and a plain gray disk simultaneously and recorded where they looked. They looked reliably longer at pattern — evidence, from a preference alone, that newborns perceive form.

Update · within-groups is a genuine mixed bag

Within-groups designs are seductive: each person is their own control, so the "groups" are perfectly equivalent, you get more statistical power, and you need fewer participants. But they carry threats between-groups designs don't (Greenwald, 1976). Order effects: experiencing one condition changes how you respond to the next — practice and fatigue effects, and carryover effects (a $90 expectation may linger into the next glass; orange juice right after toothpaste is not orange juice). And seeing every level can reveal the hypothesis — a demand characteristic: a participant who tastes "the same" wine at two prices and notices may perform the expected preference rather than feel it. Greenwald's advice still stands: when exposure to one condition genuinely changes a person — attitude change, learned skills, induced moods that linger — go between-groups. The power comes with strings.

Go deeper · counterbalancing, and when a Latin square rescues you

Order effects have a fix: counterbalancing — present the conditions in different sequences, so order effects spread evenly across conditions instead of piling onto one. With few conditions, use full counterbalancing (every possible order appears, participants randomly assigned to orders — note the coin flip sneaking back in). But orders explode factorially: 3 conditions = 6 orders, 4 = 24, 5 = 120. Enter the Latin square: a partial-counterbalancing scheme that guarantees every condition appears in every serial position exactly once, using only as many orders as there are conditions — 5 orders instead of 120. Reach for it whenever full counterbalancing would need more sequences than you have participants. (The name is Fisher-era agriculture again: the squares were literal grids of crop plots.)

Worked exampleInterrogating a real experiment end to end

Time to run the full checklist on a real study — one from the self-compassion literature this course keeps returning to. Leary, Tate, Adams, Allen, and Hancock (2007, Study 5) asked whether a brief self-compassion exercise causes people to feel better about their worst moments. Undergraduates recalled a real event involving failure, rejection, or embarrassment, then were randomly assigned to writing conditions: one group answered prompts inducing self-compassion (write kindly to yourself; note that others go through such things; describe your feelings with detachment), another answered prompts bolstering self-esteem (list your positive qualities; explain why the event doesn't reflect badly on you), and control participants either wrote about the event with no slant or did no writing at all. Then everyone reported their negative affect about the event. The self-compassion group felt less negative emotion than the self-esteem and control groups — and, strikingly, took more personal responsibility for the event while feeling better about it. Kindness, not flattery, did the work.

Now interrogate it. Covariance? Yes — conditions differed on the DV. Temporal precedence? Guaranteed by design: the writing manipulation happened before affect was measured. Internal validity? Random assignment made the groups equivalent, on average, in trait self-compassion, event severity, and everything else; the writing control condition rules out "any writing about the event helps"; and the self-esteem condition — the crucial comparison — rules out "any positive reframing helps." All three criteria met: the causal claim is earned. Then push on the other validities. Construct validity: did the prompts really induce self-compassion rather than mere distraction or mood? The prompts map onto the construct's defined components, but a direct manipulation check would strengthen the case. External validity: undergraduates recalling past events in a lab — does the effect hold for fresh wounds, older adults, other cultures? The design doesn't say; replication must. That's not a flaw; it's the standard shape of a true experiment: airtight on cause, agnostic on reach.

Interrogate any experiment · the six questions to ask every time

1. What was manipulated (IV and its levels), and what was measured (DV)? 2. Was assignment to levels truly random — or did people select their own condition (selection effect)? 3. Compared to what? Is the comparison condition doing everything the treatment does except the active ingredient? 4. Does anything else vary systematically with the IV (design confound), or is the stray variation just noise? 5. Is there evidence the manipulation actually worked (manipulation check)? 6. If within-groups: were order effects counterbalanced, and could participants have guessed the hypothesis? Answer all six and you've done a professional review.


Identify the experimental design. Tap each study to reveal its design and the tell.

Pick a study to see which of the four basic designs it uses.

Match each design to its one-line description

Click a design, then click the description that fits. Six to clear.

From manipulation to a justified causal claim

Arrange the logic chain that lets an experiment earn the word cause.

🂠

Confound, or just noise?

Decide whether the stray variable is a fatal confound or harmless noise, then flip.

Every self-compassion session was run by a warm experimenter; every control session by a brusque one.
Tap to flip
Confound (fatal)

Experimenter warmth tracks the IV systematically. You can't tell whether the prompts or the warmth drove the mood difference. Classic design confound.

Warm and brusque experimenters were spread randomly across both writing conditions.
Tap to flip
Just noise

Experimenter style doesn't favor either condition — it's unsystematic variability. It adds scatter but creates no false effect. Random assignment did its job.

The treatment group was tested in a quiet morning room; the control group in a noisy afternoon room.
Tap to flip
Confound (fatal)

Time of day and noise vary systematically with condition. Any difference could be the treatment OR the setting — internal validity is wrecked.

Participants were allowed to pick whichever writing condition sounded more appealing to them.
Tap to flip
Selection effect (fatal)

The people now differ across conditions before the study starts — those drawn to self-compassion writing may already be kinder to themselves. Not random assignment, not an experiment.

Check yourself — Logic of Experiments quiz

Eight questions. Instant feedback with the reasoning.


SourcesCited in APA 7

Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.
Fantz, R. L. (1963). Pattern vision in newborn infants. Science, 140(3564), 296–297. https://doi.org/10.1126/science.140.3564.296
Fisher, R. A. (1935). The design of experiments. Oliver & Boyd.
Greenwald, A. G. (1976). Within-subjects designs: To use or not to use? Psychological Bulletin, 83(2), 314–320. https://doi.org/10.1037/0033-2909.83.2.314
Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945–960. https://doi.org/10.1080/01621459.1986.10478354
Leary, M. R., Tate, E. B., Adams, C. E., Allen, A. B., & Hancock, J. (2007). Self-compassion and reactions to unpleasant self-relevant events: The implications of treating oneself kindly. Journal of Personality and Social Psychology, 92(5), 887–904. https://doi.org/10.1037/0022-3514.92.5.887
Plassmann, H., O'Doherty, J., Shiv, B., & Rangel, A. (2008). Marketing actions can modulate neural representations of experienced pleasantness. Proceedings of the National Academy of Sciences, 105(3), 1050–1054. https://doi.org/10.1073/pnas.0706929105
Rubin, D. B. (1974). Estimating causal effects of treatments in randomized and nonrandomized studies. Journal of Educational Psychology, 66(5), 688–701. https://doi.org/10.1037/h0037350
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Word, C. O., Zanna, M. P., & Cooper, J. (1974). The nonverbal mediation of self-fulfilling prophecies in interracial interaction. Journal of Experimental Social Psychology, 10(2), 109–120. https://doi.org/10.1016/0022-1031(74)90059-6

← Prev: Multivariate Correlational Research Next: Experimental Control →