← All lessons
Lesson 12 · Experiments

Experimental Control

Lesson 11 built the experiment. This lesson stress-tests it. An experiment can lie to you in two directions: a bogus effect can sneak in and look real, or a real effect can hide and look like nothing. Control is the craft of closing both doors — and there is a named, catalogued threat behind almost every way a study goes wrong.

The catalogueA field guide to being fooled

In 1963, Donald Campbell and Julian Stanley published a slim monograph — Experimental and Quasi-Experimental Designs for Research — that became one of the most cited works in the history of social science. Its central move was almost bureaucratic in its genius: instead of vaguely warning researchers to "be careful," it named and numbered the specific ways a study can generate a phony effect. Eight threats to internal validity, each with a definition, a signature, and a design that defeats it (Campbell & Stanley, 1963). Forty years later, Shadish, Cook, and Campbell (2002) refined and extended the list, and it remains the field's shared vocabulary. When a reviewer writes "this looks like regression to the mean," they are quoting this catalogue.

Every threat below answers the same question: could the groups (or the same group over time) have differed for some reason other than the independent variable? If yes, internal validity is compromised — the causal claim is on the table. The catalogue is easiest to see in the design where nothing protects you: the one-group pretest–posttest design. Measure a group, treat them, measure again. If scores improve, the treatment worked… or one of six imposters did. Campbell and Stanley filed this under "pre-experimental designs," which is the methodological equivalent of a restaurant hygiene grade of F.

ThreatHow it fakes an effectDesign fix
MaturationPeople change spontaneously over time — no treatment needed.A randomly assigned comparison group that matures too.
HistoryAn outside event hits everyone between pretest and posttest.A comparison group that lives through the same events.
Regression to the meanA group selected for extreme scores drifts back toward average on retest.A comparison group selected the same extreme way.
AttritionSystematic dropout reshapes the group between measurements.Track who leaves; compare dropout across conditions.
TestingTaking the measure once changes the second score (practice, sensitization).A comparison group tested on the same schedule; posttest-only designs.
InstrumentationThe measuring instrument itself changes from Time 1 to Time 2.Calibrate; keep coders blind; use identical, stable measures.
SelectionGroups differed from the start.Random assignment (Lesson 11's whole point).
Selection interactionsOne group matures faster, or experiences its own local history.Random assignment again; matched pretests in quasi-experiments.

Take the first two slowly, because they anchor everything else. Maturation is spontaneous change: people heal, settle, grow, fatigue, and adapt on their own schedule. Suppose a college runs a self-compassion journaling program for homesick first-year students each September and finds they feel dramatically better by December. Compelling — except homesickness declines across the first semester for nearly everyone, journal or no journal. The program gets credit for what the calendar did. History is different: not the passage of time, but a specific outside event that lands on the whole group mid-study. Imagine an eight-week meat-reduction intervention running exactly when a national food-safety scare dominates the news. Meat consumption drops — in your program and everywhere else. In both cases the fix is identical and beautiful: a randomly assigned comparison group ages through the same weeks and watches the same news. Whatever time and history do, they do to both groups, so any difference between groups survives the threat.

The most misunderstood threatRegression to the mean, properly

Regression to the mean is the threat students think they understand and usually don't, so let's do it right. The phenomenon is old: Francis Galton (1886) measured parents and their adult children and found that very tall parents had children who were tall, but less tall — closer to average. Not because nature punishes tall families, but because of how extreme scores are built.

Here is the machinery. Any observed score is a mix of a stable true level and transient luck: observed score = true level + chance. To land at the extreme of a distribution, you usually need both — a genuinely high (or low) true level and chance breaking your way that day. When you remeasure, true level sticks around but chance rerolls. The good luck that helped push you to the extreme doesn't repeat, so the group's average drifts back toward the mean. No treatment, no force, no "law of averages" — just selection on scores that were partly luck the first time.

Two conditions must both hold, and this is what people miss. First, the group must be selected for extreme scores — the worst sleepers, the most anxious, the highest scorers. Second, the measure must be imperfectly correlated with itself over time, which every psychological measure is. If either condition fails, no regression. A random sample doesn't regress; a perfectly reliable measure doesn't regress. But recruit the fifty most stressed students on campus for your mindfulness app, and their retest average will improve even if the app is a blank screen.

Go deeper · why regression makes punishment look effective

Regression to the mean doesn't just fool researchers — it warps everyday judgment, because life keeps "selecting on extremes" for us. A coach chews out an athlete after her worst race; next race she's better. He praises her after her best race; next race she's worse. The coach concludes criticism works and praise backfires. But her worst and best races were partly luck, and luck rerolled. Extreme performances are followed by more average ones regardless of what anyone says in between — which means regression systematically manufactures evidence that punishment works and reward fails. Any intervention aimed at people because they just scored terribly — remedial programs, crackdowns after crime spikes, therapy sought at rock bottom — will look effective for free. Only a control group selected the same extreme way can tell you whether it actually is.

People leave, practice, and driftAttrition, testing, instrumentation — and selection's accomplices

Attrition (Campbell and Stanley's grimly named "experimental mortality") is loss of participants between measurements. The danger isn't dropout itself — it's who drops out. If a demanding exercise intervention loses its least-fit participants by posttest, the survivors' average fitness rises even if the program did nothing: you're comparing everyone at pretest with the hardy remnant at posttest. The fix is bookkeeping plus honesty: report how many left, from which conditions, and whether leavers differed at pretest from stayers.

Myth · attrition always breaks internal validity

It depends on the pattern. If people drop out at similar rates for similar reasons across all conditions, the sample gets less representative — an external-validity problem — but the groups stay comparable. The killer is differential attrition: dropout that differs between conditions. If the treatment is unpleasant and drives away exactly the people it isn't helping, the survivors in the treatment group are no longer comparable to the control group — random assignment has been quietly undone after the fact. Any final difference could be the treatment or the selective exodus. This is why clinical trials obsess over retention and report a participant flow diagram before a single result.

Testing effects arise because measurement is itself an experience. Take an ability test twice and you improve just from practice; answer an attitude scale twice and the first pass may have started you thinking. Reaction-time research — including the live reaction-time studies this site hosts — sees this within a single session: people speed up noticeably over their first few dozen trials as the task becomes familiar. That's why every well-built RT experiment opens with a practice block whose data are thrown away: burn off the practice effect before the trials that count. Instrumentation is the evil twin: the measuring device changes between measurements. Human coders drift — growing lenient, or sharpening their private definition of "aggressive behavior" halfway through a study. Hardware and software drift too: an untested browser update that shifts stimulus timing mid-study changes the instrument under your feet, which is why RT researchers pin software versions and verify timing before launch.

Go deeper · testing vs. instrumentation — the pair everyone swaps

These two are constantly confused because the same numbers move. The distinction is crisp. In a testing threat, the participants changed — practice, familiarity, sensitization. In an instrumentation threat, the instrument changed — coders drifted, apparatus lost calibration, or you "improved" the question wording between waves so the versions aren't equivalent. Rule of thumb: testing = the people changed; instrumentation = the ruler changed. Different culprit, different fix: testing yields to comparison groups and posttest-only designs; instrumentation yields to calibration, blind coders, and leaving the ruler alone.

Finally, selection — groups that differed before anyone did anything — and its sneakier interactions. Random assignment kills plain selection, which is why Lesson 11 made such a fuss about it. But when assignment isn't random (Lesson 14's territory), selection teams up with the other threats: a selection–maturation interaction means one group was already changing faster (compare an honors section to a night class and the honors students improve more with or without your teaching method); a selection–history interaction means one group experienced its own local event (one campus, one dorm, one workplace gets hit by news the other never hears). Shadish et al. (2002) treat these interactions as the central headache of quasi-experimentation — keep them in your pocket for Lesson 14.

The humans in the roomDemand characteristics and experimenter expectancy

The catalogue so far blames time, luck, and rulers. The next two threats blame the people in the room — including you. Martin Orne (1962) argued that a psychology experiment is not a neutral measurement chamber but a social situation with its own etiquette. Participants arrive wanting to be good participants; they scan everything — the recruitment flyer, the lab décor, the phrasing of instructions — for clues about what the study "really" wants, and then, helpfully, deliver it. Orne called these clues demand characteristics. To show how far the "good subject" role goes, he tried to invent a task so tedious that people would refuse: adding columns of random digits, sheet after sheet, then tearing up each completed sheet and starting the next. Participants kept working for hours. If people will shred their own arithmetic to please an experimenter, they will certainly nudge a mood rating.

The experimenter's side of the handshake is just as leaky. Robert Rosenthal and Kermit Fode (1963) told student experimenters that their lab rats were specially bred to be "maze-bright" or "maze-dull." The rats were ordinary siblings, randomly labeled — yet the "bright" rats learned faster. Nothing mystical happened: handlers with high expectations treated their animals differently in dozens of tiny ways, and the data bent accordingly. This is experimenter expectancy (its scoring-room cousin is observer bias: expectations coloring how behavior gets coded). If expectations can tilt a rat study, imagine what they do when the measure is a human judging another human's "warmth."

Update · blinding is a technology, not a vibe

The engineering solution is to remove knowledge from whoever could leak it. In a double-blind design, neither participants nor the researchers who interact with them or score their behavior know who is in which condition — participants can't act out the hypothesis they can't detect, and experimenters can't nudge or score toward an expectation they don't hold. When full double-blinding is impossible (you can't hide a psychotherapy condition from the person receiving it), a masked design blinds whoever can be blinded — typically the coders and outcome assessors — which still shuts down observer bias. Modern practice adds Orne's own tools: post-experimental funnel interviews that ask participants what they thought the study was about, and pilot "simulator" runs that check whether the cover story holds. A hypothesis your participants can recite back to you is a hypothesis you have already contaminated.

Belief as an active ingredientPlacebo effects, done seriously

In 1955, anesthesiologist Henry Beecher published "The Powerful Placebo," reviewing 15 clinical studies and concluding that roughly a third of patients improved satisfactorily on placebo alone (Beecher, 1955). The paper helped institutionalize the placebo-controlled trial — one of the great methodological reforms of the century — and planted a number (35%) that got quoted for decades.

Here's the twist, and notice that you now own the tools to see it: most of the improvement Beecher attributed to placebo was never placebo at all. His studies had no no-treatment group, so "improvement on placebo" silently bundled together the placebo response plus maturation (illnesses run their natural course) plus regression to the mean (patients enroll in trials when they feel worst — an extreme-score selection if ever there was one). When Hróbjartsson and Gøtzsche (2001) analyzed over a hundred trials that included both a placebo arm and a no-treatment arm — the comparison that isolates the placebo effect itself — the powerful placebo shrank dramatically: little to no effect on most objective, binary outcomes, with modest but real effects on continuous subjective outcomes, especially pain. The famous 35% was mostly the threat catalogue wearing a lab coat.

That doesn't make placebo effects fake — it makes them a specific, measurable psychological phenomenon rather than magic. Modern placebo science treats expectation and conditioning as active ingredients with real physiological signatures, strongest for pain and other subjectively reported states. It has even produced a genuinely strange finding: open-label placebos can work. Kaptchuk et al. (2010) randomized irritable bowel syndrome patients to no treatment or to pills honestly described as inert placebos — and the openly-labeled placebo group reported meaningfully greater symptom relief. For the methodologist, the lesson is permanent: whenever a "treatment" involves belief, attention, or ritual — which in psychology is essentially always — your control group must receive the belief, attention, and ritual without the active ingredient. Otherwise you are comparing drug-plus-hope to nothing, and hope is doing unmeasured work.

When you find nothingWhat a null result can't tell you

Now flip the problem. Your groups didn't differ. It's tempting to read a null result as "the effect doesn't exist," but a flat result is ambiguous in a way a positive result isn't: maybe there's truly no effect — or maybe there is one and your study couldn't see it. Before a null result means anything, you have to rule out the ways an experiment goes blind. They come in two families: not enough true difference between groups, and too much noise within them.

Not enough between-groups difference

Weak manipulation: the IV levels were too similar to move anything — a 90-second self-compassion writing prompt versus neutral writing may simply be too small a dose to shift how people treat themselves. Insensitive measure: a DV too coarse to register real change, like classifying diet as "ate meat this month: yes/no" in a meat-reduction study — cutting from fourteen servings a week to two reads as no change. Ceiling and floor effects: everyone piles up at the top or bottom of the scale, leaving groups no room to separate.

Too much within-groups variability

Noise — error variance from individual differences, sloppy procedure, distracted participants, and above all measurement error — inflates the spread within each group until a real between-group signal drowns in it. Fixes: reliable, precise instruments (Lesson 06 pays off here); standardized procedures; attention checks to catch random responders; more participants so random errors average out; and within-groups designs, which remove stable individual differences from the error term entirely.

Two humble tools guard against most of this, and both belong in every study you ever run. A manipulation check is an extra measure that verifies the IV actually took hold: if your "self-compassion condition" doesn't score higher on a brief state self-compassion measure than the control condition, the experiment ended before the DV was ever in play — and, crucially, its null result is uninterpretable, not evidence of no effect. And pilot testing — running the full procedure on a small sample before the real study — catches weak manipulations, confusing instructions, ceilings, floors, and leaked demand characteristics while they're still cheap to fix. Pilots feel like a delay. They are the fastest thing in this lesson.

Interrogate · a study found nothing — ask in this order

1. Did the manipulation check confirm the IV took hold? 2. Was the manipulation strong enough to plausibly move the DV? 3. Was the measure sensitive — any sign of a ceiling or floor? 4. How noisy were the groups — reliable measures, standardized procedure, attention checks? 5. Did the study have the statistical power to detect a plausibly sized effect — was a power analysis reported? 6. Only if all five check out: treat the null as informative evidence against the effect. A null result earns its meaning; it doesn't get it for free.

The modern lensStatistical power, and the crisis it predicted

Statistical power is the probability that your study detects an effect given that the effect is real. You met the statistic in your statistics course; here it becomes a design tool. Power rises with larger samples, larger true effects (which is what strong manipulations buy you), less noise (which is what reliable measures buy you), and within-groups designs. Jacob Cohen spent a career pleading with psychologists to take it seriously — and being ignored. In 1962 he audited every article in a year of the Journal of Abnormal and Social Psychology and found the average study had less than a coin-flip's chance of detecting a medium-sized effect (Cohen, 1962). Thirty years later he published "A Power Primer," a two-page cheat sheet that made power analysis impossible to excuse away (Cohen, 1992). The field mostly kept not doing it.

The bill arrived. Button et al. (2013) — a paper pointedly titled "Power Failure" — estimated the median power across neuroscience literatures at around 21%, and spelled out the part that isn't intuitive: underpowered studies fail in both directions at once. They miss real effects, obviously. But the "wins" they do publish are systematically inflated, because in a small noisy study only the luckily oversized estimates clear the significance bar — a phenomenon called the winner's curse. A literature built from underpowered studies is therefore a literature of exaggerated flukes: real effects hiding among nulls, and published effects too big to be true. When the Open Science Collaboration (2015) re-ran 100 published psychology studies, fewer than half replicated, and the replication effects averaged about half the original sizes — exactly the signature the power arithmetic predicts. That story gets a full lesson (Lesson 15); for now, understand that power is where this lesson's two halves meet: it is simultaneously your defense against meaningless nulls and the field's defense against meaningless wins.

Update · a-priori power analysis is now the entry fee

The professional standard is to choose your sample size before collecting data. An a-priori power analysis asks: given the smallest effect I'd care about and the power I want (conventionally .80 or higher), how many participants do I need? Free tools like G*Power answer in seconds, and Cohen (1992) supplies the effect-size conventions (small, medium, large) that make the question tractable. Preregistration platforms and top journals now expect the calculation up front — a sample size justified before the data exist can't be quietly grown until something turns significant. When you read any modern study, ask the power question early: was this study even equipped to find what it went looking for? If not, its null tells you almost nothing — and its "significant" finding may tell you less than you think.

Control in the wildAnatomy of a reaction-time experiment

Everything in this lesson is visible in one modern package: the browser-based reaction-time experiment, the workhorse of implicit social cognition and a genre this site hosts live. Strip one down and you'll find that nearly every component is a named countermeasure to a named threat. The consent page and standardized instructions give every participant the identical framing — demand characteristics held constant. The practice block absorbs testing effects before the real trials begin. The task presents conditions in counterbalanced blocks — half of participants get block order A→B, half get B→A — so practice and fatigue push equally on both conditions instead of masquerading as an effect of one. Timing is calibrated and software versions pinned, guarding instrumentation. Attention and validity checks flag random responders, cutting within-group noise. And the sample size comes from an a-priori power analysis filed in the preregistration. None of this is decoration; it's the threat catalogue, implemented.

Design feature you can seeThreat it controls
Standardized instructions & cover storyDemand characteristics
Practice block (data discarded)Testing / practice effects
Counterbalanced block orderOrder effects (practice, fatigue)
Calibrated timing, pinned softwareInstrumentation
Attention & validity check trialsWithin-group noise (inattentive responding)
Automated scoring, no human coderObserver bias
Preregistered a-priori sample sizeLow power

Which internal-validity threat is this? Each one-group study found an "effect." Tap to reveal the imposter.

Pick a study to see which threat is faking the effect.

Match each threat to its definition

Click a threat, then click the definition it matches. Six to clear.

Diagnosing a null finding, in order

You got no effect. Arrange the questions a careful researcher asks, from first to last.

🂠

Why might this study have found nothing?

Name the obscuring factor hiding a real effect, then flip.

A word-categorization task is meant to compare accuracy across conditions, but nearly everyone scores 98–100% correct in both.
Tap to flip
Ceiling effect

Scores are jammed against the top, so the conditions have no room to separate. Make the task harder — or switch the DV to reaction time, which has no ceiling at 100%.

A 90-second self-compassion writing prompt is compared to neutral writing; the manipulation check shows the groups don't differ in state self-compassion.
Tap to flip
Weak manipulation

The dose was too small to move the construct — and the manipulation check caught it before anyone misread the null. Strengthen the induction and pilot it before trusting any result downstream.

A meat-reduction study measures diet as 'ate meat this month: yes/no.' Most participants cut from 14 weekly servings to 2 — and register as unchanged.
Tap to flip
Insensitive measure

The DV is too coarse to register a real, large change. A servings-per-week count would have caught the effect the yes/no item is blind to.

A real but small effect is studied with 12 participants per condition and comes out non-significant.
Tap to flip
Not enough power

Twelve per cell can't reliably detect a small effect — and if it does hit significance, the estimate is likely inflated (the winner's curse). An a-priori power analysis would have flagged the sample size before data collection.

Check yourself — Experimental Control quiz

Eight questions. Instant feedback with the reasoning.


SourcesCited in APA 7

Beecher, H. K. (1955). The powerful placebo. Journal of the American Medical Association, 159(17), 1602–1606. https://doi.org/10.1001/jama.1955.02960340022006
Button, K. S., Ioannidis, J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., & Munafò, M. R. (2013). Power failure: Why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience, 14(5), 365–376. https://doi.org/10.1038/nrn3475
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.
Cohen, J. (1962). The statistical power of abnormal-social psychological research: A review. The Journal of Abnormal and Social Psychology, 65(3), 145–153. https://doi.org/10.1037/h0045186
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155–159. https://doi.org/10.1037/0033-2909.112.1.155
Galton, F. (1886). Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland, 15, 246–263. https://doi.org/10.2307/2841583
Hróbjartsson, A., & Gøtzsche, P. C. (2001). Is the placebo powerless? An analysis of clinical trials comparing placebo with no treatment. New England Journal of Medicine, 344(21), 1594–1602. https://doi.org/10.1056/NEJM200105243442106
Kaptchuk, T. J., Friedlander, E., Kelley, J. M., Sanchez, M. N., Kokkotou, E., Singer, J. P., Kowalczykowski, M., Miller, F. G., Kirsch, I., & Lembo, A. J. (2010). Placebos without deception: A randomized controlled trial in irritable bowel syndrome. PLoS ONE, 5(12), Article e15591. https://doi.org/10.1371/journal.pone.0015591
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Orne, M. T. (1962). On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications. American Psychologist, 17(11), 776–783. https://doi.org/10.1037/h0043424
Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183–189. https://doi.org/10.1002/bs.3830080302
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.

← 11 · The Logic of Experiments Next: Factorial Designs →