Here's the reframe that makes this whole topic click: science isn't a body of facts, it's a set of habits for not being fooled — especially by yourself. Every term here is a countermeasure. Random assignment fights confounds. Double-blind procedures fight your own expectations. Inferential statistics fight your eagerness to see a pattern in noise. Learn each method as an answer to the question "what mistake would I make if this rule didn't exist?" and the vocabulary stops being a list to memorize and becomes a toolkit you actually understand.
Bodhi says
Think of science as a game with rules that anyone can play — and the key word is verb. Science isn't who you are or where you work; it's how you go about answering a question. A retired schoolteacher with a horse can do bad science with total sincerity; a skeptic with a stopwatch can do good science in a hallway. The method, not the credential, is what separates them.
The four canons: the terms and conditions of playing
Before any method, there are four commitments you agree to when you "play the game." Here's what each one is quietly ruling out.
Determinism says events have systematic causes — behavior isn't random magic, so it's worth hunting for the cause. Empiricism says knowledge comes from observation, not from authority or how strongly you believe something. Parsimony (Occam's Razor) says when two explanations both fit, prefer the simpler one — not because simple is always true, but because the extra machinery of the complicated one needs to earn its keep. And testability — Karl Popper's falsifiability — is the sharpest of the four: a claim is scientific only if there's some observation that could prove it wrong (Popper, 1959).
That last one trips people up, so sit with it. "Psychics are real but skeptics' bad vibes block the powers" is unfalsifiable — every failure gets explained away, so no test can ever count against it. A theory that can't lose is not strong; it's empty. This is why we say hypotheses get confirmed or disconfirmed, never "proven." Science stays permanently open to being wrong, and that openness is the source of its strength, not a weakness.
The distinction the word "theory" ruins in casual English
In everyday talk, "theory" means hunch — "it's just a theory." In science it means nearly the opposite: a well-substantiated framework that explains facts. Stephen Jay Gould's line is the fix — evolution is both a theory and a fact. The fact is that life has changed over time (the fossil and genetic record — data points reasonable observers agree on). The theory is the explanatory model — natural selection, drift, mutation — that makes those facts mean something. Same with gravity: the fact is that things fall; the theory is why. Theories don't sit below facts in some ladder of certainty. They're the structures that make facts meaningful.
Clever Hans: the founding horror story of every methods course
Berlin, turn of the 20th century. A retired schoolteacher, Wilhelm von Osten, has a horse named Hans who can seemingly add, subtract, tell time, even do square roots — all by tapping his hoof. Crowds gasp. Newspapers swarm. Here is a horse that does math. And the crucial detail: von Osten was no con man. He sincerely believed it. That sincerity is what makes the story so useful — it shows how easily we're fooled by our own expectations, even with the purest motives.
Enter Oskar Pfungst, playing the skeptic — and note a distinction worth hammering home: a skeptic is not a cynic. A cynic dismisses things out of hand; a skeptic tests them carefully. Pfungst (1911) noticed Hans got answers right only when he could see the questioner and when the questioner knew the answer. When the human didn't know, neither did Hans. The "genius horse" was reading tiny, involuntary cues — a lean forward, a release of tension — that told him when to stop tapping. Nobody was cheating. It was an inadvertent communication loop.
That loop has a name that will haunt the rest of the course: the observer-expectancy effect. The observer's expectations leak out through micro-behavior and shape the very result they expected to find — a cousin of the confirmation bias, our habit of seeing what we expect to see rather than what's in front of us. Darley and Gross (1983) put a face on it: told the same child was from a high vs. low socioeconomic background, observers "saw" the child's academic ability differently — same tape, different perceptions, driven by prior belief. As Carl Sagan put it, extraordinary claims require extraordinary evidence — and, we'd add, controlled conditions, because sincerity is not evidence.
Bodhi says
Clever Hans is why the double-blind exists. If the experimenter doesn't know which condition a participant is in, they can't leak the answer through a lean or a smile. When you meet "double-blind" later, don't memorize a definition — remember a horse tapping his hoof, and a psychologist realizing the humans were the ones giving away the answer.
Description → prediction → causation: the ladder the whole topic climbs
Every method sits on a rung of one ladder, and the ladder only goes one direction — toward stronger claims that are harder to earn. Keep this shape in your head and the "which method is this?" questions become almost automatic.
The bottom rungs: watching and asking
The observational method answers "what is happening?" through naturalistic observation (watch without interfering), ethnography (immerse in a culture), or archival analysis (mine existing records). Its guardrail is inter-rater reliability — do independent observers agree? — which is really the Clever Hans lesson turned into a number: if only your eye sees the pattern, maybe the pattern is in your eye. Self-report methods (surveys, interviews) are efficient and give you access to the inside of a head, but they buy that access with the risk of dishonesty and demand characteristics — participants guessing what you want and helpfully supplying it. Description is where science starts. But it can't tell you what goes with what, and it certainly can't tell you what causes what — which is why we climb.
Correlation: the rung where everyone slips
The correlational method asks: from knowing X, can we predict Y? The correlation coefficient (r) runs from −1.0 to +1.0. The sign is direction — positive means they move together (more IQ, more grades), negative means they move opposite (colder out, more hot chocolate). The magnitude — how close to 1 — is strength. An r of −0.8 is a stronger relationship than an r of +0.3, even though it's negative; students routinely misread the minus sign as "weak." It isn't. It's just the other direction.
Now the famous warning: correlation is not causation. You've heard it. But the slogan is useless until you can say exactly why — and there are two distinct reasons, which the short version usually leaves abstract.
The two reasons correlation ≠ causation, made concrete
1. Directionality. Suppose stress and coffee consumption correlate. Does coffee cause stress (A→B)? Or does being stressed make you reach for more coffee (B→A)? The correlation is identical either way; it cannot tell the arrow's direction. 2. The third-variable problem. Maybe neither causes the other — a lurking C causes both. Ice-cream sales correlate with drowning deaths, but ice cream doesn't drown anyone; summer heat (C) drives both swimming and ice cream. That's a spurious correlation — Tyler Vigen's charts (per-capita cheese consumption vs. bedsheet-tangling deaths) are the comedy version, but real research is littered with the serious kind. So whenever you see a correlation, ask three questions in order: could A cause B, could B cause A, or could some C be causing both?
The top rung: the experiment, and what earns causation
Only the experimental method licenses causal claims, and it earns that license with three requirements: at least two groups, a manipulated independent variable (the IV — the cause you deliberately change), and random assignment. You measure the dependent variable (the DV — the effect, the thing that "depends"). A memory trick that survives exams: the IV is what I vary; the DV is the data I collect. The enemy is the confound — any uncontrolled variable that differs between groups alongside the IV, so you can't tell which one moved the DV.
Random assignment is the quiet hero here, and it's worth understanding why it works. When you flip a coin to sort people into groups, every third variable — motivation, sleep, IQ, mood, whatever you didn't think to measure — gets scattered roughly evenly across conditions. So the groups start out equivalent on everything except the IV you're about to manipulate. That's the magic: it neutralizes confounds you didn't even know existed. It's the experimental method's answer to the third-variable problem that sinks correlation.
How to tell them apart: random ASSIGNMENT ≠ random SELECTION
Students fuse these two constantly, and they do opposite jobs. Random selection (a sampling issue) is how you recruit — drawing participants randomly from a population so your sample represents it. It buys external validity: the right to generalize your finding beyond your sample. Random assignment (a design issue) is how you sort the people you already have into conditions. It buys internal validity: the right to say the IV, not a confound, moved the DV. You can have one without the other — a lab study on 200 psych majors can have flawless random assignment (strong causal claim) yet zero random selection (shaky generalization to anyone but sophomores). The 1936 Literary Digest poll — which confidently predicted the wrong presidential winner from a huge but biased mailing list — is a random-selection failure, not an assignment one.
Two design flavors round it out. Between-groups designs give different groups different conditions (Group A hears music, Group B hears silence). Within-groups designs give the same people every condition (each person does the task with music and in silence). Within-groups is thriftier and cancels out individual differences automatically — but it opens the door to order effects (fatigue, practice), which is its own confound to manage. And to fight the Clever-Hans problem in both, we use blind and double-blind procedures so neither participant nor experimenter can nudge the result.
Reliability vs. validity: the pair that's constantly confused
These sound like synonyms and are not. Reliability is consistency — you get the same reading each time. Validity is accuracy — you're measuring the thing you claim to measure. The catch that makes it a favorite trap: a measure can be perfectly reliable and completely invalid, but it cannot be valid without being reliable.
Reliable but NOT valid
A bathroom scale that always reads exactly 5 pounds heavy. Rock-steady, repeatable — and wrong every single time. Consistency without accuracy.
Valid requires reliable
A scale can't be accurate if it flings out a different number each time you step on it. So reliability is the price of admission for validity — necessary, but not sufficient.
Zoom out from measures to whole studies and the same word splits two ways. Internal validity asks: within this study, are we sure only the IV affected the DV (no confounds)? External validity asks: does the finding generalize beyond this lab, these people, this setting? These usually trade off — the tight control of a lab study (high internal validity) can make it feel artificial (lower external validity), while a messy field study is realistic but harder to control. This is the classic control-vs-realism tension between laboratory and field settings, and it's why researchers also chase psychological realism — making the lab situation feel real enough (often via a cover story) that people behave as they would in the wild.
Statistics: telling signal from noise, and the term everyone misreads
Two jobs here. Descriptive statistics summarize the data you have — central tendency (mean, median, mode) and variability (variance, standard deviation, i.e. how spread out the scores are). Two groups can share an identical mean and be wildly different in spread, which is exactly why the mean alone lies to you. Inferential statistics do the harder job: they estimate whether a result is likely real or just chance wobble in your sample.
Which brings us to the single most misunderstood phrase in the whole vocabulary: statistical significance, the famous p < .05. Read the next box slowly — this is the one that separates people who parrot the definition from people who get it.
What p < .05 actually means (and the three things students wrongly think it means)
The real meaning: if there were truly no effect, you'd see a result this big (or bigger) by pure chance less than 5% of the time. That's it. It's a statement about how surprising your data would be in a boring, effect-free world.
Now the misreadings. It does not mean there's a 95% chance your hypothesis is true. It does not mean the effect is large or important. And it does not mean the finding is guaranteed to replicate. A tiny, trivial effect can be "significant" if your sample is huge — significance answers "is it probably real?", not "is it big enough to care about?"
That last point deserves its own name: effect size. Significance (p) tells you a result probably isn't zero. Effect size tells you how much — and therefore whether it matters in the real world. A drug that lowers blood pressure by a statistically significant 0.1 mmHg is real and useless. When you read "significant" in a headline, the sharp question isn't "did they hit p < .05?" It's "significant, sure — but how big, and who cares?" Statistical significance and practical importance are different questions, and confusing them is how bad science journalism is born.
Where the science stands now: replication and preregistration
Everything above got a hard stress-test after 2011. When large teams tried to reproduce landmark findings, a troubling share didn't hold up — the "replication crisis." A big culprit was exactly the p < .05 misunderstanding above: chase a significant result hard enough (run extra conditions, drop inconvenient participants) and you can manufacture one from noise. The fixes are now standard practice: preregistration (declaring your hypotheses and analysis before collecting data, so you can't quietly rewrite the target around the result), bigger samples, and open data. This isn't a footnote to methods — it is methods, updated. The whole story lives on Science Corrects Itself, and it'll make you read every "new study finds…" headline with sharper eyes.
The one thing to carry out of this topic
Method is humility, operationalized. Every rule in this unit exists because a smart, sincere person — a von Osten, a headline writer, you at 2 a.m. seeing a pattern — was once fooled in a predictable way, and the rule is the scar tissue. Random assignment because we can't see confounds. Double-blind because our expectations leak. Effect size because "significant" flatters us. So the reflex to build isn't memorizing which method is which; it's the reflex to ask, of any claim you meet in the wild: what would have to be true for this to be wrong, and did they check? Train that, and you'll evaluate a study — or a viral headline — better than most people who wrote the textbook.
References
Darley, J. M., & Gross, P. H. (1983). A hypothesis-confirming bias in labeling effects. Journal of Personality and Social Psychology, 44(1), 20–33. https://doi.org/10.1037/0022-3514.44.1.20
Pfungst, O. (1911). Clever Hans (the horse of Mr. von Osten): A contribution to experimental animal and human psychology (C. L. Rahn, Trans.). Henry Holt.
Popper, K. R. (1959). The logic of scientific discovery. Hutchinson.