Here's the single most useful skill you'll take from this course, more useful than any fact you'll memorize: reading a study and knowing exactly what it's entitled to claim. Every week you'll see headlines built on psychological research — "Study finds X causes Y," "New research links A to B" — and most of them are quietly overstating what the underlying design can support. This unit gives you the tools to catch that.
The horse that could do math
In the early 1900s, a horse named Clever Hans became a sensation in Germany. His owner, a retired schoolteacher, claimed Hans could solve arithmetic problems, tell time, and even spell — tapping out answers with his hoof. Investigators tested him rigorously and found no trickery: Hans got it right even with strangers asking the questions, even without his owner in the room. It looked like real equine intelligence.
Then the psychologist Oskar Pfungst ran a cleverer test. He had questioners who didn't know the answer themselves ask Hans the questions. Hans's accuracy collapsed (Pfungst, 1911). It turned out Hans was reading tiny, involuntary cues in his questioner's posture and facial tension — cues that appeared the moment the person expected the correct hoof-tap and relaxed the instant it happened. Hans wasn't doing math. He was doing something arguably more impressive: reading people. But he wasn't answering the question anyone thought they were asking him.
◆ The big idea
Without controls and blinding, you often end up measuring the experimenter, not the subject. Clever Hans is the origin story for experimenter-expectancy effects and for why researchers now build blind procedures into almost everything — so that neither the participant nor (ideally) the person running the session knows which condition is which, and can't unconsciously nudge the result.
Where you study people: lab versus field
Every study happens somewhere, and where it happens trades off two things you can't fully have at once. In a lab setting, you control the environment tightly — lighting, noise, timing, exact wording of instructions — which lets you isolate the variable you care about. That buys you internal validity: confidence that what you manipulated is actually what produced the effect. In a field setting, you study people in their everyday environments — a real classroom, a real workplace — which buys you external validity: confidence that the finding generalizes to real life, because it was observed in real life.
Neither setting is "better." A tightly controlled lab study of aggression using a contrived paradigm might nail down a causal mechanism beautifully but leave you unsure whether it plays out the same way at a bar on a Friday night. A field study of the bar might have terrific real-world relevance but a dozen uncontrolled variables tangled together. This is the internal-versus-external validity tradeoff, and it shows up constantly when you evaluate a study's setting.
Bodhi says
When you read a methods section, ask "where did this happen?" before you ask "what did they find?" A lab finding and a field finding earn different kinds of trust — neither is automatically the stronger one.
The methods ladder
Psychological methods form a rough ladder, and each rung answers a fundamentally different question. Climbing the ladder buys you more certainty about cause and effect, but it costs you something too — usually realism, ethical feasibility, or both. Work through the three rungs below.
Step 1 · Descriptive / observational — "What happens?"
This is the ground floor: case studies, naturalistic observation, surveys. You watch or ask, and you describe what you see without manipulating anything. It's how you generate hypotheses and catch phenomena you didn't know to look for. What it can't do is tell you why something happens or what would happen under different conditions — there's no comparison built in.
Step 2 · Correlational — "What's related to what?"
Now you're measuring two or more variables and asking whether they move together — does more of one tend to come with more (or less) of the other? Correlational designs let you predict and let you study things you could never ethically manipulate (you can't randomly assign people to have anxiety, but you can measure whether anxiety correlates with sleep quality). The cost: a correlation, no matter how strong, never tells you which variable caused which — or whether something else caused both. More on that below.
Step 3 · Experimental — "What causes what?"
Only the true experiment lets you make causal claims, because only here does the researcher actively manipulate a variable and randomly assign participants to conditions. This is the top of the ladder — the most powerful design for isolating cause and effect — and also the most demanding: it requires a manipulable variable, an ethical way to manipulate it, and often a lab-like level of control that can cost you realism.
Correlation is not causation
This is the phrase you'll hear repeated so often it risks becoming background noise — so let's make sure you actually understand why it's true, not just that it's true. A correlation tells you two variables move together: as one goes up, the other tends to go up (a positive correlation) or down (a negative correlation), and the strength of that relationship can range from weak to strong. What a correlation cannot tell you, no matter how strong or how statistically reliable, is which variable is doing the causing — or whether either one is causing the other at all.
There are two classic reasons a correlation can mislead you about causation. The first is the directionality problem: if A and B are correlated, maybe A causes B, but it's equally possible that B causes A. The correlation itself can't tell you which way the arrow points. The second is the third-variable problem: some hidden factor, C, might be causing both A and B independently, creating a correlation between them that has nothing to do with either causing the other.
The classic illustration: ice cream sales and drowning deaths rise and fall together across the year, and the correlation is strong. Does buying ice cream cause drowning? Obviously not. The third variable is hot weather — it drives people to both buy ice cream and go swimming, and swimming (not soft-serve) is what raises drowning risk. Whenever a correlation strikes you as surprising, your first move should be to ask what third variable might be driving both sides of it. Related to this: watch out for illusory correlation, where people perceive a relationship between two things that isn't actually there, often because a few vivid or memorable co-occurrences stick in memory while the many times the pattern didn't hold get ignored.
The myth
"A strong correlation is basically proof of causation — if the statistics are solid, you can trust the causal story."
What's actually true
Statistical strength has nothing to do with causal proof. A correlation can be enormous, highly statistically significant, and replicated across ten studies — and still tell you nothing about which variable causes which, because directionality and third variables are never ruled out by the correlation itself. Only a true experiment, with manipulation and random assignment, can rule those out.
The true experiment
An experiment earns its causal claims through a specific structure. The researcher manipulates an independent variable (the thing you change) and measures its effect on a dependent variable (the outcome you're tracking). Participants who receive the manipulation make up the experimental group; participants who don't — or who get a comparison condition — make up the control group.
The engine that makes causal inference possible is random assignment: every participant has an equal chance of ending up in any condition. Random assignment doesn't guarantee the groups are identical, but across a reasonably sized sample it equates them on average on everything you didn't measure — personality, mood, history, genetics, whatever. That's what lets you attribute a difference between groups to the manipulation rather than to some preexisting difference between the people in each group. When groups differ on something other than the intended manipulation, that something is a confound — and a good experimental design exists largely to prevent confounds from creeping in.
Two more pieces close the loop. A placebo is a fake treatment given to the control group so both groups believe they might be receiving the real thing — this controls for the power of expectation itself. In a single-blind study, participants don't know which condition they're in; in a double-blind study, neither participants nor the researchers interacting with them know — which, as Clever Hans taught us, matters enormously, because researchers who know the hypothesis can unconsciously communicate it.
◆ What the slides usually skip
Notice that a placebo control and blinding are solving two different problems. The placebo controls for participants' expectations. Blinding controls for the researcher's expectations leaking into how they treat participants or interpret ambiguous behavior. A well-designed drug trial typically needs both.
Sampling versus assignment — the mix-up everyone makes
These two phrases sound alike and get confused constantly, but they answer completely different questions and control completely different things.
Random sampling
The question: who ends up in the study at all? If participants are drawn randomly from a larger population, every member of that population had an equal chance of being selected. This is what lets you generalize findings from your sample back to the broader population. Most psychology studies do not use true random sampling — convenience samples (whoever signs up) are far more common, which limits generalizability regardless of how good the rest of the design is.
Random assignment
The question: once someone is in the study, which condition do they get? Randomly assigning already-recruited participants to experimental versus control conditions is what supports causal inference — it's what rules out confounds by equating the groups on average. You can have random assignment with zero random sampling (a convenience sample randomly split into two groups), and that's exactly what most psychology experiments actually look like.
Operational definitions, reliability, and validity
Before you can measure anything, you need an operational definition — a precise, concrete statement of exactly how a concept will be measured or manipulated in this particular study (not "aggression" in the abstract, but "number of times a participant pressed a button delivering a loud noise to a confederate"). Two more terms travel together and get conflated constantly: reliability is whether a measure gives consistent results — measure the same stable thing twice, get roughly the same answer. Validity is whether a measure actually captures the concept it claims to capture. A bathroom scale that's off by ten pounds every time is reliable (consistent) but not valid (not accurate); a measure has to be reliable to have any chance of being valid, but reliability alone doesn't guarantee it.
Whose data is this, anyway? The WEIRD problem
Even a flawless experiment inherits the limits of who's in it. Henrich, Heine, and Norenzayan (2010) pointed out that the overwhelming majority of psychology's published findings come from samples that are Western, Educated, Industrialized, Rich, and Democratic — WEIRD, for short. That's not a coincidence; it reflects who researchers have easy access to, largely undergraduates at universities in wealthy Western countries. The trouble is that WEIRD populations are, on many psychological dimensions, statistical outliers relative to humanity as a whole — not a neutral default that everyone else deviates from.
This is a generalizability problem layered on top of everything else you've learned in this unit. A study can have a large sample, random assignment, careful blinding, and still tell you relatively little about how the effect works outside a narrow slice of humanity. When you read "participants," it's worth asking who, specifically, and how far that finding is likely to travel.
Studying change over time
Some questions require tracking how people change, and there are two main ways to design that. A cross-sectional design compares different age groups all at once — measure 10-, 20-, and 40-year-olds today and compare them. It's fast and cheap, but it can't distinguish true developmental change from a cohort effect: differences between age groups that exist because they grew up in different eras with different experiences, not because of aging itself. A longitudinal design follows the same group of people over time, which sidesteps the cohort confound but costs years (sometimes decades), money, and participants who drop out along the way.
Current research: the replication crisis and open science
◆ Where the field is now
Starting in the early 2010s, psychology went through a hard reckoning. The Open Science Collaboration (2015) attempted to directly replicate 100 studies published in top psychology journals — and found that a substantial share of the original effects did not reproduce, or reproduced much more weakly than first reported. That's not a scandal so much as science working as intended: a field checking its own results and finding out where it had gotten ahead of the evidence.
The response has been constructive. Preregistration — publicly committing to your hypotheses, sample size, and analysis plan before collecting data — has moved from a niche practice to something close to a norm in many corners of the field, precisely because it closes off the flexibility that let earlier researchers unconsciously fish for significant results (Nosek, Ebersole, DeHaven, & Mellor, 2018). Data-sharing norms have shifted too, with more journals expecting raw data to be posted publicly so others can check the work. This is authentically what working researchers do now — it isn't a historical footnote, it's the water the field currently swims in, and you'll see it reflected in how modern studies report their methods.
Bodhi says
A field that revises its own famous findings isn't a broken field — it's a field taking its own methods seriously. Keep that distinction in mind before you let "some studies didn't replicate" curdle into "psychology isn't real science."
The one thing to carry out of this unit
Before you believe any claim about what causes what, ask what the design actually was. Was anything manipulated? Was assignment to conditions random? Was the sample randomly drawn, or just whoever was available — and how WEIRD is it? Was the measure operationally defined, reliable, and valid? Get in the habit of asking these questions automatically, and you'll be equipped to evaluate research claims — in this course and for the rest of your life — far more rigorously than most people ever learn to.
References
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
Pfungst, O. (1911). Clever Hans (the horse of Mr. von Osten): A contribution to experimental animal and human psychology (C. L. Rahn, Trans.). Henry Holt. (Original work published 1907)
Rosenthal, R., & Rosnow, R. L. (2008). Essentials of behavioral research: Methods and data analysis (3rd ed.). McGraw-Hill.