← All lessons
Lesson 09 · Correlational

Bivariate Correlational Research

You computed r last semester. This lesson is about believing it responsibly: what the number can and cannot say, why you must always look at the scatterplot, how big is "big," the two roads that block the path to causation — and why "correlation is worthless" is a myth that, in medicine, gets people killed.

Where this sits in the course: Lesson 03 introduced association claims — statements that two measured variables covary — and the validities you interrogate them with. Lessons 06–08 taught you to question the measures and the sample. Now we put it together for the simplest correlational design there is: two measured variables, one coefficient. Statistics taught you to compute r; this lesson teaches you to cross-examine it. (Regression, "controlling for," and mediation vs. moderation are Lesson 10 — here we stay bivariate.)

The through-liner is a scatterplot compressed into one number

A bivariate correlational study measures two variables — neither is manipulated — and asks how they travel together. The whole design rests on one act of translation: taking a cloud of dots and squeezing it into a single number between −1 and +1. That compression is powerful and lossy. Powerful, because Pearson's r captures direction and strength at a glance: the sign tells you direction (positive — the variables rise together; negative — one rises as the other falls), and the magnitude tells you strength (how tightly the dots hug a straight line). Lossy, because wildly different scatterplots can produce the identical r — a fact we'll prove in a moment with the most famous four datasets in statistics.

So be precise about what r is not. It is not a percentage — r = .50 does not mean the variables agree half the time (squaring it gives the shared variance: .50² = 25%). It is not a slope — it says nothing about how much Y changes per unit of X, only how consistently. It is not a verdict on causation — a point the second half of this lesson takes apart properly. And it is not a general-purpose relationship detector: r sees only straight-line association. Two variables can be strongly, lawfully related in a curve and hand you an r of zero with a straight face.

One more habit from Statistics worth carrying over: an r from a sample is an estimate, and estimates have wobble. A modern report gives the coefficient with its 95% confidence interval — r = .32, 95% CI [.18, .45] — so the reader can see the range of plausible true values. Two rules of thumb: small samples give wide, untrustworthy intervals (an r from n = 12 could plausibly be almost anything), and an interval that crosses zero means you can't rule out no relationship at all. A bare r with a p-value and no interval is a claim wearing sunglasses indoors — technically dressed, hiding something.

Go deeper · categorical variables correlate too

An association claim doesn't require two continuous variables. If one variable is categorical — diet group (vegan vs. omnivore) and, say, a self-compassion score — the "scatterplot" becomes a bar graph of group means, and the analysis you ran last semester was a t-test. It is still a correlational design: nobody was assigned to a diet, so the study supports an association claim, never a causal one. The design is defined by how the variables were obtained (measured, not manipulated), not by which statistic you happened to compute. This trips up more students than any formula does.

Anscombe's quartetWhy you must always, always plot

In 1973, the statistician Francis Anscombe built four small datasets as a deliberate ambush (Anscombe, 1973). Each has eleven points. Each has the same mean of X (9.0), the same mean of Y (7.5), the same correlation (r = .816), and the same regression line (Ŷ = 3 + 0.5X). By every summary statistic you learned last semester, they are quadruplets. Then you plot them.

DatasetWhat the scatterplot actually showsWhat r = .816 is hiding
IA well-behaved linear cloudNothing — this is the plot the statistics assume
IIA perfect, tight curve (arc)The relationship is curvilinear; a line is the wrong model entirely
IIIA perfect straight line — plus one outlierOne point is dragging the line off an otherwise flawless fit
IVA vertical stack of points at one X value, plus one point far rightThe entire correlation is manufactured by a single observation

Anscombe's point was not a party trick. It was an indictment of a workflow — compute the summary, skip the picture — that is still common half a century later. The quartet is the standing proof that an r you haven't seen as a scatterplot is an r you don't actually know. Every distortion in the drill below is invisible in the coefficient and obvious in the plot. That's the whole argument: plotting isn't decoration; it's inspection.

Interrogate itDiagnose what's distorting the correlation

When an exam — or a journal article — hands you a correlation, three culprits can fake, shrink, or hide the true relationship, and a fourth can make a real coefficient tell a false story. Read each scenario, make your diagnosis, then tap.

Which problem could be distorting each correlation? Tap one.

Pick a correlation to diagnose the distortion…

A real oneInterrogating a published association

Let's run the full interrogation on a real coefficient from the self-compassion literature — one you met in Lesson 03. MacBeth and Gumley (2012) meta-analyzed twenty samples and found an overall correlation of r = −.54 between self-compassion and psychopathology (depression, anxiety, and stress): people who treat themselves with more kindness under failure report substantially less distress. What does a methods student do with that number?

First, construct validity — interrogate the measures. Both variables are self-reports: the Self-Compassion Scale on one side, symptom inventories on the other. Are they reliable? Well validated? Mostly yes — but notice that when one person fills out both questionnaires in one sitting, some of the correlation can come from shared method: a respondent in a dark mood may rate everything more negatively, inflating the association between any two self-report scales. A strong association claim survives this question only if the measures have been validated against something beyond more self-report.

Second, statistical validity — interrogate the number. Is −.54 big? By any benchmark you'll meet below, yes — unusually large for individual-differences research. Is it precise? A meta-analysis pooling thousands of participants puts a tight confidence interval around it, which is exactly why meta-analytic estimates are worth more than any single study's r. Could it be an outlier or restriction-of-range artifact? Across twenty separate samples with full-range scales, far less likely than in one small study — consistency across samples is itself evidence.

And then, third: even granting that the association is real, well measured, and large — what caused what? That question the design simply cannot answer, which brings us to the two barriers.

Interrogate · the checklist for any association claim

When a headline shouts "People who do X are happier / healthier / smarter," run the list: (1) Construct validity — how well were both variables actually measured, and could shared method inflate the link? (2) Statistical validity — how big is the effect, how wide is the confidence interval, and is there more than a p-value on a huge sample? (3) The three distorters — did they show the scatterplot, or could an outlier, restricted range, or a curve be faking the number? (4) Directionality — could the arrow point the other way? (5) Third variable — what plausible C causes both? A claim that survives all five is a serious correlation. Almost no headline survives all five.

How big is big?Effect sizes, done right

Jacob Cohen gave the field its famous conventions: r ≈ .10 is "small," .30 is "medium," .50 is "large" (Cohen, 1988, 1992). They're exam gold and you should know them cold. But you should also know what Cohen himself said about them — that they were rough conventions offered for power planning when nothing better existed, not laws of nature. Fifty years of accumulated data have given us something better: context.

Funder and Ozer (2019) make the modern argument in two moves. First, judged against what psychology actually finds, Cohen's ladder is miscalibrated: surveys of published effects put the typical social-psychology correlation around r ≈ .20, so they recommend reading .05 as very small, .10 as small, .20 as medium, .30 as large — and treating an r of .40 or more in a single study less as a triumph than as a value that "likely overestimates" what will replicate. By that recalibrated scale, MacBeth and Gumley's −.54 — a meta-analytic estimate, not a lucky single study — is genuinely enormous.

Second, and more important: a "small" effect is not an unimportant effect. Effect size indexes predictive accuracy for one case at a time; importance depends on scale and stakes. The canonical example: in the Physicians' Health Study, aspirin's effect on preventing heart attack corresponded to roughly r ≈ .03 (Rosnow & Rosenthal, 2003; Funder & Ozer, 2019) — a value you'd sneer at in a lab report. The trial was stopped early because it was considered unethical to keep giving half the physicians a placebo: across millions of patients, that microscopic correlation is thousands of lives. Small effects also compound — a tiny per-instance effect, repeated daily across a population or a lifetime, accumulates into a large consequence. So the exam answer is Cohen's ladder; the thinking answer is Funder and Ozer's question: small compared to what, and consequential for whom, how often?

Update · where effect-size interpretation moved

The field's center of gravity has shifted from "is it significant?" to "how big is it, how precise is it, and does it matter?" Three habits mark a modern reader: report and interpret the effect size with its confidence interval, not just p; judge the effect against benchmarks from the actual literature (Funder & Ozer, 2019) rather than one-size-fits-all labels; and remember that with a big enough sample, a p-value goes significant for correlations too tiny to matter — statistical significance is about detectability, not importance. A study of two million people can make r = .01 "significant." Nothing about your life is different because of it.

The two barriersWhy association can't become causation

Recall the three causal criteria from Lesson 03 — covariance, temporal precedence, and the exclusion of alternative explanations (internal validity), the framework formalized by Cook and Campbell (1979) and Shadish, Cook, and Campbell (2002). A bivariate correlation establishes exactly one of the three: covariance. The two it fails deserve names you can weaponize on an exam.

The directionality problem (temporal precedence)

A correlation is symmetric — it doesn't say which came first. Take the −.54: does self-compassion protect against depression, or does depression corrode the ability to be kind to yourself? Both stories fit the same coefficient perfectly; both are psychologically plausible; the cross-sectional data cannot referee between them. Whenever you can flip the arrow and the story still makes sense, you have a directionality problem.

The third-variable problem (internal validity)

Some unmeasured C causes both A and B, manufacturing a correlation with no direct link between them. For the −.54: maybe neuroticism — a broad disposition toward negative emotion — independently lowers self-compassion scores and raises symptom scores. The association is a spurious association: real in the data, empty as a causal story. Naming a plausible, specific C is the skill; "maybe something else" earns no points.

Spurious correlations are best learned through the absurd ones. The classic: ice-cream sales track drowning deaths, courtesy of summer. The peer-reviewed joke: countries that eat more chocolate win more Nobel Prizes per capita — r = .79 across 23 countries, published in the New England Journal of Medicine as a deadpan warning about exactly this inference (Messerli, 2012); national wealth buys both chocolate and research universities. And the industrialized version: Tyler Vigen (2015) wrote software that trawls thousands of public time-series for accidental matches, yielding gems like U.S. per-capita cheese consumption correlating at roughly r ≈ .95 with deaths by bedsheet entanglement, and Nicolas Cage's annual film output tracking swimming-pool drownings. Vigen's correlations have no third variable at all — they are pure coincidence, the guaranteed by-product of comparing enough variables. That's the modern lesson stacked on the old one: with big data, meaningless correlations aren't a risk; they're a certainty. The coefficient can't tell you it's meaningless. Only thinking about mechanism can (Rohrer, 2018).

Go deeper · a preview of moderators

One more question a sharp reader asks of any association: is it the same for everyone? A moderator is a variable that changes the strength (or even the direction) of an association depending on its level — the link between A and B is strong for one group, weak or absent for another. Meta-analysts test these routinely: MacBeth and Gumley (2012) checked whether their −.54 varied across sample characteristics. Suppose the self-compassion–distress link ran strong in adults but weak in adolescents — that wouldn't debunk the association; it would specify it, telling you for whom the relationship holds and hinting at why. Moderators don't solve directionality or third variables, but they turn a flat "X relates to Y" into the more scientific "X relates to Y, under these conditions." Lesson 10 develops this properly.

Myth-busting"Correlation isn't causation" is not the end of thinking

Students over-learn the slogan and swing too far — treating every correlational finding as second-rate, dismissible on contact. That reflex has a body count.

Myth · "If you can't prove causation, the data is useless"

The evidence that smoking causes lung cancer is, to this day, correlational — no one ever randomly assigned humans to smoke for thirty years. Doll and Hill (1950) compared hundreds of lung-cancer patients with matched controls and found smokers dramatically overrepresented; the tobacco industry's response for decades was, essentially, "correlation isn't causation." (Even the great statistician Ronald Fisher, a pipe smoker, argued a constitutional third variable might cause both the craving and the cancer.) The slogan was technically true and practically catastrophic. What settled it was not an experiment but disciplined correlational reasoning: Austin Bradford Hill (1965) laid out considerations for judging when an association merits a causal interpretation — the association's strength, its consistency across studies and populations, its temporality (exposure precedes disease), a dose–response gradient (heavier smokers, more cancer), and plausibility of mechanism, among others. Smoking passed every test; no proposed third variable could. The mature skill is not chanting the slogan — it's knowing how correlational evidence is weighed when experiments are impossible or unethical. And prediction needs no cause at all: a credit score predicts default without causing it. A correlation you understand is a tool; the only mistake is claiming it proves a cause by itself.


Match the r value to its verbal description

Click an r, then click what it means (sign = direction, magnitude = strength). Get all six.

🂠

Name the third variable — if there is one

Each correlation is real in the data. What lurking variable causes both? Flip to check.

Ice-cream sales correlate with drowning deaths.
Tap to flip
Summer / hot weather

Heat drives both ice-cream buying and swimming. The correlation is real; ice cream is innocent.

Homes with more firefighters at the scene suffer more fire damage.
Tap to flip
Size of the fire

Big fires draw more firefighters AND cause more damage. Firefighters aren't the problem — and note this one also tempts a reversed causal arrow.

Countries that eat more chocolate win more Nobel Prizes (r = .79).
Tap to flip
National wealth

Rich countries afford both luxury chocolate and world-class universities. Published in the NEJM as a deadpan warning (Messerli, 2012).

Nicolas Cage's yearly film count correlates with swimming-pool drownings.
Tap to flip
Trick card: no third variable

Pure coincidence, mined from thousands of time-series (Vigen, 2015). Compare enough variables and absurd correlations are guaranteed — the coefficient can't flag its own meaninglessness.

Check yourself — Bivariate Correlation quiz

Eight questions. Instant feedback with the reasoning.


SourcesCited in APA 7

Anscombe, F. J. (1973). Graphs in statistical analysis. The American Statistician, 27(1), 17–21. https://doi.org/10.1080/00031305.1973.10478966
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum.
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155–159. https://doi.org/10.1037/0033-2909.112.1.155
Cook, T. D., & Campbell, D. T. (1979). Quasi-experimentation: Design and analysis issues for field settings. Rand McNally.
Doll, R., & Hill, A. B. (1950). Smoking and carcinoma of the lung: Preliminary report. British Medical Journal, 2(4682), 739–748.
Funder, D. C., & Ozer, D. J. (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156–168. https://doi.org/10.1177/2515245919847202
Hill, A. B. (1965). The environment and disease: Association or causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300.
MacBeth, A., & Gumley, A. (2012). Exploring compassion: A meta-analysis of the association between self-compassion and psychopathology. Clinical Psychology Review, 32(6), 545–552. https://doi.org/10.1016/j.cpr.2012.06.003
Messerli, F. H. (2012). Chocolate consumption, cognitive function, and Nobel laureates. New England Journal of Medicine, 367(16), 1562–1564. https://doi.org/10.1056/NEJMon1211064
Rohrer, J. M. (2018). Thinking clearly about correlations and causation: Graphical causal models for observational data. Advances in Methods and Practices in Psychological Science, 1(1), 27–42. https://doi.org/10.1177/2515245917745629
Rosnow, R. L., & Rosenthal, R. (2003). Effect sizes for experimenting psychologists. Canadian Journal of Experimental Psychology, 57(3), 221–237. https://doi.org/10.1037/h0087427
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Vigen, T. (2015). Spurious correlations. Hachette Books.

← Prev: Sampling All lessons Next: Multivariate Correlational Research →