← All lessons
Lesson 03 · Foundations

Claims & Validities

Every finding you will ever encounter — in a journal, a news feed, a grant proposal, an argument at dinner — reduces to two questions: what kind of claim is this? and which validity would break it? This lesson gives you the taxonomy for the first question and the four-part interrogation framework, built by Cook and Campbell, for the second. Nail this page and the rest of the semester is bookkeeping.

The core ideas this page builds on: variables vs. constants, measured vs. manipulated variables, conceptual vs. operational definitions, the three kinds of claims (frequency, association, causal), the three criteria for causation, verbs that signal each claim, and the four validities — construct, internal, external, and statistical conclusion — with sample interrogation questions. From here, we go deeper on the parts that reliably trip people up, and on where this framework actually came from.

The through-lineA claim is not an argument

This distinction is easy to breeze past — but it's the hinge of the whole lesson. A claim is a flag planted in the ground: "Electric cars are better for the environment." An argument is what makes that flag stand up: claim + evidence + reasoning. The evidence is the data; the reasoning is the logic connecting data to conclusion. A claim without an argument is just an opinion; an argument without a claim is a pile of facts with no point.

When you "interrogate" a study, you're really auditing the argument — asking whether the evidence and reasoning actually earn the flag. And here's the professional secret: researchers themselves rarely get this wrong in the journal article. The abstract of a correlational paper will say associated with, carefully. It's the press release, the news headline, and the social-media summary where a modest association quietly gets promoted to a cause. Your job as a methods-trained reader is to catch the promotion.

The taxonomyLearn to name the claim in one breath

Scientific claims come in three kinds, and this taxonomy isn't a quirk of any one textbook — it falls straight out of how many variables a claim involves and what the claim says about them. A frequency claim (sometimes called a descriptive claim) reports the level or rate of a single measured variable. An association claim says two or more measured variables covary — when one changes, the other tends to change too. A causal claim goes all the way: it says changing one variable produces a change in another. Each step up the ladder demands more of the evidence, which is exactly why naming the claim correctly is the first move in any interrogation.

The single most useful skill in this course: read a headline and classify it instantly. The trick is to count the variables and watch the verb.

ClaimVariablesSignal verbsExample headline
FrequencyOne measured variable"is / are / percent of…""4% of American adults identify as vegetarian." (real — Jones, 2023)
AssociationTwo+ measured, they covaryis linked to · is related to · is associated with · predicts"Self-compassion is linked to lower anxiety and depression." (real — MacBeth & Gumley, 2012)
CausalOne manipulated + one measuredcauses · affects · reduces · increases · leads to · makes"Writing self-compassionately reduces shame." (hypothetical)

Notice the machinery inside those examples. A variable is anything that varies across the people or observations in a study; a constant doesn't. In "4% of American adults identify as vegetarian," diet status is the one variable (vegetarian vs. not); "American adults" is a constant — it defines who was studied, it doesn't vary within the claim. And every variable exists at two levels: the conceptual definition (the abstract construct — "self-compassion" as a warm, non-judgmental stance toward one's own suffering) and the operational definition (how it was actually captured — say, scores on a 26-item self-report scale). Every claim is really a claim about constructs, made through operationalizations. Hold that thought; it becomes construct validity in a moment.

Association claims come in three flavors: positive (both go up together), negative (one up, one down — the self-compassion finding above is negative: more self-compassion, less psychopathology), and zero (no relationship). An association claim earns something a frequency claim can't: prediction. If self-compassion and anxiety covary, knowing someone's self-compassion score lets you predict their anxiety better than chance — no causal story required.

And a warning worth stating loudly: tentative language does not downgrade a causal claim. "Owning a dog may lower blood pressure" is still a causal claim — the verb lower gives it away, not the hedge. "May," "could," "suggests," and "seems to" are journalistic seatbelts, not logical demotions. The verb decides the claim type; the hedging only decides how confident the writer sounds while making it.

Interrogate · the 5 questions to ask of any claim you read

Before you believe a single headline: (1) What kind of claim is it — count the variables, watch the verb. (2) How was each variable operationalized — construct validity. (3) Who was in the sample, and does it generalize — external validity. (4) How big and how precise is the effect — statistical conclusion validity. (5) If it's causal, was it an experiment with random assignment — internal validity. Memorize these five. They are the whole course as a checklist.

Only one road to "cause"The three criteria for causation

A causal claim is the most demanding claim in science, and the demands aren't arbitrary — they descend from John Stuart Mill's (1843) canons of induction, the nineteenth-century logic that still underwrites every experiment you'll ever run. Mill argued that to infer that X causes Y, you must show that X and Y occur together, that X came first, and — his famous "method of difference" — that when everything else is held equal and only X differs, Y still differs. Translate that into modern terms and you get the three criteria every causal claim must clear, in order:

Covariance: the two variables are actually related — when X differs, Y differs. Temporal precedence: the cause came before the effect, not after and not ambiguously alongside. Internal validity (Mill's method of difference): no third variable or alternative explanation can account for the relationship. The only research design that reliably clears all three bars is the true experiment — one variable manipulated, another measured, and participants assigned to conditions at random, which scatters every confound (measured or unmeasured, known or unknown) evenly across groups.

Go deeper · why the third criterion is the hard one

Covariance and temporal precedence are the easy bars — a good correlational study can often clear both. It's internal validity that only an experiment can deliver. Take a hypothetical headline: "People who practice yoga are less stressed." The covariance may be real, and you could even measure yoga uptake before measuring stress. But you cannot rule out that people with the time, money, and disposition to take up yoga were already less stressed — or that some third variable (income, health-consciousness, flexible work schedules) drives both. No random assignment, no causal claim — full stop. This is the cash value of the mantra you've heard since Intro: correlation is not causation. It's not that correlations are worthless — they establish covariance and support prediction — it's that they can't do the one thing the verb "causes" requires.

A short historyWhere the four validities came from

The interrogation framework you're about to master has a lineage, and knowing it will make the pieces snap together. In 1957, Donald Campbell — one of the great methodologists of the twentieth century — asked a deceptively simple question: what factors make an experiment interpretable? His answer introduced the now-standard distinction between internal validity (did the treatment really cause the effect in this study?) and external validity (would the effect hold for other people, settings, and times?) (Campbell, 1957).

Two decades later, Thomas Cook and Campbell expanded the pair into a quartet in Quasi-Experimentation (Cook & Campbell, 1979), splitting each of the original two. From internal validity they carved off statistical conclusion validity — whether the study's statistical inferences about covariance are reasonable — and from external validity they carved off construct validity — whether the study's operations actually represent the abstract constructs in the claim. That four-part scheme received its canonical modern statement in Shadish, Cook, and Campbell (2002), the book methodologists simply call "Shadish, Cook, and Campbell." When this course says "the four validities," that's the source.

Construct validity has an even older root. Cronbach and Meehl (1955) confronted a hard problem: psychological constructs like intelligence, anxiety, or self-compassion have no gold-standard criterion sitting in the world to check a measure against. Their solution: a measure has construct validity to the extent that it behaves the way the theory of the construct says it should — correlating with what it should correlate with, diverging from what it shouldn't, responding to what should move it. Validation, on this view, is never finished; it's an ongoing scientific argument about a measure. We'll spend all of Lesson 06 there.

ValidityThe question it answersClassic source
ConstructDo the study's measures and manipulations actually capture the constructs in the claim?Cronbach & Meehl (1955); Cook & Campbell (1979)
InternalDid the manipulated variable — and nothing else — cause the change in the outcome?Campbell (1957)
ExternalDo the results generalize to other people, settings, and times beyond this study?Campbell (1957)
Statistical conclusionAre the statistical inferences — the existence, size, and precision of the effect — reasonable?Cook & Campbell (1979)

One terminology note so the literature doesn't confuse you: Cook and Campbell's term is statistical conclusion validity. You'll see it shortened to "statistical validity" in textbooks and in casual speech — same idea, and this page uses both.

The pairing moveEach claim has a validity it lives or dies by

The four validities are, in effect, the four ways a claim can be wrong. But you don't interrogate every claim with equal pressure on all four — each claim type has a validity it must satisfy first, because that's where its distinctive promise lives. A frequency claim promises a rate in a population, so it lives or dies on external validity: was the sample representative of the population named in the claim? Gallup can say "4% of American adults" because it draws random samples of American adults; your Instagram poll cannot. An association claim promises that two well-measured constructs genuinely covary, so it hinges on construct validity (were both variables measured well?) plus statistical conclusion validity (is the covariation real, and how big?). A causal claim promises that X produced Y, so it stands or falls entirely on internal validity — was it an experiment, and are confounds ruled out?

This claims-to-validities pairing, as a deliberate reading strategy, is a pedagogical move popularized in textbook form by Morling (2021), and it's genuinely useful: learn the claim, and you know which question to lead with. But don't mistake "priority" for "only." A frequency claim with a garbage measure fails on construct validity no matter how beautiful the sample; a flawless experiment run on 12 people can still fail on statistical conclusion validity. The priority validity tells you where to start the interrogation, not where to stop it.

ClaimPriority validityThe lead question
FrequencyExternalWho did they sample, and does it generalize to the population claimed?
AssociationConstruct + Statistical conclusionWere both variables measured well, and is the link real and strong?
CausalInternalWas it an experiment? Are confounds ruled out?
Update · the internal–external trade-off (and Mook's heresy)

Here's a tension worth naming outright: internal and external validity usually pull against each other. A tightly controlled lab experiment maximizes internal validity — but the sterile setting and convenience sample can limit external validity. A big naturalistic field study nails external validity but lets confounds sneak in. There's no free lunch; every design spends one validity to buy another. And in a famously contrarian paper, Douglas Mook (1983) argued that the trade is often worth making deliberately: many lab experiments were never trying to generalize to everyday life in the first place. They test whether a theory's prediction can happen under the right conditions — and for that purpose, "artificiality" is a feature, not a bug. The mature question isn't "is this study externally valid?" but "does this study's purpose require external validity?"

Cash in your Statistics courseStatistical conclusion validity — you already own the tools

Here's where last semester pays off. You built effect sizes, confidence intervals, and significance tests in Statistics; this course doesn't re-teach them — it puts them to work as a reading tool. Statistical conclusion validity asks: are the study's numerical conclusions reasonable, and how strong and precise is the effect? Two habits of mind matter far more than "is it significant?"

First: significance is not size. A p-value tells you how surprising the data would be if there were no effect at all — and, as Jacob Cohen (1994) spent a career pointing out, that is nearly all it tells you. It does not tell you the probability the hypothesis is true, and it says nothing about whether the effect is big enough to matter. With a huge sample, a microscopic effect will be "significant." So always ask for the effect size — the correlation r, or Cohen's d for group differences.

What counts as big? Funder and Ozer (2019) offer modern benchmarks for psychology: r ≈ .10 is small (but can be consequential when it accumulates across many people or occasions), r ≈ .20 is medium, r ≈ .30 is large — and an r of .40 or beyond in a single social-psychology study is more likely an overestimate than a miracle. Calibrate against a real example from this professor's own research area: the meta-analytic association between self-compassion and psychopathology (anxiety, depression, stress) is r = −.54 across 20 samples (MacBeth & Gumley, 2012) — a genuinely large association by any benchmark, which is precisely why self-compassion interventions are worth testing experimentally.

Significance ≠ size

A result can be statistically significant (unlikely under the null hypothesis) yet trivially small. With 40,000 participants, r = .03 will clear p < .05 — real, perhaps, but practically meaningless for most purposes (Cohen, 1994; Funder & Ozer, 2019). Always ask for the effect size, not just the p-value.

Confidence intervals

A single number lies about its own certainty. A confidence interval gives the range of values compatible with the data. "A 6-point boost, 95% CI [1, 11]" is honest; a bare "6-point boost" hides how wobbly the estimate is. Wide interval = imprecise result. Cumming (2014) calls the shift from significance rituals to estimation — effect sizes, CIs, meta-analysis — "the new statistics," and it's now mainstream practice.

Update · the replication crisis (coming attractions)

The deepest statistical-conclusion question is: would the result happen again? In the largest systematic check ever run, the Open Science Collaboration (2015) redid 100 published psychology studies. Ninety-seven percent of the originals had reported significant results; only 36% of the replications did, and replication effect sizes were about half the originals. The causes trace largely to weak statistical conclusion validity — chasing p < .05 with small samples, ignoring effect sizes, and undisclosed flexibility in analysis. The modern rule: a real result replicates. One flashy significant study is a hypothesis, not a fact. You'll meet the fixes — preregistration, open data, bigger samples — in Lesson 15.

Common trapsThree mistakes smart students make

Trap 1: letting the hedge fool you. "May," "might," and "could" never change a claim's type — only the verb does. "A plant-based diet may lower inflammation" is a causal claim wearing a raincoat.

Trap 2: treating the constant as a variable. In a hypothetical headline like "62% of college women report high stress this semester," it's tempting to see two variables (gender and stress). But gender doesn't vary within the claim — everyone described is a woman. One variable, one rate: frequency claim, interrogate the sampling.

Trap 3: dismissing association claims as worthless. "Correlation isn't causation" is true, but correlational findings establish covariance (criterion one of three), enable prediction, and are often the only ethical option — you cannot randomly assign people to childhood trauma, to a vegan identity, or to twenty years of smoking. The error isn't making association claims; it's promoting them.


Which claim is this headline? Tap one to reveal the answer and the tell.

Pick a headline to see its claim type and the interrogation it invites.

Match each validity to the question it answers

Click a validity, then click the question it's really asking. Four to clear — straight out of Shadish, Cook, and Campbell (2002).

The causal checklist, in order

A study must clear these three bars — in this sequence — to earn the word cause. Arrange them (Mill, 1843, would approve).

🂠

Name the claim — then check

Read the headline, decide the claim type in your head, then flip.

"62% of college women report high stress this semester."
Tap to flip
Frequency

One measured variable, one rate. Gender here is a constant, not a second variable — it doesn't vary within the claim. Lead with external validity: who was sampled?

"More self-compassionate people report less anxiety and depression."
Tap to flip
Association (negative)

Two measured variables that covary in opposite directions — and a real one: meta-analytic r = −.54 (MacBeth & Gumley, 2012). Large, but no manipulation, so no cause.

"Eating meat makes men feel more masculine."
Tap to flip
Causal

The verb makes implies meat consumption was manipulated to move felt masculinity. Needs a true experiment with random assignment to be believed — interrogate internal validity.

"A significant result with 40,000 people, but r = .03."
Tap to flip
Weak statistical conclusion validity

Significant but trivially tiny. A giant sample makes almost anything significant. The effect size (r = .03) says: possibly real, practically meaningless (Funder & Ozer, 2019).

Check yourself — Claims & Validities quiz

Eight questions. Instant feedback with the reasoning.


SourcesCited in APA 7

Campbell, D. T. (1957). Factors relevant to the validity of experiments in social settings. Psychological Bulletin, 54(4), 297–312. https://doi.org/10.1037/h0040950
Cohen, J. (1994). The earth is round (p < .05). American Psychologist, 49(12), 997–1003. https://doi.org/10.1037/0003-066X.49.12.997
Cook, T. D., & Campbell, D. T. (1979). Quasi-experimentation: Design and analysis issues for field settings. Rand McNally.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
Cumming, G. (2014). The new statistics: Why and how. Psychological Science, 25(1), 7–29. https://doi.org/10.1177/0956797613504966
Funder, D. C., & Ozer, D. J. (2019). Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2), 156–168. https://doi.org/10.1177/2515245919847202
Jones, J. M. (2023, August 24). In U.S., 4% identify as vegetarian, 1% as vegan. Gallup. https://news.gallup.com/poll/510038/identify-vegetarian-vegan.aspx
MacBeth, A., & Gumley, A. (2012). Exploring compassion: A meta-analysis of the association between self-compassion and psychopathology. Clinical Psychology Review, 32(6), 545–552. https://doi.org/10.1016/j.cpr.2012.06.003
Mill, J. S. (1843). A system of logic, ratiocinative and inductive (Vols. 1–2). John W. Parker.
Mook, D. G. (1983). In defense of external invalidity. American Psychologist, 38(4), 379–387. https://doi.org/10.1037/0003-066X.38.4.379
Morling, B. (2021). Research methods in psychology: Evaluating a world of information (4th ed.). W. W. Norton.
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.

← Prev: Sources of Information Next: Ethical Guidelines →