What this lesson covers: the four ways people fix their beliefs (Peirce) and why empiricism wins; theory, hypothesis, and data and the loop that connects them; Popper's falsifiability and the logic of "supported, never proven"; the working assumptions of science (determinism, empiricism, parsimony, testability); Merton's norms and why science is public, cumulative, and self-correcting; basic vs. applied vs. translational research; and the classic objections to calling psychology a science — answered.
Ways of knowingHow do you know what you know?
Start with a rude question: of all the things you're sure of right now, how did each one get into your head? The philosopher Charles Sanders Peirce asked exactly this in 1877, in an essay called "The Fixation of Belief" — and his answer is still the cleanest map of the territory (Peirce, 1877). Peirce noticed that doubt is uncomfortable. It itches. So the mind works to settle it — to "fix" belief — and there are only a few ways to do that.
The first is tenacity: pick a belief, cling to it, and steer around anything that might disturb it. Never read the other side, never sit next to the skeptic at dinner. Peirce admits, almost admiringly, that this works — for a while. Its fatal flaw is what he called the social impulse: sooner or later you meet perfectly sensible people who believe the opposite, and the itch comes back.
The second is authority: outsource the settling to an institution — a state, a church, a guru, a podcast host with great lighting. Authority scales better than tenacity; it can fix belief for millions at a time. But authorities disagree with each other, contradict their past selves, and can't rule on everything. And "believe it because someone confident said it" gives you no way to tell the trustworthy authority from the merely loud one.
The third is the a priori method: believe whatever feels most reasonable — most elegant, most agreeable to reflection. This is the philosopher's temptation, and Peirce's verdict is brutal: it makes belief a matter of intellectual fashion. What "stands to reason" in one century embarrasses the next. Reasoning from the armchair polishes beliefs; it doesn't test them.
What all three methods share is the direction of travel: they start from the belief and work to protect it. Peirce's fourth method — the method of science, which this course calls empiricism — reverses the flow. Beliefs should be fixed, he argued, by "some external permanency": something on which our thinking has no effect. Reality doesn't care what you'd prefer to be true, which is exactly what makes it a trustworthy referee. Systematic observation lets the world push back — and the push-back is the point. Science is the only method on the list with a built-in way of discovering that you are the one who's wrong.
The engineThe theory–data cycle
So science fixes belief by observation. But not by casual observation — by observation organized into a loop. You start with a theory: a general statement about how variables relate, built to explain a body of existing facts. From the theory you derive a hypothesis: a specific, testable prediction about observations you haven't made yet. You collect data. The data either fit the prediction (the hypothesis is supported) or clash with it (disconfirmed). Either way, the theory comes out sharper than it went in, and the loop runs again. That's the theory–data cycle, and everything else in this course is machinery bolted onto it.
Make it concrete. Neff (2003) proposed a theory of self-compassion: treating your own failures with the same kindness you'd offer a struggling friend should reduce the threat and self-attack that failure normally triggers. That's a theory — a general claim about how constructs relate. From it you can derive hypotheses by the dozen: people higher in self-compassion should report less anxiety after failure feedback; a brief self-compassion writing exercise should lower post-failure cortisol; self-critical perfectionists should benefit most. Each hypothesis names variables you could actually measure and an outcome that could actually fail to appear. Twenty years of data later, the theory has been supported in some places, complicated in others, and revised accordingly — which is the cycle working, not failing.
Notice the division of labor. Theories are never tested directly — they're too general. Hypotheses do the risky work of touching data. And the two outcomes of that contact are not symmetrical, which brings us to the most important idea on this page.
PopperFalsifiability: a good theory sticks its neck out
Vienna, 1919. A young Karl Popper is surrounded by three intellectual movements, each claiming the mantle of science: Freud's psychoanalysis, Adler's individual psychology, and Einstein's general relativity. What struck Popper wasn't which theory had the most supporting evidence. Freud and Adler had endless supporting evidence — that was the problem. Whatever a patient did, the theory could explain it. A man who drowns a child and a man who dies saving one could both be explained by Adler's inferiority feelings, effortlessly, after the fact. A theory that fits every possible outcome forbids nothing — and a theory that forbids nothing tells you nothing (Popper, 1963).
Einstein's theory behaved differently. It made a wild, precise prediction: starlight passing near the sun should bend by a specific amount, observable during an eclipse. If the 1919 eclipse measurements had come out otherwise, relativity was dead — publicly, quantitatively dead. It survived. And Popper realized that this was the difference that mattered: not confirmation, but risk. The scientific status of a theory is its falsifiability — whether it makes predictions that could plausibly turn out false (Popper, 1959/2002, 1963). Confirmations count only when they come from risky predictions. Every good theory is a prohibition: it forbids certain things from happening. The more it forbids, the more it says.
Students hear "falsifiable" and flinch, as if it were an insult. It's the opposite — it's the entry fee. "All swans are white" is scientific precisely because one black swan sinks it. "Everything happens for a reason" is not, because no conceivable observation could ever contradict it. And note the bar is higher than vague testability: a horoscope is technically "testable," but its predictions are so elastic ("you may face a challenge this week") that nothing can miss. Popper's demand is that the prediction be specific enough to fail. When you evaluate any claim this semester, ask the Popper question first: what observation would prove this wrong? If the honest answer is "nothing," you've left science — you're just telling stories that always fit.
Here's the logical machinery underneath, and it's worth thirty seconds of your life. Suppose your theory implies a prediction: if T, then O. Now run the two possible outcomes. You observe not-O: the theory has a real problem — logicians call this modus tollens, and it's a valid argument. You observe O: tempting to declare victory, but concluding "therefore T" is a formal fallacy (affirming the consequent), because other theories predict O too. Disconfirmation can be logically decisive; confirmation never is. That asymmetry is the entire reason for the vocabulary rule you'll be graded on all semester: data support a hypothesis, are consistent with a theory — they never prove it.
Say your theory predicts X, you run the study, and — X happens. Victory? Not quite. Your data are consistent with your theory but can't rule out every rival that would have made the same prediction. Philosophers call this underdetermination: evidence never fully pins down a single explanation. There's a mirror-image catch on the falsification side, too: a failed prediction might mean the theory is wrong, or that your measure was broken, your sample was odd, or one of a hundred background assumptions failed. So in practice single studies neither prove nor cleanly disprove. A theory earns confidence the only way it can — by surviving many risky tests, run by many independent people, any one of which could have killed it. Science is less a courtroom verdict than a long, ongoing stress test.
The four canonsThe load-bearing assumptions
Before you can play the science game at all, you have to accept four house rules. They aren't conclusions science reached; they're the starting assumptions that make the enterprise possible. Tap each to see what it's really claiming — and what collapses without it.
The Four Canons of Science. Tap one to unpack it.
Parsimony deserves one more beat, because students often mistake it for decoration. Around 1900, a horse named Clever Hans appeared to solve arithmetic problems, tapping out answers with his hoof while Europe swooned. Two theories fit the data: (1) a horse had mastered mathematics, or (2) the horse was reading the tiny, involuntary postural cues of the humans around him, who relaxed the instant his taps reached the right number. Careful testing — blindfold the horse, or use questioners who didn't know the answer — settled it in favor of the boring theory. That's parsimony doing real work: the extravagant explanation was possible, but the simple one carried fewer new assumptions and, on testing, carried the day. When a study of yours admits two explanations, the simpler rival is the one reviewers will beat you with. Learn to beat yourself with it first.
This is the single most expensive misunderstanding in science literacy. To the layperson, "theory" means hunch. In science it means something closer to the opposite: a well-organized explanation that has survived repeated risky testing. Even trickier — a theory can also be a fact. Gravity is a theory (our explanation of why things fall) and a fact (things demonstrably fall). Evolution is a theory (the mechanism) and a fact (populations demonstrably change). When someone sneers that evolution is "only a theory," they've confused the everyday word with the scientific one (Lilienfeld et al., 2010). Don't be that someone on the exam — and notice the flip side: "theory" being an honorific doesn't make any particular theory true. It makes it testable and so-far-surviving. That's the highest compliment science pays.
Before you treat any statement as science, run three quick filters. (1) Is it about the observable world? "Self-compassion predicts lower anxiety" is; "the soul is eternal" isn't. (2) Could evidence settle it? If you can't imagine data that would change your mind, it isn't an empirical question. (3) Is it falsifiable — specifically enough to fail? Does it forbid some outcome? If a claim passes all three, it's an empirical question and belongs in this course. If it fails one, it may still matter enormously — moral and aesthetic questions usually do — it's just not a question data can answer. Knowing which kind of question you're holding is half of methodological competence.
Merton's normsScience is public, cumulative, and self-correcting
Everything so far describes the logic of science. But science isn't done by logic; it's done by people — ambitious, biased, mortgage-paying people. So why should the output be trustworthy? The sociologist Robert Merton's answer is that science works not because scientists are saints but because the community is structured so that individual bias gets caught. He identified four institutional norms (Merton, 1942/1973). Communality: findings belong to everyone — you must publish your methods and results, not hoard them, so others can build on and check them. Universalism: claims are judged by the evidence, not by the fame, nationality, or credentials of whoever makes them — a first-year student's data can dethrone a full professor's theory. Disinterestedness: the institution rewards getting it right, not getting your way; conflicts of interest must be disclosed and designed around. And organized skepticism: nothing is exempt from challenge — every claim, however cherished, stays permanently on trial.
Those norms are why science, alone among Peirce's methods, is self-correcting. Tenacity, authority, and armchair reason all dig in when challenged; science is built to do the opposite. They're also why it's cumulative: because findings are published in a permanent, public record, each generation starts where the last one stopped instead of starting over. And they're why peer review exists — anonymous experts trying to poke holes in a study before publication. But hear this clearly: being published is not a finish line. It's the start of the longer trial of replication and challenge. That's also why the research-to-headline pipeline is so dangerous — a headline sells the unfinished draft as a settled verdict.
The famous headline claimed classical music raises IQ; an industry of baby-genius CDs followed. The actual study found something far narrower: college students — not babies — showed a small boost on spatial-temporal tasks that faded within about 15 minutes (Rauscher, Shaw, & Ky, 1993). No general IQ gain, ever. Then organized skepticism went to work: follow-up experiments showed the bump wasn't about Mozart at all — it tracked arousal and mood, and a lively Schubert piece or an engaging story produced the same short-lived lift (Thompson, Schellenberg, & Husain, 2001). The original finding was real; the headline was fiction; and the correction happened in the open literature, on the record. That whole arc — modest finding, inflated claim, public correction — is Merton's machinery running exactly as designed.
Does the machinery still work at scale? In 2015, 270 researchers re-ran 100 published psychology studies and got significant results in fewer than half of the replications, with effect sizes about half the originals (Open Science Collaboration, 2015). Read carelessly, that's a scandal. Read correctly, it's the immune system responding: psychology turned organized skepticism on itself, published the damage report, and rebuilt its practices — preregistration, open data, bigger samples — in response. No horoscope column has ever audited its own hit rate. The last lesson of this course returns to this story in full; for now, take the moral: self-correction isn't a slogan about science, it's a behavior, and you can watch it happen.
Why we do itBasic, applied, translational
One more distinction before the worked example, because it tells you what kind of "why" a study is answering. Basic research aims at understanding for its own sake: how does memory consolidate, why do infants attach, what is self-compassion made of? Applied research targets a practical problem in a real setting: does this eight-week program reduce burnout in nurses? Translational research is the bridge — deliberately carrying basic findings into application, the way basic work on attachment and on self-kindness eventually became structured therapeutic programs. The categories describe the goal, not the quality: Harlow's monkey studies (below) were basic research with no application in sight, and they ended up reshaping hospital visiting policies and adoption practice. Curiosity has a long fuse.
| Type | Driving question | Example |
|---|---|---|
| Basic | How does this work? | What are the components of self-compassion, and how do they relate to self-esteem? (Neff, 2003) |
| Applied | Does this fix the problem? | Does a self-compassion training program reduce anxiety in first-year nursing students? |
| Translational | How do we carry the lab finding into the clinic or classroom? | Turning basic findings on self-kindness and threat responses into a structured, testable intervention protocol |
ObjectionsSo is psychology actually a science?
Every methods professor eventually meets the raised eyebrow: chemistry is a science; is psychology, really? Take the classic objections seriously, because answering them is a test of everything above.
"It's just common sense." Then common sense should survive testing — and it routinely doesn't. People "know" that opposites attract, that we use only 10% of our brains, that venting anger releases it safely. All are false; whole books catalog such casualties (Lilienfeld et al., 2010). The venting one is instructive: catharsis "stands to reason" (the a priori method, working exactly as Peirce warned), but when Bushman (2002) randomly assigned angry participants to hit a punching bag while thinking about their provoker, they became more aggressive afterward, not less. Doing nothing beat venting. Where common sense and data collide, one of them updates — and it isn't the data.
"You can't directly observe the mind." True — and no physicist has ever directly observed an electron. Sciences routinely study unobservables by tying them to observable indicators and testing the predictions that follow. Psychology infers memory from recall performance, anxiety from behavior and physiology and report. The inference must be disciplined — that's what Lesson 06 on measurement is for — but "invisible, therefore unscientific" would take down half of physics with us.
"People are too variable — there are no laws of behavior." Psychological regularities are probabilistic, not clockwork. So is meteorology; so is epidemiology; so is genetics. "Lawful" doesn't mean "every case identical" — it means the variability itself has structure you can model and predict better than chance. Determinism, remember, requires order, not simplicity.
"Psychology's findings keep failing to replicate." Some do — and psychology found that out about itself, published it, and reformed (Open Science Collaboration, 2015). The objection accidentally concedes the point: only a real science can fail in this particular way, because only a real science makes claims specific enough to fail. What makes a field scientific was never an unblemished record. It's the method — falsifiable claims, public data, organized skepticism — and psychology plays by those rules.
The worked exampleHarlow's monkeys: theory vs. theory
Why does an infant bond to its mother? At mid-century, two theories competed. Cupboard theory — the behaviorist favorite — said the bond is really about food: the mother is loved because she feeds, and "love" is just conditioned association with milk. Contact-comfort theory said the bond is about touch and warmth, with feeding incidental. Notice the setup: two rival explanations that make different, testable predictions about the same behavior. This is exactly the configuration falsifiability wants, because whatever happens, one theory takes real damage.
Harlow (1958) gave infant rhesus monkeys two surrogate "mothers": a bare wire figure that dispensed milk, and a soft terrycloth figure that gave no milk. Cupboard theory makes a risky prediction — the infants should attach to the wire feeder. They didn't. They spent up to eighteen hours a day clinging to the cloth mother, darting to the wire mother only to feed, then racing back. When a frightening wind-up toy invaded the cage, they fled to cloth, not to milk. The prediction failed; cupboard theory was disconfirmed. And note the asymmetry doing its work: this one study genuinely wounded cupboard theory, but it didn't prove contact-comfort theory — it supported it, and decades of subsequent attachment research were needed to build the confidence we now have. One clean study, two rival theories, one prediction that failed: the theory–data cycle, run at full compression.
SourcesCited in APA 7
Bushman, B. J. (2002). Does venting anger feed or extinguish the flame? Catharsis, rumination, distraction, anger, and aggressive responding. Personality and Social Psychology Bulletin, 28(6), 724–731. https://doi.org/10.1177/0146167202289002
Harlow, H. F. (1958). The nature of love. American Psychologist, 13(12), 673–685. https://doi.org/10.1037/h0047884
Lilienfeld, S. O., Lynn, S. J., Ruscio, J., & Beyerstein, B. L. (2010). 50 great myths of popular psychology: Shattering widespread misconceptions about human behavior. Wiley-Blackwell.
Merton, R. K. (1973). The normative structure of science. In N. W. Storer (Ed.), The sociology of science: Theoretical and empirical investigations (pp. 267–278). University of Chicago Press. (Original work published 1942)
Neff, K. D. (2003). Self-compassion: An alternative conceptualization of a healthy attitude toward oneself. Self and Identity, 2(2), 85–101. https://doi.org/10.1080/15298860309032
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Peirce, C. S. (1877). The fixation of belief. Popular Science Monthly, 12, 1–15.
Popper, K. R. (1963). Conjectures and refutations: The growth of scientific knowledge. Routledge & Kegan Paul.
Popper, K. R. (2002). The logic of scientific discovery. Routledge. (Original work published 1959)
Rauscher, F. H., Shaw, G. L., & Ky, K. N. (1993). Music and spatial task performance. Nature, 365(6447), 611. https://doi.org/10.1038/365611a0
Thompson, W. F., Schellenberg, E. G., & Husain, G. (2001). Arousal, mood, and the Mozart effect. Psychological Science, 12(3), 248–251. https://doi.org/10.1111/1467-9280.00345