Every lesson so far has been about getting one study right: a valid measure, a fair sample, a clean manipulation, threats ruled out. This lesson zooms out to the question that decides whether any of it matters. Science is not a collection of studies; it's a collection of findings that survive being checked. So we end the course with the checking: what replication is, how psychology discovered that much of its literature couldn't pass the check, exactly why honest researchers were producing findings that weren't there, and the reforms — preregistration, Registered Reports, open data, big collaborative samples — that rebuilt the field's norms in about a decade. Then we push past "would it happen again?" to the harder question: "would it happen again somewhere else, with someone else?"
Three ways to say "again"Direct, conceptual, and replication-plus-extension
Replication isn't one thing. A direct replication repeats the original study as closely as possible — same procedure, same measures, new participants — to ask the bluntest question available: is the effect real? Suppose a published experiment reports that a ten-minute self-compassionate writing exercise reduces rumination after a failure. A direct replication runs that exact protocol again with fresh people. If the effect shows up, the fact is confirmed; if it doesn't, something is wrong with the original, the replication, or both. What a direct replication can't do is tell you whether the idea travels — it never leaves the original context, and it inherits every design flaw the original had.
A conceptual replication keeps the same theoretical variables but changes how they're operationalized. Same hypothetical: instead of a writing exercise and a rumination questionnaire, a researcher tests whether a two-week self-compassion phone practice changes how long people persist on a task after failing it. The constructs (self-compassion, response to failure) are identical; the operations are new. If the effect appears anyway, you've learned something a hundred direct replications couldn't teach you — the theory is doing the work, not some quirk of one procedure. Conceptual replication is how construct validity and external validity get built.
A replication-plus-extension reproduces the original and adds something — a new condition, a new moderator, a new population. Re-run the writing study exactly, but add a third group that writes self-critically, and now you replicate the original comparison while testing a new one in the same design. Most of what gets called "building on prior work" is replication-plus-extension: the field's normal engine of progress, when it's running honestly.
| Type | What changes | Question it answers | What it builds |
|---|---|---|---|
| Direct | Only the participants | Is the effect real? | Confidence in the fact |
| Conceptual | The operationalizations | Is the theory right? | Construct & external validity |
| Plus-extension | Original kept + something added | Does it hold, and what else is true? | Cumulative knowledge |
Here's the line students most often get backwards: a big sample does not automatically buy external validity. Generalization depends on how participants were selected and on how many settings, populations, and operationalizations an effect has survived — not on raw N. A direct replication with 10,000 people still hasn't left the original context, so it says nothing new about generalizability. Only conceptual replications and extensions — which vary the measures, settings, and samples — tell you whether an effect travels. External validity comes from how, not how many.
The reckoningThe replication crisis, told as the story it was
The crisis didn't start in psychology. In 2005, epidemiologist John Ioannidis published a paper with the most alarming title in modern science: "Why most published research findings are false" (Ioannidis, 2005). His argument was purely statistical: when studies are underpowered, when tested hypotheses are unlikely to begin with, and when researchers have flexibility in how they analyze and report, the math guarantees that a large share of published "discoveries" are false positives. No fraud required — just ordinary incentives operating on ordinary flexibility. Medicine mostly shrugged. Psychology filed it away.
Then psychology got its wake-up call from an impeccable source. In 2011, Daryl Bem — a respected social psychologist — published nine experiments in the field's flagship journal claiming evidence for precognition: participants' responses were apparently influenced by events that hadn't happened yet, with eight of nine studies reaching statistical significance across more than 1,000 participants (Bem, 2011). The paper followed the field's standard methods and passed peer review at its top journal. That was precisely the problem. If business-as-usual methodology could "demonstrate" that the future reaches backward in time, then business-as-usual methodology could demonstrate anything — and the field had to figure out how.
The answer arrived within months. Simmons et al. (2011) named the mechanism researcher degrees of freedom: all the small, defensible-looking choices an analyst can make after seeing the data — when to stop collecting, which outcome to report, which participants to exclude, which covariate to add. Their simulations showed that combining just four such flexibilities pushed the false-positive rate from the advertised 5% to roughly 61%. Then they proved it with a real experiment, "finding" (p < .05) that listening to the Beatles' "When I'm Sixty-Four" made undergraduates nearly a year and a half younger — chronologically. An impossible conclusion, manufactured entirely with standard tools, undisclosed flexibility, and a straight face.
Was this rare? John et al. (2012) surveyed over 2,000 research psychologists with a truth-telling incentive and found that these questionable research practices (QRPs) were the norm, not the exception: roughly two thirds admitted failing to report all of a study's dependent measures, and in the incentivized condition about nine in ten admitted at least one QRP. The same year the field got its fraud scandal too: Diederik Stapel, a prominent Dutch social psychologist, was exposed by three junior colleagues for outright fabricating data — 58 papers eventually retracted, including a press-released "finding" that thinking about meat makes people more selfish and antisocial, from a study that had never been run at all. Stapel was fraud, and fraud is rare. But the field's deeper realization was uglier in a way: you didn't need fraud. QRPs alone could fill journals with effects that weren't there.
So psychology did something almost no field had done: it audited itself. The Open Science Collaboration (2015) — 270 researchers — ran high-powered direct replications of 100 studies sampled from three top journals. By the significance criterion, only 36% replicated, and the replication effect sizes averaged about half the originals. The Many Labs projects then sharpened the picture by running each of a set of findings across dozens of labs at once: Many Labs 1 found 10 of 13 classic effects replicated consistently across 36 samples (Klein et al., 2014), while Many Labs 2 — 28 findings, 125 samples, more than 15,000 participants — replicated about half (Klein et al., 2018). Many Labs 2 also quietly demolished the field's favorite excuse. If failed replications were caused by "context sensitivity" — hidden moderators, different populations — effects should have varied a lot from sample to sample. They mostly didn't. Whether a finding replicated depended on the finding, not on where or with whom it was retested.
Don't walk away with "psychology is 36% true." The audited studies came from 2008 — peak pre-reform practices — and replication rates differ sharply by subfield and by how strong the original evidence was. The honest summary: under the old incentives, a large minority of published effects — plausibly more — were false positives or badly inflated, and a single study, however exciting, was never the unit of knowledge. That last clause was always true. The crisis just made it impossible to ignore.
The machineryWhy p-hacking and HARKing manufacture false positives
It's worth seeing the mechanics, because none of this requires villainy. A significance test run exactly once, as planned, has a 5% false-positive rate — one lottery ticket. Every researcher degree of freedom is another ticket drawn from the same lottery. Measure seven outcomes and you have seven chances for one to hit p < .05 by luck; report only the winner and the published record shows a clean "predicted" effect that was, in truth, one lucky draw out of seven. Peek at the data and collect more only if you're "almost significant," try the analysis with and without outliers, with and without a covariate — each fork multiplies the tickets while the paper still advertises a 5% error rate. That's p-hacking: not fabricating data, just quietly buying extra lottery tickets and reporting the win as if it were the only ticket you ever held.
HARKing — hypothesizing after the results are known (Kerr, 1998) — is the same trick applied to theory. You run an exploratory study, notice an unexpected pattern, then write the paper as though you predicted that pattern all along. The damage is subtle and severe: a confirmed a priori prediction is strong evidence precisely because the hypothesis took a risk — it could have failed. A hypothesis reverse-engineered from the data took no risk and can't fail, so "confirming" it confirms nothing; the p-value's meaning quietly evaporates. Exploration is legitimate and necessary science. The sin is relabeling it as confirmation. And crucially, researchers who did these things mostly weren't cheating on purpose — they were reasoning motivatedly inside a system that paid for novel, significant, tidy results and paid nothing for honest nulls.
Fraud and QRPs differ in kind. Fabrication — Stapel inventing datasets at his kitchen table — is rare, criminal in spirit, and was caught by junior colleagues doing exactly what a healthy scientific culture trains people to do. QRPs are the honest-person's failure mode: defensible-looking choices, made one at a time, by researchers who believed their own hypotheses. That's why the fixes below are structural rather than moral. You don't solve a lottery-ticket problem with sincerity; you solve it by making people commit to their ticket before the drawing.
The missing nullsPublication bias and the file drawer
Underneath all of it sat a quieter distortion. Journals overwhelmingly published novel, statistically significant results — so studies that found nothing went unpublished and unseen. Rosenthal (1979) named it the file drawer problem and sketched the nightmare version: in principle, journals could be filled with the 5% of studies that are Type I errors while the file drawers back in the labs hold the 95% that found nothing. Reality was never quite that grim, but the direction was right, and the consequences compound. Publication bias doesn't just hide information; it selects for inflated effects (small studies only get published when they luck into a big estimate), punishes the researchers who test rather than confirm, and poisons the well for every literature review that follows. A field that only prints its hits cannot know what it knows.
The fixesHow psychology rebuilt its norms
The reforms map onto the problems one to one. Preregistration attacks flexibility: you post your hypotheses, design, sample size, exclusion rules, and analysis plan in a public, timestamped registry — typically the Open Science Framework — before collecting data (Nosek et al., 2018). The point isn't bureaucracy; it's restoring the distinction that p-values depend on, between prediction and postdiction. You can still explore — you just have to say that's what you're doing. A Registered Report goes further and attacks publication bias at the root: the introduction and method are peer-reviewed and provisionally accepted before any data exist, and the journal commits to publishing the results no matter how they turn out (Chambers, 2013). Launched at the journal Cortex in 2013, the format has spread to over 300 journals — and it's the strongest structural defense against the file drawer ever devised, because a null result can no longer be rejected for being a null result.
Open data and materials attack unverifiability: when your data, code, and stimuli are public, errors get caught, analyses get re-run, and other labs can replicate without guessing what you did. Large samples and a priori power analysis attack the small-study noise that made single labs so unreliable — including big multi-lab collaborations that treat the study, not the lab, as the unit. And journal policy itself changed: the TOP guidelines (Transparency and Openness Promotion) gave journals a graded standard for requiring disclosure, data sharing, and preregistration (Nosek et al., 2015), and many journals now award open-science badges, publish replication attempts, and require all measures and exclusions to be reported — closing off the many-DVs trick by rule.
| Problem | Fix | How it works |
|---|---|---|
| p-hacking / researcher degrees of freedom | Preregistration | Analysis choices are locked before the data can tempt you |
| HARKing | Preregistration | The timestamp separates prediction from postdiction |
| File drawer / publication bias | Registered Reports | Acceptance before results — nulls get published |
| Unverifiable claims | Open data & materials | Anyone can check, re-run, and reuse |
| Noisy single-lab estimates | Power analysis & multi-lab studies | Sample size set in advance; labs pooled |
| Weak journal norms | TOP guidelines & badges | Transparency becomes policy, not virtue |
First: preregistration vs. Registered Report. Preregistration is you timestamping your plan — it kills p-hacking and HARKing, but the journal can still reject your finished paper for being boring, so publication bias survives. A Registered Report adds the journal's advance commitment, which is why it's the stronger fix. Second: reproducibility vs. replicability. Reproducibility means re-running the same analysis on the same data and getting the same numbers — a computational check that open data makes routine. Replicability means collecting new data and finding the same result — a scientific check. A study can be perfectly reproducible and still fail to replicate; the spreadsheet was fine, the effect wasn't.
The showcaseA 26-lab replication your professor helped run
Here's what the new era looks like in practice — and it's homegrown. A few years ago, Professor Magee and a student, Alia Rohan, joined a Registered Replication Report targeting one of social cognition's foundational findings: hostile priming. In the original study, Srull and Wyer (1979) had participants unscramble sentences laced with hostility-related words (punch, fight, argue), then read about "Donald," a man whose behavior is carefully ambiguous. Primed participants rated Donald about three scale points more hostile — an enormous effect, taught to psychology students for four decades.
The replication project rebuilt the paradigm carefully — fresh stimulus materials, independently rated for hostility, and a protocol peer-reviewed and accepted before data collection, exactly as the Registered Report format requires. Then 26 labs tested over 7,300 participants (McCarthy et al., 2018). The result: essentially nothing. In the primary analyses across 22 labs and more than 5,600 people, hostile-primed participants rated Donald a mere 0.08 scale points more hostile than neutral-primed participants — against the original's three. And because the study was pre-accepted, that null went straight into the published record instead of a file drawer.
Look at the spread across the 26 labs and everything in this course snaps into focus. If any one of those labs had run the study alone, it might have landed at the high end and "found" hostile priming — or at the low end and confidently found nothing. Either single lab would have published a wrong-ish conclusion. Only many high-powered, preregistered attempts, pooled, reveal that there is no reliable effect to find. Ask of every finding you ever encounter: How many labs? How many operationalizations? Was the analysis locked before the data arrived? Where are the nulls? A finding that survives those questions deserves your belief. One that has never faced them is a hypothesis wearing a fact's clothing.
Weighing everythingMeta-analysis — cumulative evidence, with a blind spot
If no single study is the unit of knowledge, what is? A literature — and the tool for weighing one is meta-analysis: a quantitative synthesis that collects every study testing the same relationship and computes a weighted average effect size, with bigger, more precise studies counting more. Done well, meta-analysis is the closest thing science has to a verdict: it tells you the best current estimate of an effect and how much it varies across contexts. But it has a built-in weakness you can now name on sight — it can only average the studies it can find. If the file drawer ate the nulls, the meta-analysis averages a biased sample and overestimates the effect, with false precision. Meta-analysts fight back by hunting unpublished studies and using diagnostic tools (a funnel plot of effect sizes against study precision goes tellingly lopsided when small null studies are missing), but the real cure is upstream: preregistration and Registered Reports get the nulls onto the record where a meta-analysis can see them. Garbage in, garbage out — and open science is how the field stopped manufacturing garbage.
GeneralizationWEIRD samples and the other crisis
Suppose an effect replicates beautifully. There's still a second question: who is it true of? Henrich et al. (2010) put a number on psychology's embarrassment: 96% of participants in top-journal studies came from Western industrialized countries — 68% from the United States alone — places holding about 12% of humanity. They named the sample WEIRD: Western, Educated, Industrialized, Rich, Democratic. The deeper finding wasn't the percentage; it was that on measure after measure — visual illusions, fairness in economic games, styles of reasoning, even how the self is construed — WEIRD people are frequent outliers, among the least representative populations you could pick for generalizing about the species. A finding built on American undergraduates is a finding about American undergraduates until shown otherwise.
Yarkoni (2022) pushed the point further and made it uncomfortable for everyone: the generalizability crisis. Our verbal conclusions routinely outrun our evidence — a study using one vignette, one word list, one lab task quietly licenses a claim about "moral judgment" or "priming" in general, as though the particular stimuli were interchangeable with all possible stimuli. Statistically, they aren't: the specific stimuli, like the specific participants, are a narrow sample from a huge space, and effects often shrink or vanish when that space is sampled properly. The remedy is partly technical (treat stimuli as sampled, vary them) and partly rhetorical honesty: say what you actually studied. "Ten minutes of self-compassionate writing reduced rumination in U.S. undergraduates after a lab failure task" is a smaller sentence than "self-compassion heals" — and it has the advantage of being what the data showed.
RealismEcological validity, field experiments, and why the lab isn't "fake"
The last generalization question is about settings. A field experiment runs the full experimental logic — manipulation, random assignment — out in the world, and earns instant generalization to that world along with all the noise and lost control that come with it. But don't conclude that lab research is automatically artificial. The key distinction is between experimental realism — does the situation psychologically grip participants, producing real emotion, motivation, and behavior? — and mundane realism, or ecological validity — does it superficially resemble daily life? A rigged, infuriating lab game looks nothing like anyone's Tuesday, yet the anger it produces is entirely real, so it's a legitimate test of a theory about anger. What a causal theory needs is the genuine engagement of the process it describes, not set dressing. The mature posture is triangulation: tight lab experiments to isolate the mechanism, field experiments and diverse samples to prove it survives contact with the world.
The endingAn honest, optimistic close
It would be easy to end a methods course cynical, and it would be a misreading of everything you just learned. Notice what actually happened: a field discovered — using its own tools — that its literature was contaminated; published the damning audit of itself in the world's most visible journal; diagnosed the mechanism precisely; and then rebuilt its incentives, in public, in about a decade. Fraudsters were caught by their own students. The excuses were tested and, where they failed, retired. That is not a scandal; that is science working — self-correction running exactly as advertised, just slower and more embarrassingly than anyone would like. Fields that never have a replication crisis aren't cleaner. They're just not checking.
And the reforms aren't aspirational anymore — they're simply how research is done now. The live studies hosted on this site run under the norms this lesson describes: preregistered on the OSF before launch, sample sizes set by a priori power analysis, data posted publicly for anyone to check. You're entering psychology after the audit, inheriting the rebuilt version — better norms, better tools, and a literature that is finally learning to say "we were wrong" out loud. Your job, as the person who now knows how to interrogate a claim, is to keep it that way. Ask how they know. Ask who they studied. Ask where the nulls are. And when your own study someday finds nothing — publish it.
Direct, conceptual, or replication-plus-extension? Tap a study to reveal which type it is and why.
SourcesCited in APA 7
Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524
Chambers, C. D. (2013). Registered Reports: A new publishing initiative at Cortex [Editorial]. Cortex, 49(3), 609–610. https://doi.org/10.1016/j.cortex.2012.12.016
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. https://doi.org/10.1017/S0140525X0999152X
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), Article e124. https://doi.org/10.1371/journal.pmed.0020124
John, L. K., Loewenstein, G., & Prelec, D. (2012). Measuring the prevalence of questionable research practices with incentives for truth telling. Psychological Science, 23(5), 524–532. https://doi.org/10.1177/0956797611430953
Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196–217. https://doi.org/10.1207/s15327957pspr0203_4
Klein, R. A., Ratliff, K. A., Vianello, M., Adams, R. B., Jr., Bahník, Š., Bernstein, M. J., Bocian, K., Brandt, M. J., Brooks, B., Brumbaugh, C. C., Cemalcilar, Z., Chandler, J., Cheong, W., Davis, W. E., Devos, T., Eisner, M., Frankowska, N., Furrow, D., Galliani, E. M., . . . Nosek, B. A. (2014). Investigating variation in replicability: A "Many Labs" replication project. Social Psychology, 45(3), 142–152. https://doi.org/10.1027/1864-9335/a000178
Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Jr., Alper, S., Aveyard, M., Axt, J. R., Babalola, M. T., Bahník, Š., Batra, R., Berkics, M., Bernstein, M. J., Berry, D. R., Bialobrzeska, O., Binan, E. D., Bocian, K., Brandt, M. J., Busching, R., . . . Nosek, B. A. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443–490. https://doi.org/10.1177/2515245918810225
McCarthy, R. J., Skowronski, J. J., Verschuere, B., Meijer, E. H., Jim, A., Hoogesteyn, K., Orthey, R., Acar, O. A., Aczel, B., Bakos, B. E., Barbosa, F., Baskin, E., Bègue, L., Ben-Shakhar, G., Birt, A. R., Blatz, L., Charman, S. D., Claesen, A., Clay, S. L., . . . Yıldız, E. (2018). Registered Replication Report on Srull and Wyer (1979). Advances in Methods and Practices in Psychological Science, 1(3), 321–336. https://doi.org/10.1177/2515245918777487
Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D., Breckler, S. J., Buck, S., Chambers, C. D., Chin, G., Christensen, G., Contestabile, M., Dafoe, A., Eich, E., Freese, J., Glennerster, R., Goroff, D., Green, D. P., Hesse, B., Humphreys, M., . . . Yarkoni, T. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425. https://doi.org/10.1126/science.aab2374
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), Article aac4716. https://doi.org/10.1126/science.aac4716
Rosenthal, R. (1979). The "file drawer problem" and tolerance for null results. Psychological Bulletin, 86(3), 638–641. https://doi.org/10.1037/0033-2909.86.3.638
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. https://doi.org/10.1177/0956797611417632
Srull, T. K., & Wyer, R. S. (1979). The role of category accessibility in the interpretation of information about persons: Some determinants and implications. Journal of Personality and Social Psychology, 37(10), 1660–1672. https://doi.org/10.1037/0022-3514.37.10.1660
Yarkoni, T. (2022). The generalizability crisis. Behavioral and Brain Sciences, 45, Article e1. https://doi.org/10.1017/S0140525X20001685