You keep landing on this page from other topics — a little asterisk on Zimbardo, a caution flag on social priming, a "check this before you cite it" note on terror management. So let's put the whole thing in one place and get it right. The replication crisis is not the story of a field that turned out to be fake. It's the story of a field that built better tools to catch its own mistakes — and then had the nerve to use them on its own greatest hits. That's not failure. That's the immune system working.
Bodhi says
The one sentence to leave with: a single dramatic study is a hypothesis, not a fact. Facts in science are things that keep happening when different people, in different labs, run the test again. A finding you saw once — even a famous one — is a promising lead, not a settled truth. The whole crisis is really the field learning to take that sentence seriously.
First: what is a "replication," and why does it matter?
A replication is when someone runs a study again to see if the result shows up a second time. The most demanding kind is a direct (or "close") replication: you copy the original method as faithfully as you can — same procedure, same measures — and just collect fresh data. If the effect is real and general, it should reappear. If it only shows up in the original lab and nowhere else, that tells you something too.
Why insist on this? Because any single study can get a "positive" result by sheer luck. Run enough coin-flips and eventually you'll get ten heads in a row that means nothing. Direct replication is how science separates the real signal from the lucky fluke. A conceptual replication (testing the same idea with a different method) is valuable for showing an effect generalizes, but it's a weaker safety check — if it fails, you can always argue the method was too different. Direct replication removes that escape hatch, which is exactly why the field had quietly avoided doing much of it for decades.
2011: the year the alarm went off
Three things happened in close succession that made the problem impossible to ignore.
First, a respected researcher published a rigorous-looking paper claiming evidence for precognition — that people could, in effect, feel the future (Bem, 2011). It appeared in a top journal, used standard methods, and reached statistical significance. And that was the point that landed like a slap: if our normal, accepted methods can produce a publishable "proof" of ESP, then something is wrong with the methods, not with physics. Bem's paper became an accidental gift — a reductio ad absurdum the whole field could see.
Second, Simmons, Nelson, and Simonsohn (2011) published "False-Positive Psychology," and it named the machinery. They showed — with real experiments and simulations — that ordinary, well-intentioned analytic choices could push the odds of a false positive from the assumed 5% to over 60%. The culprit was researcher degrees of freedom: all the small decisions a researcher makes along the way (which participants to exclude, when to stop collecting data, which of several measures to report, which covariates to include). Make those choices after peeking at the data, chasing whatever looks significant, and you're p-hacking — often without feeling like you're cheating at all. In their most memorable demonstration, they "proved" that listening to a Beatles song literally made people younger. It's a joke result, produced by entirely common practices.
Third — and this is the big one — the Open Science Collaboration (2015) did the obvious thing nobody had done at scale: it directly replicated 100 published psychology studies with large samples and preregistered plans. The result was sobering. Depending on how you counted, only about 36–39% of the studies produced a statistically significant effect the second time, and the replication effect sizes were on average roughly half the size of the originals. Not zero. Not fraud. But a lot smaller and shakier than the published record suggested.
Read the number correctly
"39% replicated" does not mean "61% of psychology is false." A failed replication can happen because the original was a fluke — or because the effect is real but smaller than reported, or fragile, or depends on conditions nobody wrote down. The honest takeaway is humbler and more useful: published significance is weaker evidence than we treated it as. One study earns a raised eyebrow, not a citation in your worldview.
Why it happened — and why "fraud" is the wrong word
Here's the part students often get backwards. The replication crisis was overwhelmingly not about lying. Outright fabrication happens and is rightly a scandal when caught, but it's rare and it's not what these numbers are about. The crisis was mostly produced by an ecosystem of ordinary practices and bad incentives that quietly manufactured false positives. Five interlocking causes:
- Publication bias / the file drawer. Journals published exciting, positive, significant results and rejected null results as "boring." So studies that found nothing got stuffed in a file drawer and vanished. The published literature became a highlight reel, not a full record — and a highlight reel of coin-flips looks a lot like a superpower.
- Small samples, low power. Many classic studies ran on tiny numbers of people. Small samples don't just miss real effects; when they do hit significance, they tend to overestimate the effect size. A field addicted to small, cheap studies was practically designed to publish inflated effects.
- P-hacking. The researcher-degrees-of-freedom problem above — trying analyses until one crosses p < .05, then reporting only that one (Simmons et al., 2011). Usually done sincerely, by people who thought they were just "cleaning the data."
- HARKing. "Hypothesizing After the Results are Known" (Kerr, 1998): running an exploratory study, finding an unexpected pattern, then writing the paper as if you'd predicted it all along. It converts a lucky accident into a fake confirmed prediction — and hides how much the result was really a surprise.
- Incentives. Underneath it all: careers were built on novel, positive, surprising results. "We replicated an old finding" or "we found nothing" didn't get you hired, funded, or into the good journals. When you reward flashy over solid, you get flashy — and flashy is exactly what fails to replicate.
Notice how gentle that list is toward the people involved. Almost none of it requires a villain. That's the genuinely unsettling insight — and, if you squint, a very social-psychological one: put well-meaning people in a system that rewards the wrong thing, and the situation produces the behavior. (Sound familiar? It's the power of the situation, aimed back at scientists themselves.)
Bodhi says
Don't let anyone talk you into cynicism here. "See, psychology is junk!" is the lazy read — and it's wrong, because the people who discovered these problems were psychologists, using better psychology. A field that can't find its own mistakes is scary. A field that goes looking for them, out loud, is exactly the kind you should trust more.
The fixes — this is the hopeful part
What makes this a success story rather than a scandal is what came next. Within a few years the field reengineered how it works. If you understand these reforms, you understand why post-2015 psychology is genuinely more trustworthy than what came before:
How the field rebuilt itself
Preregistration. Before collecting data, you publicly timestamp your hypothesis, your sample size, and your exact analysis plan. Now there's a paper trail — you can't quietly p-hack or HARK, because everyone can see what you said you'd do. It draws a bright line between confirmatory tests (planned) and exploratory ones (fishing), and both are fine as long as you label them honestly.
Registered Reports. The clever one. Journals review and accept your study based on the question and method — before you have results. Publication no longer depends on the results coming out "significant," which kills publication bias at the root and removes the incentive to p-hack. Null results get published just as readily.
Bigger samples and real power analysis. Studies got dramatically larger, so results are less at the mercy of luck and effect sizes are estimated more honestly.
Open data and open materials. Post your data, code, and stimuli so anyone can check your work — or reuse it. Sunlight as disinfectant.
Many Labs projects & meta-analysis. Instead of one lab's word, dozens of labs run the same study simultaneously (the "Many Labs" collaborations), and meta-analysis pools everything to estimate the truth across all of it. Trust migrates from any single result to the accumulated record.
You'll notice these aren't cosmetic. They attack the exact causes from the last section, one for one. That's what a maturing science looks like: not "we were right all along," but "we found the leak and we fixed the plumbing."
The scorecard: what held up, what shrank
So which of the greatest hits survived the audit? This is where honesty matters most, in both directions. Plenty of psychology came through just fine — and pretending everything collapsed is as false as pretending nothing did. Read this fairly: several entries on the right are contested, not "debunked," and their original authors often dispute the failed replications. The point isn't a demolition. It's calibration.
✓ Held up well
The Stroop effect. Say the ink color, not the word — one of the most bulletproof findings in all of psychology. Replicates endlessly.
Asch conformity. The line-judgment conformity effect is robust and well-replicated across decades and cultures (meta-analyzed by Bond & Smith, 1996), though its size varies with the situation.
The contact hypothesis. That well-structured intergroup contact reduces prejudice held up powerfully — a meta-analysis of 500+ studies confirmed it (Pettigrew & Tropp, 2006).
Cognitive dissonance (core effect). The basic phenomenon — inconsistency creates discomfort that drives attitude change — remains solid, even as details get refined.
Careful self-report attitude measurement. Well-built survey and attitude measures have generally proven reliable and reproducible.
✗ Shrank, contested, or fell
Behavioral social priming. The famous claim that reading "elderly"-related words made people walk more slowly (Bargh) failed a careful direct replication (Doyen et al., 2012). This behavioral priming literature is among the hardest-hit.
Ego depletion. The idea that willpower is a limited fuel that runs down. A large multi-lab Registered Replication found the effect near zero (Hagger et al., 2016) — still debated, but far weaker than believed.
Facial feedback. That holding a pen in your teeth (forcing a smile) makes things funnier failed a 17-lab replication (Wagenmakers et al., 2016). Later work suggests a small effect under some conditions — contested, not settled.
Terror Management (basic mortality-salience effect). The core "remind people of death, watch worldview defense" effect did not replicate in Many Labs 4 (Klein et al., 2022). Proponents dispute the methods; treat it as an open question. See Terror Management.
Many stereotype-threat replications. The signature effect appears real in some settings but has proven smaller and more fragile than the early literature implied.
The Stanford Prison "Experiment." Archival work shows guards were coached toward cruelty and it lacked real experimental controls (Le Texier, 2019) — best read as a demonstration, not an experiment. See History.
And one special case that isn't really a replication failure at all, but a storytelling failure — worth its own paragraph because it teaches a slightly different lesson.
The "38 witnesses" that never were
You've heard the Kitty Genovese story: 38 people supposedly watched a 1964 murder from their windows and did nothing, and it launched the bystander-effect research. Careful reexamination showed the "38 witnesses who saw and ignored it" narrative was largely a journalistic exaggeration — the number and the passivity were overstated in the original news coverage (Manning, Levine, & Collins, 2007). Crucially, this does not debunk the bystander effect itself: Latané and Darley's actual experiments on diffusion of responsibility are solid and replicate. The lesson is subtler — a shaky anecdote can sit on top of perfectly good science. Keep the two separate. See Prosocial Behavior.
The meta-lesson: skepticism, not cynicism
Here's the payoff, and it's the same distinction you met in Research Methods. A cynic hears "39% replicated" and throws the whole field in the trash. A skeptic hears the same number and updates: single studies are weak evidence, accumulation is strong evidence, and I should ask "has this replicated?" before I believe it. Cynicism is lazy and, ironically, just as uncritical as gullibility — it's a way of never having to weigh evidence at all. Skepticism is the actual skill worth building.
So when a topic page on this site sends you here with an asterisk — on social priming, terror management, the Genovese story, prejudice research (stereotype threat), or the fast-and-frugal side of heuristics — that asterisk isn't there to make you distrust psychology. It's there to make you good at it: to hold findings at the confidence they've earned. A dramatic result is a hypothesis. Trust is something a finding accumulates, one replication at a time.
Carry this into every study you meet
Ask three questions of any finding: How big was the sample? Has it been directly replicated? Is there a meta-analysis? If the answers are "small, no, and no," treat the result as an interesting lead — worth remembering, not worth betting on. That single habit will make you a sharper reader of psychology than most people who've never heard the word "replication."
The one thing to carry out of this page
The replication crisis is the best evidence that psychology is a real science — precisely because a real science is the only kind that would run this audit on itself, publish the embarrassing numbers, and then rebuild its own methods in response. You should leave here trusting good science more, not less: the preregistered, well-powered, replicated, meta-analyzed findings that survived this reckoning are on firmer ground than almost anything in the field's first century. Skepticism isn't the enemy of trust. Done right, it's how trust gets earned.
References
Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425. https://doi.org/10.1037/a0021524
Bond, R., & Smith, P. B. (1996). Culture and conformity: A meta-analysis of studies using Asch's (1952b, 1956) line judgment task. Psychological Bulletin, 119(1), 111–137. https://doi.org/10.1037/0033-2909.119.1.111
Doyen, S., Klein, O., Pichon, C.-L., & Cleeremans, A. (2012). Behavioral priming: It's all in the mind, but whose mind? PLoS ONE, 7(1), e29081. https://doi.org/10.1371/journal.pone.0029081
Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., Anggono, C. O., Batailler, C., Birt, A. R., Brand, R., Brandt, M. J., Brewer, G., Bruyneel, S., Calvillo, D. P., Campbell, W. K., Cannon, P. R., Carlucci, M., Carruth, N. P., Cheung, T., Crowell, A., De Ridder, D. T. D., Dewitte, S., … Zwienenberg, M. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science, 11(4), 546–573. https://doi.org/10.1177/1745691616652873
Kerr, N. L. (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review, 2(3), 196–217. https://doi.org/10.1207/s15327957pspr0203_4
Klein, R. A., Cook, C. L., Ebersole, C. R., Vitiello, C., Nosek, B. A., Hilgard, J., Ahn, P. H., Brady, A. J., Chartier, C. R., Christopherson, C. D., Clay, S., Collisson, B., Crawford, J. T., Cromar, R., Gardiner, G., Gosnell, C. L., Grahe, J., Hall, C., Howard, I., … Ratliff, K. A. (2022). Many Labs 4: Failure to replicate mortality salience effect with and without original author involvement. Collabra: Psychology, 8(1), 35271. https://doi.org/10.1525/collabra.35271
Le Texier, T. (2019). Debunking the Stanford Prison Experiment. American Psychologist, 74(7), 823–839. https://doi.org/10.1037/amp0000401
Manning, R., Levine, M., & Collins, A. (2007). The Kitty Genovese murder and the social psychology of helping: The parable of the 38 witnesses. American Psychologist, 62(6), 555–562. https://doi.org/10.1037/0003-066X.62.6.555
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Pettigrew, T. F., & Tropp, L. R. (2006). A meta-analytic test of intergroup contact theory. Journal of Personality and Social Psychology, 90(5), 751–783. https://doi.org/10.1037/0022-3514.90.5.751
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. https://doi.org/10.1177/0956797611417632
Wagenmakers, E.-J., Beek, T., Dijkhoff, L., Gronau, Q. F., Acosta, A., Adams, R. B., Jr., Albohn, D. N., Allard, E. S., Benning, S. D., Blouin-Hudon, E.-M., Bulnes, L. C., Caldwell, T. L., Calin-Jageman, R. J., Capaldi, C. A., Carfagno, N. S., Chasten, K. T., Cleeremans, A., Connell, L., DeCicco, J. M., … Zwaan, R. A. (2016). Registered replication report: Strack, Martin, & Stepper (1988). Perspectives on Psychological Science, 11(6), 917–928. https://doi.org/10.1177/1745691616674458