Foundations · Unit 02

Research Methods

How psychology earns the word "science" — and how to tell what a study can and can't actually claim.

~17 min · pairs with the Research Methods lecture

Here's the single most useful skill you'll take from this course, more useful than any fact you'll memorize: reading a study and knowing exactly what it's entitled to claim. Every week you'll see headlines built on psychological research — "Study finds X causes Y," "New research links A to B" — and most of them are quietly overstating what the underlying design can support. This unit gives you the tools to catch that.

The horse that could do math

In the early 1900s, a horse named Clever Hans became a sensation in Germany. His owner, a retired schoolteacher, claimed Hans could solve arithmetic problems, tell time, and even spell — tapping out answers with his hoof. Investigators tested him rigorously and found no trickery: Hans got it right even with strangers asking the questions, even without his owner in the room. It looked like real equine intelligence.

Then the psychologist Oskar Pfungst ran a cleverer test. He had questioners who didn't know the answer themselves ask Hans the questions. Hans's accuracy collapsed (Pfungst, 1911). It turned out Hans was reading tiny, involuntary cues in his questioner's posture and facial tension — cues that appeared the moment the person expected the correct hoof-tap and relaxed the instant it happened. Hans wasn't doing math. He was doing something arguably more impressive: reading people. But he wasn't answering the question anyone thought they were asking him.

Historic photograph of Clever Hans, a horse, facing his owner Wilhelm von Osten beside a number board in a study room.
Clever Hans with his owner, Wilhelm von Osten, and the number board he tapped answers on — before Pfungst traced his accuracy to unconscious cues from his questioners. Public domain, via Wikimedia Commons.

◆ The big idea

Without controls and blinding, you often end up measuring the experimenter, not the subject. Clever Hans is the origin story for experimenter-expectancy effects and for why researchers now build blind procedures into almost everything — so that neither the participant nor (ideally) the person running the session knows which condition is which, and can't unconsciously nudge the result.

Theories, hypotheses, and the loop between them

Clever Hans is a story about a claim that outran its evidence, so let's be precise about what evidence-based claims look like. Start with the distinction between a fact and an opinion: a fact is an observable reality that can be checked, while an opinion is a personal judgment that may or may not be accurate. Science is the machinery for converting the second into the first — or discarding it. That machinery runs on two kinds of idea. A theory is a well-developed set of ideas that explains a body of observations and has held up against repeated testing; it is the best current account of some slice of the world, not a guess. A hypothesis is a specific, testable prediction derived from a theory, often phrased as an if–then statement: if the James-Lange theory is right that emotion is built from bodily arousal, then a person who cannot feel that arousal should feel less fear. Theories are too big to test all at once; hypotheses are the bite-sized pieces you can actually put on the table.

The pieces connect in a loop. From a theory you derive a hypothesis, collect data, and either confirm the theory or modify it — and a modified theory generates new hypotheses, so the cycle never really closes. Two directions of reasoning power it. Deductive reasoning runs from the general to the specific: start with a hypothesis and work out what must be observed if it's true, then look. Inductive reasoning runs the other way: pile up observations and build a generalization from them. Induction is how case studies and naturalistic observation generate new theories; deduction is how experiments test them. Neither is safe on its own — a deductive argument from a false premise yields a perfectly logical wrong answer, and an inductive leap from apples, bananas, and oranges to "all fruit grows on trees" is demolished by the first strawberry.

One property separates a scientific hypothesis from a merely interesting one: it must be falsifiable — capable, at least in principle, of being shown wrong. This is the philosopher Karl Popper's test (Popper, 1959), and it explains why Freud shows up in this unit as a cautionary example. There is no observation that could disprove the existence of the id, ego, and superego, so however rich those ideas are, they can't be tested, which means they can't be corrected. The James-Lange hypothesis above, by contrast, can be — and was: people with spinal-cord injuries that block feedback from the body do still feel emotions, if somewhat less intensely, which forced the theory to be refined rather than accepted whole. A hypothesis that risks nothing tells you nothing.

Where you study people: lab versus field

Every study happens somewhere, and where it happens trades off two things you can't fully have at once. In a lab setting, you control the environment tightly — lighting, noise, timing, exact wording of instructions — which lets you isolate the variable you care about. That buys you internal validity: confidence that what you manipulated is actually what produced the effect. In a field setting, you study people in their everyday environments — a real classroom, a real workplace — which buys you external validity: confidence that the finding generalizes to real life, because it was observed in real life.

Neither setting is "better." A tightly controlled lab study of aggression using a contrived paradigm might nail down a causal mechanism beautifully but leave you unsure whether it plays out the same way at a bar on a Friday night. A field study of the bar might have terrific real-world relevance but a dozen uncontrolled variables tangled together. This is the internal-versus-external validity tradeoff, and it shows up constantly when you evaluate a study's setting.

Jane Goodall crouching in a forest, extending her hand toward a young wild chimpanzee that reaches back to touch it.
Field research at its purest: Jane Goodall observing wild chimpanzees at Gombe. Naturalistic observation buys external validity — behavior seen in the real world — at the cost of experimental control. Photo: Hugo van Lawick / National Geographic; from lecture slides.

Bodhi says

When you read a methods section, ask "where did this happen?" before you ask "what did they find?" A lab finding and a field finding earn different kinds of trust — neither is automatically the stronger one.

The methods ladder

Psychological methods form a rough ladder, and each rung answers a fundamentally different question. Climbing the ladder buys you more certainty about cause and effect, but it costs you something too — usually realism, ethical feasibility, or both. Work through the three rungs below.

Step 1 · Descriptive / observational — "What happens?"

This is the ground floor: case studies, naturalistic observation, surveys. You watch or ask, and you describe what you see without manipulating anything. It's how you generate hypotheses and catch phenomena you didn't know to look for. What it can't do is tell you why something happens or what would happen under different conditions — there's no comparison built in.

Step 2 · Correlational — "What's related to what?"

Now you're measuring two or more variables and asking whether they move together — does more of one tend to come with more (or less) of the other? Correlational designs let you predict and let you study things you could never ethically manipulate (you can't randomly assign people to have anxiety, but you can measure whether anxiety correlates with sleep quality). The cost: a correlation, no matter how strong, never tells you which variable caused which — or whether something else caused both. More on that below.

Step 3 · Experimental — "What causes what?"

Only the true experiment lets you make causal claims, because only here does the researcher actively manipulate a variable and randomly assign participants to conditions. This is the top of the ladder — the most powerful design for isolating cause and effect — and also the most demanding: it requires a manipulable variable, an ethical way to manipulate it, and often a lab-like level of control that can cost you realism.

Experimental What causes what? Correlational What goes with what? Descriptive What happens? causal certainty
Each rung of the methods ladder answers a different question. Climbing it buys causal certainty — but usually costs you realism, ethical feasibility, or both.

The ground floor has more rooms than the three named above, and each trades depth against reach in its own way. A case study goes deep on one person or a handful — the conjoined Hogan twins, joined at the thalamus and apparently sharing some sensations, have been followed for years precisely because no other method could tell us what their brains can. The richness is unmatched; the catch is that people worth a case study are, by definition, unusual, so generalizing from them to everyone else is hazardous. Naturalistic observation watches behavior in its own setting, and its whole value depends on the observer staying inconspicuous, since people who know they are watched stop behaving naturally (think of your driving when a police car appears behind you). Its cousin, structured observation, sets people a specific task and watches how they handle it — Mary Ainsworth's Strange Situation, in which a caregiver leaves and returns to a toddler in a room of toys, is the classic example (Ainsworth et al., 1978). Both are exposed to observer bias, the tendency for an observer who knows the hypothesis to unconsciously see what they expect; the remedies are clear coding rules set in advance and checking that independent observers agree, which is where inter-rater reliability (below) earns its keep. Archival research skips participants entirely and mines existing records — transcripts, census data, old surveys. It's cheap and fast, but you get only the data someone else decided to collect, in whatever form they collected it.

Surveys and other self-report methods buy breadth: hundreds or thousands of people, quickly, which is exactly what you want when the goal is generalizing to a population. What they lose is depth, and they inherit a weakness that runs through every method where people describe themselves: social desirability bias, the pull to answer in the way that makes you look good. Ask a lecture hall who always washes their hands after using the restroom and every hand goes up; station an unobtrusive observer by the sinks and the number drops. Well-built surveys work around this with indirect questions — one study on post-9/11 attitudes found participants would not admit prejudice toward Arab Americans when asked directly, yet reported markedly less willingness to interact with them than with other groups when asked about specific situations. The lesson isn't that surveys are worthless; it's that the answer to a question is a behavior too, and behaviors have causes beyond the truth.

One important design lives in the gap between the correlational and experimental rungs: the quasi-experiment. Here the researcher compares groups that differ on some variable of interest — smokers versus non-smokers, students at two different schools, people before and after a natural disaster — but without random assignment, because the grouping already exists in the world and can't ethically or practically be assigned. Quasi-experiments are enormously useful for questions you could never study any other way, but they inherit the correlational rung's core weakness: because the groups weren't randomly formed, they may differ in ways beyond the variable you care about, so causal claims stay tentative. A related point worth filing away is that individual studies are not the final word either. When many studies test the same question, researchers can combine their results statistically in a meta-analysis, which pools the evidence to estimate an effect more precisely than any single study can. A well-conducted meta-analysis is often the most trustworthy answer psychology can offer, because it averages out the flukes and biases of individual experiments.

Correlation is not causation

This is the phrase you'll hear repeated so often it risks becoming background noise — so let's make sure you actually understand why it's true, not just that it's true. A correlation tells you two variables move together: as one goes up, the other tends to go up (a positive correlation) or down (a negative correlation), and the strength of that relationship can range from weak to strong. What a correlation cannot tell you, no matter how strong or how statistically reliable, is which variable is doing the causing — or whether either one is causing the other at all.

There are two classic reasons a correlation can mislead you about causation. The first is the directionality problem: if A and B are correlated, maybe A causes B, but it's equally possible that B causes A. The correlation itself can't tell you which way the arrow points. The second is the third-variable problem: some hidden factor, C, might be causing both A and B independently, creating a correlation between them that has nothing to do with either causing the other.

It helps to know how correlation is actually quantified. Researchers summarize the relationship between two variables with a correlation coefficient, symbolized r, that ranges from −1.00 to +1.00. The sign tells you direction: a positive r means the variables rise together, a negative r means one rises as the other falls. The magnitude — how far from zero — tells you strength: an r near ±1.00 is a tight, highly predictable relationship, while an r near 0 means knowing one variable tells you almost nothing about the other. A crucial subtlety hides in that number, though: r measures only the strength of a straight-line relationship. Two variables can be powerfully related in a curved pattern — think of the way arousal helps performance up to a point and then hurts it — and still produce an r close to zero, because the relationship isn't linear. So a small correlation coefficient doesn't always mean "no relationship"; sometimes it means "not a straight-line relationship."

The classic illustration: ice cream sales and drowning deaths rise and fall together across the year, and the correlation is strong. Does buying ice cream cause drowning? Obviously not. The third variable is hot weather — it drives people to both buy ice cream and go swimming, and swimming (not soft-serve) is what raises drowning risk. Whenever a correlation strikes you as surprising, your first move should be to ask what third variable might be driving both sides of it. Related to this: watch out for illusory correlation, where people perceive a relationship between two things that isn't actually there, often because a few vivid or memorable co-occurrences stick in memory while the many times the pattern didn't hold get ignored.

The full moon is the textbook case. Ask emergency-room nurses, police officers, or teachers, and many will swear that people act strangely when the moon is full. A meta-analysis pooling nearly forty studies found no such relationship at all — rates of odd behavior are flat across the lunar cycle (Rotton & Kelly, 1985). The belief survives anyway, and the reason is worth naming: confirmation bias, the tendency to notice and remember evidence that fits a hunch while letting the misses slide past. A wild night under a full moon gets filed as proof; a wild night under a crescent moon is just a wild night. Illusory correlations aren't harmless curiosities, either — the same mechanism, applied to groups of people, is one of the engines of stereotyping, where a few vivid instances get welded to a whole category.

The myth

"A strong correlation is basically proof of causation — if the statistics are solid, you can trust the causal story."

What's actually true

Statistical strength has nothing to do with causal proof. A correlation can be enormous, highly statistically significant, and replicated across ten studies — and still tell you nothing about which variable causes which, because directionality and third variables are never ruled out by the correlation itself. Only a true experiment, with manipulation and random assignment, can rule those out.

Check yourself
Headline: "Cities with more coffee shops per capita have higher rates of anxiety disorders." What kind of design does this headline describe?

The true experiment

An experiment earns its causal claims through a specific structure. The researcher manipulates an independent variable (the thing you change) and measures its effect on a dependent variable (the outcome you're tracking). Participants who receive the manipulation make up the experimental group; participants who don't — or who get a comparison condition — make up the control group.

The engine that makes causal inference possible is random assignment: every participant has an equal chance of ending up in any condition. Random assignment doesn't guarantee the groups are identical, but across a reasonably sized sample it equates them on average on everything you didn't measure — personality, mood, history, genetics, whatever. That's what lets you attribute a difference between groups to the manipulation rather than to some preexisting difference between the people in each group. When groups differ on something other than the intended manipulation, that something is a confound — and a good experimental design exists largely to prevent confounds from creeping in.

Two more pieces close the loop. A placebo is a fake treatment given to the control group so both groups believe they might be receiving the real thing — this controls for the power of expectation itself. In a single-blind study, participants don't know which condition they're in; in a double-blind study, neither participants nor the researchers interacting with them know — which, as Clever Hans taught us, matters enormously, because researchers who know the hypothesis can unconsciously communicate it.

◆ What the slides usually skip

Notice that a placebo control and blinding are solving two different problems. The placebo controls for participants' expectations. Blinding controls for the researcher's expectations leaking into how they treat participants or interpret ambiguous behavior. A well-designed drug trial typically needs both.

Two subtler threats round out the picture, because participants are not passive measuring instruments — they're people who notice they're being studied and form guesses about what the researcher wants. Demand characteristics are cues in a study that tip participants off to the hypothesis, prompting them to helpfully (or contrarily) act it out rather than behave naturally. Closely related is the Hawthorne effect: the tendency for people to change their behavior simply because they know they're being observed, independent of any actual manipulation. And the experimenter's own expectations can do real work, not just at the level of unconscious cues but on outcomes themselves. In one famous demonstration, Rosenthal and Jacobson (1968) told teachers that certain randomly chosen students were poised to "bloom" intellectually; months later those students had in fact gained more on IQ tests than their classmates — the teachers' expectations had quietly shaped how they taught and evaluated the children. That is precisely why modern designs work so hard to keep expectations, on both sides of the table, from contaminating the result.

Sampling versus assignment — the mix-up everyone makes

These two phrases sound alike and get confused constantly, but they answer completely different questions and control completely different things.

Random sampling

The question: who ends up in the study at all? If participants are drawn randomly from a larger population, every member of that population had an equal chance of being selected. This is what lets you generalize findings from your sample back to the broader population. Most psychology studies do not use true random sampling — convenience samples (whoever signs up) are far more common, which limits generalizability regardless of how good the rest of the design is.

Random assignment

The question: once someone is in the study, which condition do they get? Randomly assigning already-recruited participants to experimental versus control conditions is what supports causal inference — it's what rules out confounds by equating the groups on average. You can have random assignment with zero random sampling (a convenience sample randomly split into two groups), and that's exactly what most psychology experiments actually look like.

Check yourself
A researcher recruits whichever introductory psychology students sign up (a convenience sample), then flips a coin to decide who gets the real treatment versus the placebo. What does this design have — and what is it missing?

Operational definitions, reliability, and validity

Before you can measure anything, you need an operational definition — a precise, concrete statement of exactly how a concept will be measured or manipulated in this particular study (not "aggression" in the abstract, but "number of times a participant pressed a button delivering a loud noise to a confederate"). Two more terms travel together and get conflated constantly: reliability is whether a measure gives consistent results — measure the same stable thing twice, get roughly the same answer. Validity is whether a measure actually captures the concept it claims to capture. A bathroom scale that's off by ten pounds every time is reliable (consistent) but not valid (not accurate); a measure has to be reliable to have any chance of being valid, but reliability alone doesn't guarantee it.

Both terms come in several flavors that are worth recognizing. Reliability can be assessed as test–retest reliability (does the same person score about the same on two occasions?), inter-rater reliability (do two independent observers coding the same behavior agree?), and internal consistency (do the items on a scale that supposedly measure one thing actually hang together?). Validity likewise splits apart: construct validity asks whether the measure truly captures the underlying concept; internal validity — which we met with the lab setting — asks whether the study's design supports its causal claim; and external validity asks whether the finding generalizes beyond the specific sample and setting. When you evaluate a study, it's rarely enough to say a measure is "reliable and valid" — the sharper question is which kind, and whether that's the kind the claim actually depends on.

Three more flavors of validity turn up often enough to know by name. Face validity is the weakest: does the measure look like it measures the thing, on the surface? A questionnaire about test anxiety that asks about sweaty palms before exams has face validity; that alone proves little. Ecological validity is external validity's realism component — how closely the study situation resembles real life, which is exactly what naturalistic observation buys and a contrived lab task gives up. Predictive validity asks whether a measure forecasts what it should: the SAT is defended on the grounds that scores predict first-year college GPA, and it's criticized on the grounds that the prediction is weaker than advertised and tilted against students from historically marginalized groups. When you hear that a test is "valid," the useful follow-up is: valid for predicting what, in whom?

Check yourself
A new "test anxiety" questionnaire gives students wildly different scores when they retake it a week later, with no reason to think their actual anxiety changed. What's the problem?

Whose data is this, anyway? The WEIRD problem

Even a flawless experiment inherits the limits of who's in it. Henrich, Heine, and Norenzayan (2010) pointed out that the overwhelming majority of psychology's published findings come from samples that are Western, Educated, Industrialized, Rich, and Democratic — WEIRD, for short. That's not a coincidence; it reflects who researchers have easy access to, largely undergraduates at universities in wealthy Western countries. The trouble is that WEIRD populations are, on many psychological dimensions, statistical outliers relative to humanity as a whole — not a neutral default that everyone else deviates from.

This is a generalizability problem layered on top of everything else you've learned in this unit. A study can have a large sample, random assignment, careful blinding, and still tell you relatively little about how the effect works outside a narrow slice of humanity. When you read "participants," it's worth asking who, specifically, and how far that finding is likely to travel.

Studying change over time

Some questions require tracking how people change, and there are two main ways to design that. A cross-sectional design compares different age groups all at once — measure 10-, 20-, and 40-year-olds today and compare them. It's fast and cheap, but it can't distinguish true developmental change from a cohort effect: differences between age groups that exist because they grew up in different eras with different experiences, not because of aging itself. A longitudinal design follows the same group of people over time, which sidesteps the cohort confound but costs years (sometimes decades), money, and participants who drop out along the way.

What "statistically significant" really means

You'll see the phrase statistically significant attached to almost every result, and it's routinely misunderstood — including, for a long time, by researchers themselves. In the standard framework, a result is called significant when its p-value falls below a threshold, conventionally .05. The p-value is the probability of getting a result at least this extreme if the effect were actually zero — that is, if nothing but chance were operating. A small p-value means the data would be surprising under the "nothing's here" assumption, which gives you some license to doubt that assumption. What "significant" does not mean is that the effect is large, important, or even necessarily real. It is a statement about surprise-under-chance, not about the size or practical meaning of the finding.

That's why researchers increasingly report effect size alongside significance. Effect size asks the question significance ignores: how big is the difference? With a large enough sample, a trivially tiny difference can clear the .05 bar and be labeled "significant," while a genuinely important effect studied in a small sample can miss it. Significance tells you roughly whether an effect is there; effect size tells you whether it matters. Reading both together is a mark of statistical literacy — and it protects you from headlines that dress up a microscopic effect in the language of a breakthrough.

This distinction becomes urgent once you see how much wiggle room researchers have. Simmons, Nelson, and Simonsohn (2011) showed, using both simulations and real experiments, that ordinary and seemingly innocent choices — peeking at the data and deciding whether to collect more, dropping "outliers," trying several outcome measures and reporting the one that worked — can inflate the false-positive rate far above 5% without anyone lying or even feeling dishonest. They named these choices researcher degrees of freedom, and the paper's title made the danger plain: undisclosed flexibility "allows presenting anything as significant." This is the intellectual bridge to the crisis in the next section: if significance can be manufactured through flexible analysis, then a literature full of just-barely-significant findings is exactly the literature you'd expect to have trouble replicating.

Current research: the replication crisis and open science

◆ Where the field is now

Starting in the early 2010s, psychology went through a hard reckoning. The Open Science Collaboration (2015) attempted to directly replicate 100 studies published in top psychology journals — and found that a substantial share of the original effects did not reproduce, or reproduced much more weakly than first reported. That's not a scandal so much as science working as intended: a field checking its own results and finding out where it had gotten ahead of the evidence.

The response has been constructive. Preregistration — publicly committing to your hypotheses, sample size, and analysis plan before collecting data — has moved from a niche practice to something close to a norm in many corners of the field, precisely because it closes off the flexibility that let earlier researchers unconsciously fish for significant results (Nosek, Ebersole, DeHaven, & Mellor, 2018). Data-sharing norms have shifted too, with more journals expecting raw data to be posted publicly so others can check the work. This is authentically what working researchers do now — it isn't a historical footnote, it's the water the field currently swims in, and you'll see it reflected in how modern studies report their methods.

Two older safeguards sit underneath all of this. The first is peer review: before a study appears in a scientific journal, several other researchers with expertise in the topic read it — usually anonymously — and report to the editor on whether the rationale holds, the method is sound, the statistics are right, the ethics were observed, and the conclusions actually follow from the data. The editor then publishes, asks for revisions (the usual outcome), or rejects. Peer review is quality control, not a guarantee of truth; its quieter job is forcing authors to describe their methods clearly enough that someone else could replicate the study, which is the check that ultimately matters. The second safeguard is retraction. When a published paper turns out to rest on fabricated or falsified data, or on a design too flawed to support its claims, the journal can formally withdraw it. The most consequential retraction in recent memory is the 1998 paper that claimed a link between childhood vaccines and autism: it was withdrawn once it emerged that the data had been manipulated and that the lead author had an undisclosed financial stake in the result (The Editors of The Lancet, 2010). Large epidemiological studies have since found no link — but the paper had a decade's head start, and the measles outbreaks that followed the drop in vaccination rates are the cost of a false finding that got loose before the system caught it.

Bodhi says

A field that revises its own famous findings isn't a broken field — it's a field taking its own methods seriously. Keep that distinction in mind before you let "some studies didn't replicate" curdle into "psychology isn't real science."

Check yourself

The one thing to carry out of this unit

Before you believe any claim about what causes what, ask what the design actually was. Was anything manipulated? Was assignment to conditions random? Was the sample randomly drawn, or just whoever was available — and how WEIRD is it? Was the measure operationally defined, reliable, and valid? Get in the habit of asking these questions automatically, and you'll be equipped to evaluate research claims — in this course and for the rest of your life — far more rigorously than most people ever learn to.

References

Ainsworth, M. D. S., Blehar, M. C., Waters, E., & Wall, S. (1978). Patterns of attachment: A psychological study of the strange situation. Lawrence Erlbaum.

The Editors of The Lancet. (2010). Retraction—Ileal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children. The Lancet, 375(9713), 445. https://doi.org/10.1016/S0140-6736(10)60175-4

Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83.

Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.

Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.

Pfungst, O. (1911). Clever Hans (the horse of Mr. von Osten): A contribution to experimental animal and human psychology (C. L. Rahn, Trans.). Henry Holt. (Original work published 1907)

Popper, K. R. (1959). The logic of scientific discovery. Hutchinson. (Original work published 1935)

Rosenthal, R., & Jacobson, L. (1968). Pygmalion in the classroom. The Urban Review, 3(1), 16–20. https://doi.org/10.1007/BF02322211

Rosenthal, R., & Rosnow, R. L. (2008). Essentials of behavioral research: Methods and data analysis (3rd ed.). McGraw-Hill.

Rotton, J., & Kelly, I. W. (1985). Much ado about the full moon: A meta-analysis of lunar-lunacy research. Psychological Bulletin, 97(2), 286–306. https://doi.org/10.1037/0033-2909.97.2.286

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366. https://doi.org/10.1177/0956797611417632