← All lessons
Lesson 13 · Experiments

Factorial Designs

What happens when you run two experiments at once? You get to ask a bigger, more human question: does the effect of one thing depend on another? That "it depends" is the interaction — the single most important idea in this lesson, and the reason factorial designs are how modern psychology actually gets done.

This lesson covers: why Fisher invented the factorial and what it buys you; factors, levels, and cells (and how to decode "2 × 3"); the 2×2 as the workhorse design; main effects from marginal means and interactions from cell patterns; crossover vs. spreading interactions; saying interactions correctly; participant variables as measured factors and what causal language they license; three-way designs and their power problem; and why interactions — moderation — are where psychological theories actually live.

The whole pointWhy cross factors at all?

For most of scientific history, the official advice was: change one thing at a time. Vary the fertilizer or the seed variety or the watering schedule — never two at once, or you won't know what did what. It sounds like discipline. Ronald Fisher, working through mountains of crop data at the Rothamsted agricultural station in the 1920s, decided it was superstition. "We must ask Nature few questions, or, ideally, one question, at a time," he wrote, summarizing the dogma — and then rejected it flat: "The writer is convinced that this view is wholly mistaken. Nature… will best respond to a logical and carefully thought out questionnaire" (Fisher, 1926, p. 511). Cross your factors, Fisher argued, and every plot of land does double duty: the same fields that test fertilizer also test seed variety, at no extra cost. He formalized the idea in The Design of Experiments (Fisher, 1935), and the factorial design — two or more independent variables, every level of each combined with every level of the others — became standard equipment across the experimental sciences.

But efficiency is only half of Fisher's insight, and honestly the boring half. The deeper payoff is that crossing factors lets you ask a question a one-IV experiment cannot ask at any price: does the effect of one variable change depending on the other? A drug that helps adults but not kids. A persuasion tactic that works on strangers and backfires on friends. A masculinity threat that changes men's behavior and leaves women's untouched. Each of those is an interaction, and you cannot see a single one of them without crossing two factors in the same study. A one-IV experiment answers "does X work?" A factorial answers something richer: "does X work the same way for everyone, everywhere?" — and in psychology, the answer to that second question is almost always no, in ways that turn out to be the actual finding.

The vocabulary, tightenedCells, factors, levels — and how to read "2 × 3"

Every factorial has a shorthand, and it's fully decodable. In "2 × 3": the count of numbers = the number of IVs (here, two factors); the value of each number = how many levels that factor has (one factor has 2 levels, the other has 3); and when you multiply the numbers you get the number of cells — the distinct condition combinations (2 × 3 = 6 cells). A cell is one specific recipe: "male participant, masculinity-threatening feedback." Memorize that decoding and you'll never be stumped by a Method section again.

How participants fill the cells is a separate design choice, and it inherits everything you learned in Lessons 11 and 12. In a between-groups factorial, each cell is a separate group of people — clean, but expensive (a 2×2 with 50 per cell needs 200 participants). In a within-groups (repeated-measures) factorial, the same people pass through every cell — 50 participants total, but now order effects are on the table and you must counterbalance. A mixed factorial splits the difference: one factor between, one within. And notice one lovely piece of Fisherian arithmetic: in a 2×2 between-groups design, all 200 participants contribute to the test of each factor — 100 vs. 100 for factor A, a different 100 vs. 100 for factor B. Two experiments' worth of answers, one experiment's worth of people, plus the interaction thrown in free. That's the factorial efficiency Fisher was selling.

Go deeper · the one rule that saves you on the exam

Say it until it's automatic: "An interaction means it depends." That's the whole idea in four words. If the effect of one factor changes depending on the level of the other factor, you have an interaction. On a line graph this becomes a shape rule you can apply in half a second: parallel lines → no interaction (the effect is the same at every level, so nothing depends on anything). Non-parallel lines → probably an interaction (the gap between the lines changes, so the effect depends on where you look). Lines that cross form an X — a perfect crossover — and, notably, a perfect X cancels out both main effects. Parallel or not: that's the fastest read in the lesson.

The workhorseA real 2×2: masculinity threat × participant gender

The 2×2 is to social psychology what the fruit fly is to genetics: small, cheap, and capable of carrying an astonishing share of the field's theory. Here is a real one. Precarious manhood theory holds that manhood — unlike womanhood — is culturally defined as a status that is hard to earn and easy to lose, so men respond to threats against it with public displays of masculinity (Vandello et al., 2008). Willer et al. (2013) tested the idea directly. Participants — men and women — completed a gender-identity questionnaire and then received randomly assigned feedback: your responses place you in the masculine range (identity confirmed) or the feminine range (identity threatened). The feedback was fake; the random assignment was real. Then everyone rated a set of masculine-coded attitudes, including how desirable they found an SUV. That's a 2 (feedback: threat vs. confirm — manipulated) × 2 (participant gender: men vs. women — measured) factorial with four cells.

The result: men told they were feminine cranked up the masculine posturing — more support for war, more homophobic attitudes, more desire for an SUV — while women's answers didn't budge no matter which feedback they got (Willer et al., 2013). The table below shows the pattern with illustrative values (a 1–7 SUV-desirability rating, simplified for easy arithmetic — the shape matches the published result; the numbers are teaching numbers).

Threat ("feminine")
Confirm ("masculine")
Men
6.0
4.0
Women
5.0
5.0

Reading a tableMarginal means and the "difference of differences"

You don't need a graph to find the effects — a 2×2 table is enough, if you know where each effect lives. Main effects live in the marginal means: average each row (or column) across the other factor, and ask whether those averages differ. A main effect of feedback = "averaging over men and women, did threat change the DV?" A main effect of gender = "averaging over both feedback conditions, did men and women differ?" The interaction lives somewhere subtler: the difference of differences. Compute the effect of one factor at each level of the other — threat minus confirm for men, then threat minus confirm for women — and ask whether those two effects are the same. If the two differences differ, you have an interaction. Work the table above before you open the box below; the box has the answers.

Interrogate · work this exact table

Main effect of feedback? Column marginals: threat (6.0 + 5.0)/2 = 5.5 vs. confirm (4.0 + 5.0)/2 = 4.5. A full point apart → yes, a main effect of feedback. Main effect of gender? Row marginals: men (6.0 + 4.0)/2 = 5.0 vs. women (5.0 + 5.0)/2 = 5.0. Identical → no main effect of gender. Interaction? Difference of differences: for men, 6.0 − 4.0 = +2.0; for women, 5.0 − 5.0 = 0. The differences differ → yes, an interaction: threat raised masculine posturing among men and did nothing among women. And notice what the marginals alone would have told you — "feedback had an effect; men and women didn't differ." Both technically true, both missing the entire story. The interaction is the finding. Never stop at the margins.

The two shapesCrossover vs. spreading, in plain English

Interactions come in two visual flavors, and it pays to be able to name them on sight. A crossover interaction is "it depends" at full strength — the direction of the effect flips. You want coffee hot and iced tea cold: the effect of temperature on your enjoyment reverses depending on the drink, and on a graph the two lines cross into an X. A spreading interaction is "only when" — the effect appears at one level of the other factor and vanishes at the other. The Willer et al. (2013) result is a textbook spreading pattern: threat moved men (steep line) and left women flat (level line), so the lines fan apart without crossing. Both count as interactions — in both, the difference of differences is nonzero. The distinction is just whether the effect reverses (crossover) or merely switches on and off (spreading). One more shape fact worth owning: a perfect crossover — +2 in one row, −2 in the other — produces marginal means that are dead equal, meaning zero main effects while both cells are doing something dramatic. That's not a paradox; it's a warning about what averages hide.

Myth · "the main effect is the main finding"

The name lies. A main effect is not the "most important" effect — it's just the average effect of one factor, computed while ignoring the other. When a study finds a significant interaction, the interaction is the headline, and the main effects should be interpreted with caution or not at all, because averaging across an interaction can flatly misrepresent what happened: in the Willer et al. pattern, "no main effect of gender" is true and useless; in a perfect crossover, "no main effect of anything" is true while the phenomenon screams from every cell. Interaction first, main effects second — always.

Say it rightInteraction language done properly

Students lose more exam points on interaction sentences than on interaction arithmetic, so here is the template. A correct interaction statement always has the same skeleton: "the effect of X on the DV depends on the level of Y" — followed by the specifics. For our worked example: "Masculinity threat increased masculine-coded attitudes among men but had no effect among women; the effect of threat depended on participant gender." Compare the common failed attempts. "There was an interaction between feedback and gender" — true but empty; it names the pattern without describing it. "Men were more affected" — closer, but vague about affected by what, on what. "Threat worked and gender didn't matter" — worst of all, because it's two main-effect claims wearing an interaction's coat. The test for your own sentence: does it state one factor's effect at each level of the other factor? If yes, you've described an interaction. If your sentence could be true of a design with only one IV, you haven't.

The causal fine printParticipant variables: the IV × PV design

Look back at the Willer et al. design and notice something quietly important: only one of the two factors was actually an experiment. Feedback was manipulated — randomly assigned, so the threat and confirm groups were equivalent on everything else, and the causal claim goes through. But nobody was randomly assigned to be a man or a woman. Gender is a participant variable (a person's pre-existing characteristic — gender, age, diet, personality), and when you build it into a factorial as a measured factor, it becomes a quasi-independent variable. The design is called an IV × PV design, and it is everywhere in social and personality psychology, because the field's favorite questions have exactly this shape: does a manipulation land differently on different kinds of people?

The fine print is about what causal language each factor licenses. The manipulated factor earns full causal claims: threat caused the shift in men's attitudes. The participant-variable factor earns none on its own: men and women differ in a hundred tangled ways (socialization, identity, hormones, cultural pressure — take your pick), so a "main effect of gender" is a correlational observation dressed in factorial clothes. What the design does license is a precise conditional claim: the causal effect of the manipulation differs across groups. That's not a consolation prize — it's often exactly what the theory predicts. Rothgerber (2013), for instance, documented that men endorse pro-meat justifications far more than women and that this tracks masculine-identity concerns — correlational groundwork that practically begs for the IV × PV follow-up: manipulate a masculinity threat (or a meat-framing message) and test whether it moves men and women differently. One factor carries the cause; the other tells you for whom.

Go deeper · this is Lesson 10's moderation, wearing a lab coat

If "the effect of X depends on Y" sounds familiar, it should — it's the definition of moderation from the multivariate lesson (Baron & Kenny, 1986). A moderator is a variable that changes the strength or direction of a relationship, and a factorial interaction is the same statistical object with better credentials: when X is randomly assigned, "the X→DV effect depends on Y" becomes a claim about a causal effect changing across levels of Y. A factorial with a measured moderator (an IV × PV design) is the experimental twin of correlational moderation — same "it depends," stronger inference on the manipulated side. This is also why interactions are where psychological theories live. Serious theories rarely predict "X raises Y for everyone, always" — they predict for whom and under what conditions. Precarious manhood theory doesn't predict a main effect of gender; it predicts the threat × gender interaction (Vandello et al., 2008). Testing an interaction is usually a sharper theory test than testing a main effect, because far more rival explanations predict "some overall effect" than predict "an effect exactly here and not there."

More factors, more problemsThree-way designs, briefly and honestly

Nothing stops you at two factors. A 2×2×2 crosses three two-level factors — say, feedback (threat/confirm) × participant gender × audience (public/private responses) — and yields 8 cells, three main effects, three two-way interactions, and one three-way interaction: the two-way interaction itself changing across levels of the third factor ("threat boosts men's posturing, but only when responses are public"). These designs exist, get published, and are sometimes exactly what a theory demands. They are also brutally hard to do well, for one compounding reason: cells multiply while the signal you're chasing shrinks.

Update · the power problem, quantified

Adding factors is intoxicating, but the arithmetic is unforgiving. A 2×2 has 4 cells; a 2×2×2 has 8; a 3×4 has 12. Fifty per cell in a between-groups 2×2×2 is already 400 participants — and that understates the problem, because interactions are typically smaller effects than the main effects they qualify. Sommet et al. (2023) worked out the modern guidance: a spreading ("only when") interaction — an effect present in one group and absent in the other — needs roughly four times the sample required to detect the simple effect itself, and interactions that merely dampen an effect (rather than switch it off) need many times more than that. Only the perfect crossover comes cheap. Underpowered high-order factorials produce exactly the flashy, fragile results you'll meet again in the replication lesson (Lesson 15). The professional norm now: preregister the specific interaction your theory predicts, power for that, and treat any unpredicted three-way that "turned up" as a hypothesis for the next study, not a finding in this one. Design lean; don't cross factors just because you can.

Reading the literatureSpotting factorials in the wild

Once you can decode the notation, Method sections open up. "A 2 (self-compassion induction: compassionate writing vs. neutral writing) × 2 (diet: vegan vs. omnivore) between-subjects design" — you should now be able to say, without slowing down: two factors; the induction is manipulated and carries causal claims, diet is a participant variable and doesn't; four cells; and the authors are almost certainly hunting an interaction, because if they only cared about the induction's average effect they wouldn't have crossed it with diet. Then the two-step read of any result they show you: marginals for main effects, difference of differences for the interaction, interaction first. Graph version: parallel or not. Table version: do the differences differ. That's the entire skill, and it transfers to every factorial you will ever meet.


Main effect, interaction, or both? Each button is a 2×2 result pattern. Tap one to reveal what it shows and how you'd read it.

Pick a pattern to see whether it's a main effect, an interaction, both, or neither.

Match the pattern to its interpretation

Click a pattern, then the interpretation it fits. Five to clear.

How to interrogate a 2×2 table, in order

Put the analysis steps in the sequence a careful reader actually follows.

🂠

Name that design (or effect)

Read the description, decide in your head, then flip.

"A 2 × 3 mixed factorial: gender (between-subjects) crossed with three doses of caffeine (within-subjects)."
Tap to flip
Mixed factorial, 6 cells

Two factors → two numbers. One factor is between-groups (gender, also a participant variable) and one is within-groups (dose), so it's mixed. 2 × 3 = 6 cells.

"Masculinity threat raised support for war among men, but women's support didn't move either way."
Tap to flip
Spreading interaction ("only when")

The threat effect appears in one group (men) and not the other. Lines fan apart. Gender is a participant variable — the causal claim belongs to the manipulated threat, conditional on gender (Willer et al., 2013).

"You want coffee hot and iced tea cold — the effect of temperature reverses by drink."
Tap to flip
Crossover interaction ("it depends")

The effect of temperature flips direction depending on the drink. Lines cross into an X, and a perfect crossover cancels both main effects.

"Same 50 participants each complete all four conditions of a 2 × 2."
Tap to flip
Within-groups (repeated-measures) factorial

One group runs through every cell — the most participant-efficient design (50 people, not 200). Watch for order effects; counterbalance.

Check yourself — Factorial Designs quiz

Eight questions, including two "read this 2×2 table." Instant feedback with the reasoning.


SourcesCited in APA 7

Baron, R. M., & Kenny, D. A. (1986). The moderator–mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality and Social Psychology, 51(6), 1173–1182. https://doi.org/10.1037/0022-3514.51.6.1173
Fisher, R. A. (1926). The arrangement of field experiments. Journal of the Ministry of Agriculture of Great Britain, 33, 503–513.
Fisher, R. A. (1935). The design of experiments. Oliver and Boyd.
Rothgerber, H. (2013). Real men don't eat (vegetable) quiche: Masculinity and the justification of meat consumption. Psychology of Men & Masculinity, 14(4), 363–375. https://doi.org/10.1037/a0030379
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Sommet, N., Weissman, D. L., Cheutin, N., & Elliot, A. J. (2023). How many participants do I need to test an interaction? Conducting an appropriate power analysis and achieving sufficient power to detect an interaction. Advances in Methods and Practices in Psychological Science, 6(3). https://doi.org/10.1177/25152459231178728
Vandello, J. A., Bosson, J. K., Cohen, D., Burnaford, R. M., & Weaver, J. R. (2008). Precarious manhood. Journal of Personality and Social Psychology, 95(6), 1325–1339. https://doi.org/10.1037/a0012453
Willer, R., Rogalin, C. L., Conlon, B., & Wojnowicz, M. T. (2013). Overdoing gender: A test of the masculine overcompensation thesis. American Journal of Sociology, 118(4), 980–1022. https://doi.org/10.1086/668417

← Prev: Experimental Control All lessons Next: Quasi-Experiments & Small-N →