This page covers two big ideas that live outside the randomized lab. First, the quasi-experimental toolkit — nonequivalent control groups, interrupted time series, regression discontinuity, and natural experiments — and the discipline of interrogating each design against the internal-validity threat list from Lesson 12. Second, small-N and single-case designs, a genuinely different philosophy of evidence in which one subject, measured hundreds of times under tight control, becomes its own experiment. The through-line: causal language is earned by ruling out alternatives, and there is more than one way to earn it.
When you can't randomizeThe one missing ingredient, and what it costs
Lesson 11 made the case that random assignment is the single most powerful move in all of research design: one procedure that neutralizes every preexisting difference between groups at once, measured and unmeasured alike. But whole territories of psychology are closed to it. Sometimes ethics forbids randomizing — you cannot assign children to grow up poor, or assign adults to twenty years of smoking. Sometimes nature or biography has already done the assigning — strokes, bereavement, religious upbringing, dietary identity. And sometimes policy moves in units too big to randomize — a law lands on an entire city or state on the same day. A quasi-experiment keeps the skeleton of an experiment — a real treatment or event, a real comparison — but the groups are formed by nature, law, self-selection, or a schedule rather than by a coin flip (Campbell & Stanley, 1963).
Be clear-eyed about what that one absence costs. Randomization gives you probabilistic equivalence: any confound, known or unknown, gets scattered roughly evenly across conditions. Take it away and every preexisting difference between the groups rides along with the treatment, fully intact. Internal validity stops being a property the design hands you and becomes a case you must build — alternative explanation by alternative explanation, threat by threat. What you buy in exchange is reality: real patients, real laws, real lifelong habits, studied where they actually live. Quasi-experiments routinely beat lab experiments on external validity; the question is always whether they can claw back enough internal validity to say anything causal.
The intellectual tradition here has a definite address. Donald Campbell and Julian Stanley's (1963) slim monograph — originally a handbook chapter — gave the field its two enduring gifts: a shorthand notation for designs (X for a treatment, O for an observation, a row per group) and the famous checklist of threats to internal validity you met in Lesson 12: selection, history, maturation, regression to the mean, attrition, testing, instrumentation. Cook and Campbell (1979) expanded the program for messy field settings, and Shadish, Cook, and Campbell (2002) remains the modern bible. The tradition's core attitude is worth internalizing: the design does not rule out threats for you. You rule them out, one plausible alternative at a time, using comparison groups, pretests, and hard looks at the data.
Here's how close to home this lives: a good chunk of your professor's own research compares vegans and omnivores on psychological outcomes like self-compassion and moral emotion. Nobody is randomly assigned a diet they've kept for ten years — so every vegan–omnivore comparison is, by construction, a nonequivalent-groups design. If vegans score differently on some measure, a selection story is always lurking: the kind of person who goes vegan may have differed on that measure before ever touching a lentil. That doesn't make the comparison worthless; it means the causal claim has to be argued, not assumed.
The toolkitFive designs, two questions
Every design in the quasi-experimental family is an answer to two questions: do you have a comparison group? and do you measure more than once? Each added ingredient patches specific threats and leaves others exposed.
| Design | Skeleton | Best weapon against | Achilles heel |
|---|---|---|---|
| Nonequivalent control group, posttest-only | Two unrandomized groups, one measurement each, after the fact | Nothing much — it documents a difference | Selection: the groups may have differed all along |
| Nonequivalent control group, pretest/posttest | Two unrandomized groups, measured before and after | Lets you check whether groups started apart, and compare change | Selection–maturation: groups changing at different natural rates |
| Interrupted time series | One group, many measurements before and after an "interruption" | Maturation and regression — the long baseline shows the trend | History: something else happening at the same moment |
| Time series + comparison series | Treated and untreated series tracked in parallel | History — the outside event should hit both series | Local events that touch only the treated series |
| Regression discontinuity | Treatment assigned by a sharp cutoff on a known score | Selection — the assignment rule is fully known | Answers only apply near the cutoff; needs lots of cases there |
The posttest-only nonequivalent control group design is the floor of the family. Survey vegans and omnivores once on self-compassion and you have exactly this: two self-selected groups, one snapshot. It can establish that a difference exists — a perfectly good association claim — but it cannot say why. Adding a pretest upgrades you considerably: now you can see whether the groups started in the same place and analyze change rather than raw endpoints. Imagine one company office adopts a mindful self-compassion program while a sister office doesn't, with burnout measured in both offices before and after. If the offices started equal and only the program office improved, selection gets much harder to argue — though a subtle version survives: maybe the office that chose the program was already on a different trajectory. That residual threat, groups maturing at different rates, is why pretests patch selection only partially.
The interrupted time-series design swaps the comparison group for the comparison of a group with itself, over time. Measure the outcome at many points before and after some interruption — a law, a policy, a treatment onset — and ask whether the series shows a break in level or slope right at the interruption. The long baseline is the weapon: maturation and regression to the mean predict smooth continuations of an existing trend, not a sharp discontinuity timed to the event. What the single series cannot beat is history — some other event landing at the same moment. For that you add a comparison series, and here the classic comes from health psychology. In June 2002, Helena, Montana banned smoking in all public buildings and workplaces. During the six months the ban was in force, hospital admissions for heart attacks in Helena fell from an average of about 40 (in the same months of surrounding years) to 24 — roughly a 40% drop — while admissions in the surrounding region outside the ban showed no such decline (Sargent, Shepard, & Glantz, 2004). Then opponents got the ban suspended in court, and admissions climbed back. Notice what the design pieces each do: the comparison region rules out a statewide history threat (there was no regional heart-health miracle that winter), and the rebound after suspension is a community-scale reversal — the outcome tracked the intervention on and off. Hold that thought; it returns at n = 1.
Buried in a 1960 education-journal article is the cleverest idea in the whole family. Thistlethwaite and Campbell (1960) studied students who received a Certificate of Merit — awarded by a strict cutoff on a national exam score — and compared students just above the line to students just below it. The insight: whether you scored 119 or 121 against a cutoff of 120 is close to luck. A point either way reflects the noise in any test — which question you guessed, how you slept. So right at the cutoff, the two groups are nearly identical on everything, measured and unmeasured — the local equivalent of random assignment — except that one group got the treatment. If the outcome jumps discontinuously exactly at the cutoff, the treatment is far and away the most plausible cause. What makes regression discontinuity special among quasi-experiments is that the selection process, the usual unknowable menace, is completely known: assignment was the cutoff rule, full stop. Economists rediscovered the design decades later and made it a workhorse of modern policy evaluation — scholarship effects, class-size rules, legal drinking ages — anywhere a bright line decides who gets what. Its honest limits: the causal estimate applies to people near the cutoff (a scholarship's effect on a 119-scorer may not tell you much about a 90-scorer), you need plenty of cases in that neighborhood, and nobody can be nudging themselves across the line.
Nature runs the studyNatural experiments
Sometimes the world performs a manipulation no ethics board would ever approve, and the researcher's job is to be standing there with a clipboard. A natural experiment is a real change in people's circumstances, imposed by nature or policy, in a way plausibly unrelated to the outcome you care about — which is precisely what makes it useful. The textbook-beautiful case: Costello, Compton, Keeler, and Angold (2003) had been tracking the mental health of 1,420 rural North Carolina children for years — an ordinary longitudinal study — when, partway through, a casino opened on the Eastern Band of Cherokee reservation and began distributing a share of profits to every tribal household. About 14% of the study families rose out of poverty for a reason that had nothing whatsoever to do with their children's psychology. That's the crucial ingredient. In an ordinary comparison of poor and non-poor children, selection is everywhere — families differ in a hundred ways that travel with income. Here, the income change was dropped on families from outside. And the children whose families left poverty showed a marked drop in conduct and oppositional symptoms, landing at the level of children who had never been poor, while children whose families stayed poor showed no such change. Anxiety and depression, interestingly, didn't budge — a specific pattern that itself argues against generic feel-good confounds. No randomization anywhere, yet the causal claim — poverty relief reduced these children's behavioral problems — stands on unusually firm ground, because the "assignment" mechanism was external, abrupt, and documented.
Economists ran into these same problems, sharpened the tools, and renamed them — and the names are worth knowing because you will meet them constantly in policy research. A time-series-plus-comparison-series analysis like the Helena study is close kin to difference-in-differences: take the change in the treated group and subtract the change in the comparison group over the same window, so the comparison's trend estimates "what would have happened anyway." Regression discontinuity kept its name and became a star of the so-called credibility revolution in empirical economics. The traffic runs both ways: Campbell's 1960s psychology gave economics its designs, and economics gave psychology back a harder-nosed standard for what counts as a defensible causal claim without randomization.
Don't memorize threats — ask them, as questions, of every design you meet. Selection: could the groups have differed before the treatment? (Every vegan–omnivore comparison, forever.) History: did some outside event hit during the study window? (Helena's answer: the comparison region.) Maturation: would people have changed anyway on their own schedule? (The casino study's answer: the still-poor children didn't.) Regression to the mean: was anyone selected for extreme scores that would drift back regardless? (Time-series answer: a long baseline reveals the drift.) Attrition: did a systematic kind of person drop out? (Check whether dropouts differ from completers.) Testing/instrumentation: did repeated measurement or a changed instrument do the work? Then note which patch fixes which hole: a pretest answers "did they start equal?", a comparison group absorbs history and maturation, many time points expose regression and trends, and data checks handle attrition and confounds. A design earns causal language exactly to the extent that every question on this list has a good answer.
Worked exampleA quasi-experiment defends itself — then keeps defending
Watch the interrogation happen to a famous finding. Danziger, Levav, and Avnaim-Pesso (2011) tracked about 1,100 parole rulings by Israeli judges across the working day — structurally an interrupted time series, with the judges' two food breaks as the interruptions. The proportion of favorable rulings started near 65% at the beginning of each session, sagged toward zero as the session wore on, and snapped back after each break — a striking pattern the authors read as mental depletion, eased by rest and food. They did the defensive work you'd hope for: the obvious design confound is that easier cases might be scheduled after breaks, so they checked whether case seriousness, prior incarcerations, and sentence length varied with position in the session, and found no relationship.
Then the story kept going — which is the real lesson. Weinshall-Margel and Shapard (2011) went to the Israeli courts and found that case ordering was not random after all: the boards tended to finish all of one prison's cases before a break, and prisoners without attorneys — whose cases fare worse — tended to come later within each block. Danziger and colleagues replied that the pattern survived controlling for legal representation. A decade on, methodologists still argue about it. Notice what this is: not a scandal, but quasi-experimentation working as designed. A randomized experiment settles the assignment question by construction; a quasi-experiment's causal claim stays permanently open to a newly discovered selection story, and honest researchers treat "we checked" as an ongoing obligation rather than a one-time stamp.
The N=1 turnA different philosophy of evidence
Now flip the entire strategy. Everything so far — and honestly, most of this course — assumes the road to truth runs through many participants, with statistics averaging away the noise. Small-N designs reject the premise. Instead of a little information about many people, collect an enormous amount of information about a few — sometimes a single-case design with one person or animal. B. F. Skinner (1938) built the experimental analysis of behavior this way: individual rats and pigeons, responding thousands of times under exquisitely controlled conditions, with the cumulative record of each animal's behavior serving as the data. Murray Sidman's Tactics of Scientific Research (1960) made the philosophy explicit: if you can bring one organism's behavior under tight experimental control — turn it on and off at will — you have demonstrated causation more directly than any between-group average can, and you replicate by doing it again within the same subject and then in the next subject. Group statistics, on this view, can even mislead: a smooth average learning curve may describe no individual learner in the group.
Here's the arithmetic worth sitting with. A between-groups study with n = 1,000 typically has one or two observations per person — all its power comes from the crowd. A single-case study may have one person and a thousand observations — its power comes from stability, repetition, and control. When measurement is precise and behavior can be brought to a steady state, n = 1 with 1,000 observations can support stronger causal conclusions than n = 1,000 with one observation each, because every alternative explanation has to survive contact with a long, stable, repeatedly-perturbed record. The oldest proof is Hermann Ebbinghaus (1885/1913), who taught himself thousands of nonsense syllables and measured his own forgetting with chronometric patience. His forgetting curve — steep early loss, then a long slow tail — came from a sample of one, and it has replicated for 140 years.
Neuropsychology supplies the other legendary case. In 1953, a young man known for fifty years only as H.M. underwent bilateral removal of medial temporal lobe structures, including most of the hippocampus, to treat intractable epilepsy. Scoville and Milner (1957) documented the devastating and weirdly specific result: his intelligence was intact, his old memories largely intact, but he could form almost no new lasting memories of facts and events. Decades of subsequent testing sharpened the picture — he could acquire new motor skills across sessions while having no recollection of ever practicing them — splitting memory into systems ("knowing how" versus "knowing that") and tying declarative memory to the hippocampus. One patient, studied with experimental care against matched comparisons, reorganized an entire field. That's the small-N wager paying off: the value wasn't that H.M. was representative; it was that his case, measured deeply, was decisive.
This is the misconception this whole section exists to kill. A testimonial — "I tried it and I feel great!" — is a single unmeasured observation with no baseline, no control, and no way to rule out coincidence. A well-run single-case experiment is its opposite in every particular: it establishes a stable baseline (many measurements before anything changes), introduces the treatment, and then either reverses it (ABAB) to show the behavior tracks the treatment on and off, or staggers it across behaviors, settings, or people (multiple baseline) so that mere coincidence would have to strike each target exactly on schedule. That repeated, controlled measurement gives a good single-case study high internal validity — the very thing an anecdote has none of. One is data; the other is a story. The difference is the design, not the sample size.
The small-N toolkit maps neatly onto the threats it fights, and every piece should now look familiar (Kazdin, 2011). A stable baseline is the n = 1 version of a pretest — but better, because it's many pretests: if a client's daily panic-attack count sat flat at four per day for three weeks and dropped only when treatment began, maturation and regression to the mean are already in trouble. An ABAB reversal design is the n = 1 version of the Helena rebound: baseline (A), treatment (B), withdraw the treatment (A) and watch the behavior return, reinstate it (B) and watch it improve again. A behavior that obeys the switch four times in a row leaves coincidence nowhere to hide — and ending on B means the participant keeps the working treatment. When reversal is impossible (you can't unlearn a skill) or unethical (you don't withdraw a treatment that's stopping self-injury), the multiple-baseline design staggers instead: introduce the intervention for three clients at different start dates, or for one client across three settings, and show that each behavior changes precisely when — and only when — its treatment begins. An outside event would have to arrive on three different schedules, each perfectly timed. Coincidence isn't that punctual.
How do single-case researchers analyze all this? Mostly by visual analysis — trained inspection of the plotted data for changes in level, trend, variability, and how quickly the change follows the phase switch (Kazdin, 2011). The tradition's logic is bracing: if an effect is real and under experimental control, you should be able to see it, and an effect too faint to survive eyeballing may be too faint to matter clinically. (Statistical techniques for single-case data exist and are increasingly used alongside the graphs, but the graph remains the primary exhibit.) Applied behavior analysis — the treatment science built on Skinner's foundation — still runs on these designs today, precisely because a clinician's real question is about this client, not the average client.
Internal validity: high. Repeated measurement with reversals or staggered onsets within the same subject is a genuine experiment — assignment-to-phases is under the researcher's control even though there's one participant. External validity: the weak spot. One person may represent no one else, so the moves are replication across new subjects (Sidman's answer) and triangulation — H.M.'s lesions-and-memory findings generalize because neuroimaging in ordinary adults converges on the same hippocampal story. Sometimes generalization isn't even the goal; the clinician treating this client wants a conclusion about this client. Construct validity: fine when the outcome is objective (a digit string recalled, a panic attack logged), but judgment calls ("was that rumination?") demand multiple observers and interrater reliability checks. Statistical validity: often carried by the graph plus an effect-size question — by what margin, and how immediately, did the behavior change?
The verdictWho earns causal language?
Line the whole lesson up on one axis: how thoroughly does the design answer the threat checklist? A randomized experiment answers it by construction. Regression discontinuity comes surprisingly close, because its selection mechanism is fully known. A time series with a comparison series — especially one with a lucky reversal, like Helena — answers most questions well. A natural experiment's strength rides entirely on how external and abrupt the "assignment" was; the casino study scores high. A pretest/posttest nonequivalent design answers some questions; a posttest-only design answers almost none, and should stick to association language. And a single-case design with stable baselines and reversals earns full causal language about its subject, cashing in external validity it must buy back through replication and triangulation. The skill this lesson leaves you with is not reverence for any design — it's the interrogation. Next lesson asks the uncomfortable follow-up: when whole literatures of properly-run studies get re-examined, how much survives?
Name that design. Tap a scenario to reveal the design label and what to interrogate.
SourcesCited in APA 7
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.
Cook, T. D., & Campbell, D. T. (1979). Quasi-experimentation: Design and analysis issues for field settings. Houghton Mifflin.
Costello, E. J., Compton, S. N., Keeler, G., & Angold, A. (2003). Relationships between poverty and psychopathology: A natural experiment. JAMA, 290(15), 2023–2029. https://doi.org/10.1001/jama.290.15.2023
Danziger, S., Levav, J., & Avnaim-Pesso, L. (2011). Extraneous factors in judicial decisions. Proceedings of the National Academy of Sciences, 108(17), 6889–6892. https://doi.org/10.1073/pnas.1018033108
Ebbinghaus, H. (1913). Memory: A contribution to experimental psychology (H. A. Ruger & C. E. Bussenius, Trans.). Teachers College, Columbia University. (Original work published 1885)
Kazdin, A. E. (2011). Single-case research designs: Methods for clinical and applied settings (2nd ed.). Oxford University Press.
Sargent, R. P., Shepard, R. M., & Glantz, S. A. (2004). Reduced incidence of admissions for myocardial infarction associated with public smoking ban: Before and after study. BMJ, 328(7446), 977–980. https://doi.org/10.1136/bmj.38055.715683.55
Scoville, W. B., & Milner, B. (1957). Loss of recent memory after bilateral hippocampal lesions. Journal of Neurology, Neurosurgery & Psychiatry, 20(1), 11–21. https://doi.org/10.1136/jnnp.20.1.11
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Sidman, M. (1960). Tactics of scientific research: Evaluating experimental data in psychology. Basic Books.
Skinner, B. F. (1938). The behavior of organisms: An experimental analysis. Appleton-Century.
Thistlethwaite, D. L., & Campbell, D. T. (1960). Regression-discontinuity analysis: An alternative to the ex post facto experiment. Journal of Educational Psychology, 51(6), 309–317. https://doi.org/10.1037/h0044319
Weinshall-Margel, K., & Shapard, J. (2011). Overlooked factors in the analysis of parole decisions. Proceedings of the National Academy of Sciences, 108(42), E833. https://doi.org/10.1073/pnas.1110910108