In this lesson: population vs. sample vs. sampling frame; census; why representativeness beats size; the probability methods (simple random, systematic, stratified, cluster, multistage); the non-probability methods (convenience, purposive, snowball, quota) and what they can still legitimately support; self-selection and nonresponse bias; when external validity is make-or-break and when it's deliberately deprioritized; the WEIRD problem; and data quality in the online-panel era.
The vocabularyPopulations, samples, and the list in between
Start with three terms that look interchangeable and aren't. The population is the entire group your claim is about — every U.S. adult, every first-year nursing student, every person who has ever quit eating meat. The sample is the subset you actually measured. And hiding between them is the term students skip and researchers obsess over: the sampling frame, the actual, operational list you drew your sample from. Your population might be "U.S. adults," but your frame is whatever you could really get your hands on — a voter file, a phone directory, the user base of an app. Whenever the frame doesn't match the population, bias has already entered the study before a single person is contacted.
Why sample at all? Because the alternative — a census, measuring literally everyone — is almost always impossible. Even the U.S. government, with constitutional authority and billions of dollars, struggles to count every resident once a decade, and known undercounts of renters, young children, and marginalized groups persist anyway. Sampling isn't the budget option; it's the only option. The entire statistical machinery of survey research exists because a well-chosen fraction can stand in for the whole.
The everyday proof is a blood test. Your doctor doesn't drain you to check your cholesterol; a few milliliters will do, because blood is continuously mixed — any spoonful is like any other. That mixing is doing real work in the metaphor. A sample represents a population only when something like mixing has occurred: when chance, not convenience or self-interest, decided who got scooped up. Take the spoonful from an unstirred pot and you get the top layer, however big your spoon.
The through-lineRepresentativeness beats size. Every time.
Here is the single most important idea in this lesson, and the one most reliably gotten wrong: how well a sample generalizes depends almost entirely on how people were selected, and almost not at all on how many. Sample size and sample bias are different dials. Size controls precision — a bigger probability sample gives you a tighter margin of error around your estimate. Selection controls accuracy — whether that estimate is centered on the truth in the first place. A biased sample is a rifle with a bent sight: more shots just give you a tighter cluster around the wrong spot.
This is why the phrase "margin of error" on a poll is only meaningful for probability samples. The math that produces "±3 points" assumes random selection; apply it to a self-selected internet poll and it describes the precision of an estimate about the wrong population. Generalizing from sample to population is the heart of external validity — one of the four validities in the Cook and Campbell tradition (Shadish, Cook, & Campbell, 2002) — and no quantity of biased data buys you any of it.
The proof1936: the biggest sample ever assembled, wrong by 46 states
If you remember one study from this page, make it a magazine poll. In 1936, Literary Digest — which had correctly called every presidential election since 1916 — mailed out roughly ten million straw-poll ballots and got about 2.4 million back, a sample so enormous that no modern poll comes within two orders of magnitude of it. The Digest confidently predicted that Alf Landon would defeat Franklin Roosevelt with 57% of the vote. Roosevelt won 46 of 48 states, in one of the great landslides in American history. Meanwhile, George Gallup called the election correctly with a sample in the tens of thousands — a fraction of a percent of the Digest's. Size lost to selection, publicly and fatally; the magazine itself folded within two years (Squire, 1988).
The textbook version blames the sampling frame: the Digest mailed ballots to people on automobile-registration lists and in telephone directories, and in the depths of the Depression, owning a car or a phone skewed wealthy — and wealthy skewed Republican. That's true, but Squire (1988) found the fuller story using a 1937 Gallup survey that asked people whether they had received a ballot and whether they had returned it. The frame bias was real, but the bigger killer was nonresponse: even among people who received ballots, Landon supporters were far more likely to mail them back than Roosevelt supporters. Two separate biases, stacked. And note the eeriest detail: the same biased method had called four straight elections correctly — a biased procedure can keep working right up until the bias finally lines up with the thing you're measuring. Being right is not evidence that your method is sound.
Modern pollsters don't sample from car registries, so what happened in 2016? The professional postmortem (Kennedy et al., 2018) found the national polls were actually fine — they predicted the popular-vote margin about as accurately as usual. The misses were concentrated in state polls, especially the Upper Midwest, and traced to two causes: a genuinely late swing among undecided voters, and a weighting failure — many state polls didn't adjust for education, so college graduates (who answer surveys at higher rates and broke for Clinton) were overrepresented. The moral is the 1936 moral, updated: the danger has migrated from who's on the list to who bothers to answer — and the fix is better modeling of nonresponse, never simply "more people."
Probability methodsWhen chance does the choosing
A probability sample is any sample in which every member of the population has a known, nonzero chance of being selected, and chance — not the researcher, not the participant — makes the final call. This family of methods is the only route to a sample you can defend as representative, and the classic engineering of it is laid out in Kish's (1965) Survey Sampling, still the field's reference manual. The members of the family differ mainly in how they cope with an awkward fact: for most interesting populations, no complete list of members exists.
| Method | How it works | When it's actually used |
|---|---|---|
| Simple random | Every member of the frame has an equal chance; names out of a (digital) hat | The conceptual gold standard — but rare in practice, because it requires a complete frame of the whole population, which usually doesn't exist |
| Systematic | Pick a random start, then take every kth person on the list | A practical stand-in for simple random when you have an ordered list; equivalent to random unless the list has a hidden repeating pattern |
| Stratified random | Divide the population into meaningful strata (e.g., region, class year), then sample randomly within each | When you must guarantee that key subgroups show up in the right proportions — or show up at all; also squeezes more precision from the same N |
| Cluster | Randomly select whole pre-existing groups (schools, clinics, city blocks), then take everyone in the chosen clusters | When there's no list of individuals but there is a list of groups; trades some precision for enormous practicality |
| Multistage | Random selection at two or more levels: randomly pick clusters, then randomly pick people within them | The workhorse of big national surveys — how face-to-face studies of "American households" are really done |
Keep the classic confusion pair straight: strata are meaningful, clusters are arbitrary. You stratify by class year because class year matters to your question; you cluster by classroom because classrooms are simply where the students happen to be stored. And file away a warning for Lesson 11: random sampling (who gets into the study — an external-validity issue) is a completely different operation from random assignment (who gets which condition — an internal-validity issue). A study can have either, both, or neither, and exams love to swap the labels.
Name that methodSort the scenario into a sampling technique
The exam won't ask you to define "stratified random sampling" — it'll describe a researcher doing something and ask what it's called. Build that reflex. Read each scenario, decide, then tap.
What sampling method is this? Tap one.
Non-probability methodsWhen chance can't do the choosing
Now the honest half of the lesson: most psychology studies don't use probability samples, and never have. A non-probability sample is any sample where selection probabilities are unknown — where who ends up in the study is determined by convenience, referral, or volunteering rather than chance. The family portraits: convenience sampling studies whoever is easiest to reach (historically, the introductory-psychology subject pool). Purposive sampling deliberately recruits a specific kind of person because the research question demands them. Snowball sampling lets participants recruit other participants, which is often the only way into hidden or stigmatized populations. Quota sampling sets demographic targets and fills them non-randomly — stratified sampling's cheaper cousin.
Here's a modern, realistic example of purposive recruitment done well. Suppose a researcher wants to compare vegans and omnivores on self-compassion. Vegans are a small minority of the population, so a random sample of 400 adults might contain a dozen of them — useless for a group comparison. Instead, the researcher uses an online panel's prescreening filters to deliberately recruit, say, 100 vegans and 100 omnivores in the United States and the same in the United Kingdom: purposive selection with quota-style balancing across diet and country. Nothing about this is random, and no one should pretend otherwise. But notice what the design can support: a claim about how the groups differ and whether that difference replicates across two countries. What it cannot support is a prevalence claim — "X% of vegans are…" — because nobody randomly sampled the world's vegans. The method isn't a flaw being hidden; it's a scope condition being honored.
That's the general rule for the whole non-probability family: they are legitimate tools for association and theory-testing questions, where the bet is that the psychological process under study works similarly in the sampled and unsampled alike. They are illegitimate tools for frequency claims, where the entire claim is a statement about the population's composition. A convenience sample can tell you whether two variables travel together; it cannot tell you how common anything is.
Who shows upSelf-selection and nonresponse
Two biases deserve their own billing because they can wreck even a well-framed study. Self-selection bias (volunteer bias) means the people who choose to participate differ systematically from those who don't — and they do: volunteers tend to be more educated, more sociable, higher in need for approval, and more interested in the topic than refusers. An online restaurant rating is the purest case: the diners motivated to write reviews are disproportionately the delighted and the furious. Nonresponse bias is the same demon wearing survey clothing — the frame was fine, the invitations went out properly, but the people who answered differ from the people who didn't.
These get conflated on every exam, so pin the difference. A biased sampling frame means the wrong list to begin with — the Digest's car owners and telephone subscribers. Nonresponse bias means the list was fine, but the subset who responded differs from the subset who didn't — the Digest's Landon voters eagerly mailing ballots back while Roosevelt voters tossed them (Squire, 1988). Both wreck external validity, but only nonresponse survives a perfect frame — which is why it is the harder, more modern problem, and the one behind the 2016 state-poll misses (Kennedy et al., 2018). Diagnostic question: could this bias exist even if the list were perfect? If yes, you're looking at nonresponse or self-selection, not the frame.
The plot twistWhen external validity is beside the point
By now you might conclude that any study without a probability sample is quietly worthless. Here is the twist: for large stretches of psychology, external validity is deliberately deprioritized — and the field's best defense of that choice is Mook's (1983) wonderfully titled "In defense of external invalidity." Mook's argument: most experiments aren't trying to estimate what the population does; they're testing whether a theory's prediction holds. The laboratory is artificial on purpose, the way a vacuum chamber is — you strip away real-world noise precisely to see whether the predicted process can occur at all. When a memory researcher shows that a mnemonic works in 40 undergraduates, the claim being tested is about how human memory works, not about the national prevalence of anything. What generalizes is the theory, not the percentage.
When you read any study, ask two questions in order. First: what kind of claim is being made? If it's a frequency claim — how common, what percentage, how prevalent — external validity is make-or-break, and only a probability sample will do. Polling and prevalence research live or die here. If it's an association or causal claim, a non-probability sample is often acceptable. Second: is the sample's bias plausibly connected to the variables under study? A sample of sophomores is a poor place to study attitudes toward retirement, but a perfectly reasonable place to study visual attention. The sample has to be biased with respect to the process being studied before the bias becomes fatal (Mook, 1983). "Convenience sample" is a description, not a verdict.
The WEIRD problemPsychology's narrow database
There is, however, a version of the sampling problem that theory-testing can't shrug off. Sears (1986) warned that social psychology had built its entire picture of "human nature" from one peculiar demographic: college sophomores, tested in laboratories, on academic-flavored tasks. Sophomores aren't miniature adults — they have less crystallized attitudes, less settled identities, stronger cognitive skills, and a greater tendency to comply with authority than the population at large — and a science built on them will quietly mistake late adolescence for the human condition.
Henrich, Heine, and Norenzayan (2010) blew the critique open. Surveying the top psychology journals, they found that roughly 96% of research participants came from Western, Educated, Industrialized, Rich, and Democratic — WEIRD — societies, home to about 12% of humanity, with American undergraduates dominating even that slice. Worse, wherever cross-cultural data existed, WEIRD samples turned out to be among the least typical humans available: outliers on visual illusions like the Müller-Lyer, on fairness in economic games, on moral reasoning, even on spatial cognition. The deep threat isn't that our samples are unrepresentative — it's that a discipline can't always tell which of its findings are human universals and which are local curiosities of one unusual tribe. That question runs all the way to Lesson 15.
The modern eraFrom the sophomore pool to the online panel
The field's answer to the sophomore problem was to move recruitment online. Buhrmester, Kwang, and Gosling (2011) introduced psychologists to Amazon's Mechanical Turk with a startling report card: samples noticeably more diverse than the campus pool, data as reliable as traditional methods, collected in days for pocket change. A gold rush followed — and then the quality problems: "professional" survey-takers completing thousands of studies, farms of bots, and a small hyperactive worker pool answering everyone's surveys. Purpose-built research platforms emerged next, and head-to-head comparisons found that participants on Prolific were more naive, more diverse, and at least as honest and attentive as MTurk workers (Peer, Brandimarte, Samat, & Acquisti, 2017; Peer, Rothschild, Gordon, Evernden, & Damer, 2022).
The online-panel era: better, but not solved
| Platform | What it is | The catch |
|---|---|---|
| MTurk | Amazon's crowd-work marketplace; the field's first mass online pool (Buhrmester et al., 2011) | "Professional" survey-takers, bots, and a small over-sampled worker base; quality varies wildly without screening |
| Prolific | Purpose-built for research, with rich prescreening (diet, nationality, occupation…) and fair-pay rules | Still self-selected online volunteers; representativeness must be checked, never assumed |
| CloudResearch | A vetting layer over MTurk that pre-screens for attentive, verified responders | Screening raises data quality substantially (Peer et al., 2022) but cannot turn a panel into a probability sample |
| Opt-in panels (Qualtrics, Dynata…) | Commercial pools recruited for market research, rentable for science | Scored worst on attention, comprehension, and honesty in head-to-head tests (Peer et al., 2022) |
The upgrade is real — an online panel is far more diverse in age, region, and occupation than a room of sophomores. But note the deep continuity: these are still convenience samples. No one on any panel was randomly drawn from the population. The sophomore problem didn't disappear; it grew up and moved online — and stayed WEIRD.
The online era added a problem the Literary Digest never faced: participants who are present but not paying attention. The standard countermeasure is the attention check — Oppenheimer, Meyvis, and Davidenko (2009) called it the instructional manipulation check, an item whose instructions quietly tell you to ignore the obvious answer ("to show you are reading, select 'strongly disagree'"). In some of their samples, nearly half of participants failed it, and screening the failures out measurably increased statistical power. Modern practice treats data quality as something you measure and report — attention checks, completion-time screens, bot detection, open-ended responses read by human eyes — rather than something you hope for. When you read an online study, look for that paragraph; its absence is informative.
① What is the claim — frequency, association, or causal? (Frequency ⇒ representativeness is everything.) ② What population is the claim about, and what was the actual frame? ③ How were people selected — did chance ever get a vote? ④ Who declined, dropped out, or never responded — and could they differ from those who stayed? ⑤ Is the sample's bias plausibly related to the variables under study, or incidental to them? ⑥ If it's an online sample: what data-quality screening is reported? A study can pass all six with a convenience sample — and a 2.4-million-person sample can flunk at question ③.
SourcesCited in APA 7
Buhrmester, M., Kwang, T., & Gosling, S. D. (2011). Amazon's Mechanical Turk: A new source of inexpensive, yet high-quality, data? Perspectives on Psychological Science, 6(1), 3–5. https://doi.org/10.1177/1745691610393980
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. https://doi.org/10.1017/S0140525X0999152X
Kennedy, C., Blumenthal, M., Clement, S., Clinton, J. D., Durand, C., Franklin, C., McGeeney, K., Miringoff, L., Olson, K., Rivers, D., Saad, L., Witt, G. E., & Wlezien, C. (2018). An evaluation of the 2016 election polls in the United States. Public Opinion Quarterly, 82(1), 1–33. https://doi.org/10.1093/poq/nfx047
Kish, L. (1965). Survey sampling. John Wiley & Sons.
Mook, D. G. (1983). In defense of external invalidity. American Psychologist, 38(4), 379–387. https://doi.org/10.1037/0003-066X.38.4.379
Oppenheimer, D. M., Meyvis, T., & Davidenko, N. (2009). Instructional manipulation checks: Detecting satisficing to increase statistical power. Journal of Experimental Social Psychology, 45(4), 867–872. https://doi.org/10.1016/j.jesp.2009.03.009
Peer, E., Brandimarte, L., Samat, S., & Acquisti, A. (2017). Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70, 153–163. https://doi.org/10.1016/j.jesp.2017.01.006
Peer, E., Rothschild, D., Gordon, A., Evernden, Z., & Damer, E. (2022). Data quality of platforms and panels for online behavioral research. Behavior Research Methods, 54(4), 1643–1662. https://doi.org/10.3758/s13428-021-01694-3
Sears, D. O. (1986). College sophomores in the laboratory: Influences of a narrow data base on social psychology's view of human nature. Journal of Personality and Social Psychology, 51(3), 515–530. https://doi.org/10.1037/0022-3514.51.3.515
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Houghton Mifflin.
Squire, P. (1988). Why the 1936 Literary Digest poll failed. Public Opinion Quarterly, 52(1), 125–133. https://doi.org/10.1086/269085