OrientingEverything downstream depends on this unit
Developmental psychology is not a collection of facts to memorize; it is a collection of inferences, and every inference is only as trustworthy as the study that produced it. Before you can responsibly believe a single claim about how children think, how adolescents change, or how the aging mind works, you have to know how that claim was tested — and how it could be wrong.
This is why methods come first rather than last. A striking result reported in the news, a confident sentence in an older textbook, a viral graph about screens and teenagers — each of these is a conclusion, and conclusions can be manufactured by a weak design just as easily as by a strong one. The skills in this unit are the tools you use to tell the difference: to ask which people were studied, whether the comparison was fair, how the key variable was measured, and whether anyone else has been able to find the same thing twice. Treat them as a permanent filter you run on every subsequent chapter, because the rest of the course rests on them.
The core disciplineCorrelation is not causation — but why not?
Everyone can recite the slogan. Fewer can say what the two failure modes actually are. A correlation is simply a statistical statement that two variables move together in a patterned way — as one goes up, the other tends to go up (a positive correlation) or down (a negative one). Its strength is summarized by a correlation coefficient, r, that runs from −1.00 through 0 to +1.00, where values near the ends signal a tight linear relationship and values near zero signal none. What that number can never do is tell you why the two variables travel together.
Correlational evidence is genuinely valuable — it is often the only ethical or practical way to study questions like prenatal stress, childhood poverty, or bereavement, where an experimenter cannot assign people to conditions. But its interpretive weakness is fixed and unavoidable: when two things move together, there are always three live possibilities, and a good developmentalist keeps all three on the table at once rather than leaping to the causal story that happens to be the most dramatic.
The directionality problem — you don't know which way the arrow points. Kids who read more have bigger vocabularies. Does reading build vocabulary, or do kids who already know more words enjoy reading more and do it more? Probably both, but a correlation can't tell you.
The third-variable problem — some lurking factor drives both. Ice-cream sales correlate with drowning; the lurker is summer. In development: children spanked more often show more aggression — but harsh temperament, chaotic homes, and poverty can drive both the spanking and the aggression, so the correlation alone proves nothing about spanking causing aggression.
"Teens with more screen time report more depression." Tap each interpretation of that correlation.
Earning the arrowThe experiment, and why causation is expensive
If a correlation cannot license a causal claim, what can? The answer is the experiment, and its defining feature is not lab coats or equipment but a single decisive move: random assignment. In a true experiment the researcher manipulates one variable — the independent variable — and measures its effect on another — the dependent variable — while assigning participants to conditions purely by chance. Random assignment is what dissolves the third-variable problem. Because coin-flips distribute temperament, wealth, prior ability, and every other lurking factor roughly evenly across the groups, any difference that emerges afterward can be attributed to the manipulation rather than to a pre-existing difference between the people in each group.
A well-run experiment therefore aims for high internal validity — confidence that the independent variable, and nothing else, produced the change in the dependent variable (Campbell & Stanley, 1963). Threats to internal validity are the confounds an experimenter works to design out: uncontrolled differences between conditions, expectancy effects, and the placebo response. This is why the strongest experimental medicine uses double-blind procedures, in which neither the participant nor the person collecting the data knows who is in which condition, so that neither hope nor bias can quietly nudge the outcome. The randomized controlled trial that governs modern drug approval is simply this logic taken to its rigorous limit.
The catch, and it is a deep one for developmental science, is that the most important questions in this field cannot be studied experimentally without violating ethics or reality. You cannot randomly assign infants to secure versus abusive caregiving, assign teenagers to smoke or not, or assign one group of children to grow up in poverty. Where manipulation is impossible, researchers fall back on the quasi-experiment, which compares groups that already differ — children in poverty versus not — while statistically controlling for as many confounds as possible. Quasi-experiments buy causal plausibility, never causal certainty, and reading them well means holding that distinction firmly in mind.
The other toolkitCorrelational and observational designs
Because so much of development resists the experiment, the field leans heavily on non-experimental designs, and each carries its own trade-off between richness and control. Naturalistic observation watches behavior unfold in its ordinary setting — a playground, a classroom, a dinner table — sacrificing control for ecological validity, the assurance that what you see resembles real life rather than an artificial lab task. Its enemies are observer bias and the possibility that being watched changes the very behavior under study.
The case study examines a single person or small group in extraordinary depth, and it has repeatedly seeded developmental science with ideas no group study could have surfaced — the tragic case of Genie, deprived of language into adolescence, sharpened every later debate about critical periods. Yet a single case can never establish what is typical, and generalizing from one vivid example is one of the most seductive errors in the field. Survey and interview methods reach large samples cheaply but inherit the frailties of self-report: social desirability, faulty memory, and the gap between what people say and what they do. None of these designs can nail down causation, but together they map the terrain of development that experiments are forbidden to touch.
The bigger pictureMost of psychology studied the world's WEIRDest people
Here's a question worth asking of every study. When you read "studies show children do X," ask: which children? For decades the answer was overwhelmingly the same narrow slice of humanity — and the field is only now reckoning with how much that skew distorted its supposedly universal conclusions.
A landmark analysis found that the vast majority of psychology's participants come from Western, Educated, Industrialized, Rich, and Democratic societies — often literally American undergraduates earning course credit. That's roughly 12% of humanity, and on many measures (visual perception, moral reasoning, notions of fairness, even how the "self" is construed) WEIRD samples are outliers, not a neutral default. So a finding about "human development" may really be a finding about middle-class Western development. When you meet the cross-cultural attachment data later, this is why it matters.
The scale of the problem is not subtle. A content analysis of the field's flagship journals found that samples were drawn almost entirely from the United States, a country holding roughly 5% of the world's population, prompting the blunt conclusion that American psychology needed to become less American (Arnett, 2008). A decade on, the pattern had barely shifted, renewing the call to build a psychology of the whole species rather than of the undergraduates who happen to be convenient (Rad et al., 2018). The stakes are not merely about fairness in sampling; they are about validity. When measures of moral reasoning, visual illusions, cooperation, and self-concept all come out differently outside WEIRD samples, a "finding about human nature" may be a finding about a historically peculiar way of living. Keep this in your pocket for the attachment, moral-development, and self-concept chapters, where the cross-cultural data force exactly this reckoning.
Measurement literacyReliability vs. validity, made intuitive
These two get memorized and immediately confused, so it is worth slowing down. A measure is only as good as these two properties, and they are independent of each other: a measure can have either, both, or neither. The dartboard picture fixes them for life.
Reliability = consistency
Does the measure give the same answer on repeat? A bathroom scale that reads 150, 150, 150 is perfectly reliable. Darts all landing in the same tight cluster — even if that cluster is in the wrong corner. Reliability is necessary but not sufficient: you can be consistently wrong.
Validity = accuracy
Does it measure what you think it measures, and hit the truth? Darts clustered on the bullseye. A scale reading your true weight is valid. The subtle one is construct validity: does your questionnaire really tap "self-esteem," or just "willingness to say nice things about yourself"? A reliable measure of the wrong construct is a very consistent mistake.
Each property comes in several practical flavors. Reliability can be assessed as test–retest consistency (the same people score similarly weeks apart), inter-rater agreement (two independent observers code the same behavior the same way), or internal consistency (the items on a scale hang together). Validity fans out even further, into content validity (do the items cover the whole construct?), criterion validity (does the measure predict a real outcome?), and the deepest of all, construct validity. The foundational treatment held that establishing what a test measures is not a single check but an ongoing accumulation of evidence, gathered by embedding the measure in a web of theoretical predictions and seeing whether they hold (Cronbach & Meehl, 1955). This is why a reliable questionnaire can still be worthless: consistency guarantees only that you are measuring something the same way each time, not that the something is what you named it. A validity claim, unlike a reliability claim, is never finished.
Time-span designsThe confound worth unpacking
Development happens over time, which forces a choice no other area of psychology must make so sharply: how do you fit time itself into the design? Cross-sectional and longitudinal designs answer differently, each has a signature flaw, and understanding those flaws is the whole reason sequential designs exist.
A cross-sectional study compares 20-year-olds to 70-year-olds today and finds the 70-year-olds are worse with smartphones. Is that aging — or is it that they grew up without touchscreens? That confound is a cohort effect: a group difference caused by the historical era, not by getting older. A longitudinal study follows one group across decades, so it separates age from cohort — but it's slow, expensive, loses participants (selective attrition), and its findings may be locked to that one cohort's era. Sequential designs combine both — several age groups, each followed over time — precisely to tease age, cohort, and time-of-measurement apart. That's the payoff of the design.
The conceptual breakthrough that organizes all of this was the recognition that any observed developmental difference is a tangle of three separable influences: the person's age, the cohort they were born into, and the historical time of measurement when the data were collected (Schaie, 1965). Because these three are mathematically confounded — knowing any two fixes the third — no single design can isolate all of them, which is precisely why the sequential logic that layers cross-sectional and longitudinal sampling together became the methodological signature of life-span developmental psychology (Baltes et al., 1977). Longitudinal work carries its own hazards beyond cost and time: selective attrition, in which the participants who drop out differ systematically from those who stay, and practice effects, in which people improve simply from taking the same test repeatedly. A finding that survives a sequential design has cleared a genuinely high bar.
The reckoningThe replication crisis & open science
This is the most important methods story of the last fifteen years, and it is worth telling properly rather than as a slogan. Around 2011 psychology confronted evidence that a large share of its famous, textbook-ready findings did not reproduce when independent labs ran them again — and the diagnosis of why turned out to be more unsettling than any single failed study.
The theoretical alarm had actually been sounded earlier, in a widely read argument that, given small samples, small effects, and flexible analysis, most published research findings in a field could be false without anyone committing fraud (Ioannidis, 2005). The mechanism was made concrete by a demonstration that undisclosed flexibility in how data are collected and analyzed — quietly adding participants, dropping conditions, choosing among outcome measures after the fact — lets a determined researcher present almost any result as statistically significant (Simmons et al., 2011). These "researcher degrees of freedom" rarely feel like cheating from the inside; they feel like reasonable judgment calls, which is exactly what makes them dangerous.
Just how routine these practices were became clear when a large anonymous survey of psychologists found that a striking proportion admitted to at least some questionable research practices, such as selectively reporting the studies that worked or deciding whether to collect more data after peeking at the results (John et al., 2012). Compounding the problem at the level of the whole literature is publication bias: journals prefer positive, novel results, so null findings languish unpublished in what has long been called the file-drawer problem, leaving the visible record systematically overstated (Rosenthal, 1979).
The Open Science Collaboration (2015) tried to reproduce 100 published psychology studies; only about 36–39% yielded a significant result again, and effect sizes were roughly half the originals. The culprit wasn't mostly fraud — it was p-hacking (trying analyses until one crosses p < .05), tiny underpowered samples, and publishing only the studies that "worked." The fixes are now standard practice: preregistration — committing to your hypothesis and analysis before seeing the data (Nosek et al., 2018) — reporting effect sizes and confidence intervals, sharing data and materials openly, and valuing direct replications. When you evaluate any claim in this course, ask: is it preregistered? Has it replicated? How big is the effect, really?
Science doesn't do proof — it does disproof. Following Karl Popper's (1959) falsification, we don't confirm hypotheses; we test the null hypothesis and either reject it (evidence against "no effect") or fail to reject it. We never "accept" or "prove" our own hypothesis, because the next study could always overturn it. A theory earns respect by surviving serious attempts to break it, not by racking up confirmations. So the honest phrasing is always "the data support" or "we reject the null" — never "this proves."
The other reckoningResearch ethics and the IRB
Method is not only about being right; it is about being permitted. Developmental research studies children, patients, and other vulnerable people, and the modern ethical framework governing it grew directly out of historical abuses — most infamously the Tuskegee syphilis study, in which treatment was withheld from Black men for decades so researchers could observe the disease's course. The framework that emerged rests on three principles: respect for persons (people are autonomous agents whose informed consent must be sought), beneficence (maximize benefits and minimize harms), and justice (the burdens and benefits of research are distributed fairly), articulated in the report that still anchors American research ethics (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research, 1979).
In practice these principles are enforced by the Institutional Review Board (IRB), a committee that must approve any study before data collection begins, weighing its scientific value against its risk to participants. Its core requirements are informed consent — a clear, voluntary agreement based on a genuine understanding of what participation involves — together with the right to withdraw at any time, protection of confidentiality, and a debriefing that explains any concealment once the session ends. Working with children adds a layer: because a young child cannot legally consent, researchers obtain a parent's or guardian's informed consent and, wherever the child is old enough to understand, the child's own assent. These are not bureaucratic hurdles bolted onto science; they are part of what separates research from exploitation, and a study that ignores them is invalid in a sense no statistic can repair.
SynthesisReading the evidence like a developmentalist
Pull the threads together and a single disposition emerges — a way of receiving any claim about human development that neither swallows it whole nor rejects it reflexively. It asks what kind of design produced the claim, and therefore what the claim is entitled to say: an experiment can speak of causes, a correlation only of associations, a case study only of possibilities. It asks who was actually studied, and whether "human development" has quietly been narrowed to a WEIRD slice of it. It asks how the central construct was measured, and whether that measure is merely consistent or genuinely valid. It asks whether an age difference might be a cohort difference in disguise. And it asks the question the last fifteen years taught the field to ask of itself: has anyone found this twice?
Held together, these habits are not cynicism but its opposite — the discipline that lets you believe the well-supported findings in the chapters ahead precisely because you know how the weak ones fall apart. Every subsequent unit in this course is an application of what you have just built here. The interesting question is never simply "what did the study find?" but "what is this study, given its design and its sample and its measures, actually in a position to tell us?"
SourcesCited in APA 7
Arnett, J. J. (2008). The neglected 95%: Why American psychology needs to become less American. American Psychologist, 63(7), 602–614.
Baltes, P. B., Reese, H. W., & Nesselroade, J. R. (1977). Life-span developmental psychology: Introduction to research methods. Brooks/Cole.
Campbell, D. T., & Stanley, J. C. (1963). Experimental and quasi-experimental designs for research. Rand McNally.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83.
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLoS Medicine, 2(8), e124.
John, L. K., Loewenstein, G., & Prelec, D. (2012). Measuring the prevalence of questionable research practices with incentives for truth telling. Psychological Science, 23(5), 524–532.
National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. (1979). The Belmont report: Ethical principles and guidelines for the protection of human subjects of research. U.S. Department of Health, Education, and Welfare.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
Popper, K. R. (1959). The logic of scientific discovery. Hutchinson.
Rad, M. S., Martingano, A. J., & Ginges, J. (2018). Toward a psychology of Homo sapiens: Making psychological science more representative of the human population. Proceedings of the National Academy of Sciences, 115(45), 11401–11405.
Rosenthal, R. (1979). The file drawer problem and tolerance for null results. Psychological Bulletin, 86(3), 638–641.
Schaie, K. W. (1965). A general model for the study of developmental problems. Psychological Bulletin, 64(2), 92–107.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.