The naive theoryAsking feels easy — that's the trap
Here's the folk theory of surveys: people have opinions and memories stored in their heads like files; a question retrieves the file; the answer is the file's contents. If that theory were true, survey design would be clerical work. It is not true, and the evidence against it is some of the most entertaining in psychology. Ask Americans whether the government should forbid public speeches against democracy and about half say yes; ask whether it should not allow them and three quarters say no. Forbidding and not allowing are the same policy — yet the two wordings produced answers about twenty points apart (Rugg, 1941, as cited in Schwarz, 1999). Nothing about the respondents' "true attitude" changed between versions. The question changed, and the answer followed.
Norbert Schwarz's (1999) landmark review put it bluntly in its subtitle: self-reports are shaped by how the questions are asked. That's not a reason to abandon surveys — self-report remains the only direct route to subjective experience, and most of what we know about attitudes, emotions, and beliefs came through it. It's a reason to treat every item the way Lesson 06 taught you to treat every measure: as an instrument whose construct validity must be earned, not assumed. A survey question is a ruler. Rulers can be miscalibrated.
Inside the respondent's headThe four-stage model: answering is a cognitive task
The best framework for understanding why questions misfire comes from Tourangeau, Rips, and Rasinski's The Psychology of Survey Response (2000). Their model says answering even a mundane question requires four cognitive stages, and error can enter at every one:
1 · Comprehension. The respondent must figure out what you're asking — not what you meant, what your words convey. Ask "how many drinks did you have last week?" and respondents must decide: does a shared bottle of wine count as one drink or three? Does "last week" mean the previous seven days or the previous calendar week? Respondents resolve ambiguity using every cue available, including the response scale itself — which is how the scale becomes part of the question.
2 · Retrieval. They must pull relevant information from memory. For attitudes, there is often no stored "file" to retrieve; people construct a judgment on the spot from whatever is mentally accessible at that moment — which is why the questions that came before matter so much. For behavioral frequencies, nobody has a mental ledger of, say, vegetable servings; they retrieve a few instances and go to stage three.
3 · Judgment and estimation. They convert what they retrieved into an answer — usually by estimating. "I probably watch TV most evenings, call it two hours" is not recall; it's inference, built from a rate and a rough calculation, and open to every bias inference is heir to.
4 · Response. Finally they map their internal answer onto your response options — and, at this last step, decide whether to edit it. This is where social desirability lives: the answer that leaves the head is not always the answer that was formed.
Keep this model in your pocket. Every classic survey flaw is a failure at one of the four stages, and diagnosing which stage is what separates a survey methodologist from someone who just writes questions.
Rensis Likert (rhymes with "lick-ert," not "lie-kurt") introduced his attitude-scaling method in 1932, and a technical distinction is worth pinning down: a single agree–disagree statement is a Likert item; you only have a true Likert scale when you sum or average several items measuring the same construct. That aggregation is the whole point — it's the survey version of Lesson 06's reliability logic, letting the idiosyncratic noise in any one item cancel out. The semantic differential (rating a target between bipolar adjective anchors like weak … strong) is a cousin format, best when you want the connotation of a concept rather than agreement with a claim. And always label your endpoints: an unanchored "1 to 7" forces every respondent to invent their own private ruler.
The questions shape the answersThree demonstrated ways to bend a survey
Wording. The forbid/allow gap above is the classic, but the family is large: leading questions smuggle the desired answer into the question ("Don't you agree that…?"), loaded terms attach evaluations to the topic ("greedy corporations," "hardworking teachers"), and double-barreled items ask two questions while permitting one answer — if a respondent agrees that they "eat healthy and exercise regularly," you have learned nothing interpretable about either half. Negations, especially doubled ones ("Do you disagree that the policy should never…?"), turn the item into a reading-comprehension test, which means you're partly measuring verbal ability instead of your construct. Every one of these is a comprehension-stage failure: the words did work you didn't intend.
Response anchors. The scale is not a neutral container — respondents read it as information about what's normal. Schwarz, Hippler, Deutsch, and Strack (1985) asked Germans how much television they watched daily. One group got a scale running from "up to half an hour" to "more than two and a half hours"; the other got a scale running from "up to two and a half hours" to "more than four and a half hours." On the low-range scale, 16% reported watching more than two and a half hours; on the high-range scale, 38% did. Same behavior, same question stem — the scale itself told respondents what a "typical" amount was, and they estimated accordingly. The lesson generalizes: whenever people must estimate (stage 3), they lean on whatever frame you hand them.
Question order. Strack, Martin, and Schwarz (1988) asked students two questions: how happy they were with life in general, and how often they were dating. Asked in that order — general first — the two answers were essentially uncorrelated. Reverse the order, asking about dating first, and the correlation leapt to about .66. Why? The dating question made romantic life mentally accessible right when students had to construct a judgment of their life overall (stage 2), so they built the general judgment partly out of it. Context isn't contamination sitting around the measurement; on a survey, context is part of the measurement. The working defenses: randomize or counterbalance item order across respondents, and put the most general questions before the specific ones that could infect them.
The wording traps, in one table
| Trap | What it looks like | Stage it breaks | The fix |
|---|---|---|---|
| Leading / loaded | "Don't you agree that hardworking teachers deserve a raise?" | Comprehension → judgment | Neutral framing; strip evaluative adjectives |
| Double-barreled | "Do you eat healthy and exercise regularly?" | Comprehension | One idea per item — split it |
| Negation | "Do you disagree that it should never be repealed?" | Comprehension | Word positively |
| Skewed anchors | Frequency scale that makes heavy use look typical | Judgment / estimation | Pilot the scale; use open-ended frequency when possible |
| Order / priming | Specific question makes content accessible for the next one | Retrieval | Randomize or counterbalance order; general before specific |
Spot the trapWhat's wrong with this question?
Train the diagnostic reflex: read each item, name the flaw and the stage it breaks, then tap to check.
Which flaw does each question commit? Tap one.
Response setsThe respondent's habits, not the question's flaws
Even a perfectly worded survey meets respondents who answer in patterned ways that have nothing to do with the construct. These response sets are properties of people, so the fixes live in the design of the whole questionnaire, not in any single item.
Acquiescence — "yea-saying" — is the tendency to agree with statements regardless of content. It's not laziness so much as politeness plus low effort: agreeing is the path of least resistance. Unchecked, it inflates scores on any scale where "agree" always means "high on the construct." The fix is reverse-coding: write some items so that disagreement indicates more of the construct ("I rarely feel warmth toward myself," on a self-compassion scale). A pure yea-sayer now agrees their way to a middling score, and their pattern becomes detectable.
Fence-sitting is the mirror habit: parking on the neutral midpoint to avoid commitment, especially on sensitive topics. The blunt fix is forced choice — remove the midpoint or use formats that require picking between options. The cost is real (some people genuinely are neutral, and you've just forbidden them from saying so), which is why methodologists treat midpoint removal as a judgment call rather than a commandment.
Socially desirable responding is the big one: editing answers at stage four to look good. Delbert Paulhus (1984) showed it isn't one thing but two. Impression management is deliberate self-presentation to an audience — faking good, on purpose. Self-deceptive enhancement is an honestly held but inflated self-view: the respondent isn't lying to you, they're passing along the flattering story they've already told themselves. The distinction matters because the standard remedies only reach the first kind.
Anonymity is genuinely powerful — but it only removes the audience, which means it only fixes impression management. Paulhus's (1984) two-component model predicts, and his data showed, that self-deceptive enhancement survives anonymity intact: people who sincerely believe they're unprejudiced, generous, and calm under pressure report exactly that to an anonymous form, because as far as they know it's true. That residue is a core reason researchers built indirect and implicit measures — instruments that don't route through the respondent's self-story at all.
Ask or watch?When self-report is fine — and when behavior disagrees
So should you trust self-report? The honest answer is: it depends on what you're asking about. For current subjective states — how anxious you feel, how much pain you're in, how satisfied you are with your job — the person is the only instrument with access, and well-built scales measuring such states behave well by every standard in Lesson 06. Self-report earns its keep constantly. The trouble starts in two specific places.
First: people cannot reliably report the causes of their own behavior. Nisbett and Wilson's (1977) famous review, "Telling more than we can know," collected demonstration after demonstration. In one, shoppers evaluated what they were told were four different nylon stockings and showed a strong preference for the right-most pair — the items were actually identical, and the preference was a pure position effect. Asked why they chose as they did, shoppers confidently cited quality, knit, and sheerness; when the position effect was suggested, they denied it, sometimes looking at the interviewer with concern. The unsettling conclusion: when you ask "why did you do that?", people don't introspect on the actual causal process — they generate a plausible theory about themselves and report that. The report is sincere. It is also, often, wrong.
Second: self-reported behavior and actual behavior can part ways, especially where the behavior is socially charged — diet, prejudice, exercise, drinking, screen time. Baumeister, Vohs, and Funder (2007) audited what personality and social psychology actually measures and found direct observation of real behavior had grown rare enough for them to rename the field, acidly, "the science of self-reports and finger movements." Their argument wasn't that self-report is worthless; it was that a psychology which never checks reports against conduct has quietly changed its subject matter. The methodological moral for you: when the claim is about what people do, find a way to measure doing.
One modern answer to both problems is to measure attitudes from behavior itself — usually millisecond-scale behavior. The Implicit Association Test family infers associations from how fast people can sort concepts that are paired compatibly versus incompatibly. The Affect Misattribution Procedure (AMP; Payne, Cheng, Govorun, & Stewart, 2005) briefly flashes an attitude object, then asks you to judge a neutral Chinese ideograph as pleasant or unpleasant — and the feeling evoked by the prime leaks into the judgment. Neither task asks you to describe yourself, so neither routes through impression management or your self-story. Your professor runs studies with exactly these instruments; you can try a live demonstration AMP at mwmagee.com/studies/amp/. Implicit measures have their own reliability and validity debates — they're instruments, not oracles — but they're the field's standing reply to "what if the ruler asks the object to measure itself?"
Watching wellObservation: coding schemes and interrater reliability
Observational research replaces the question with a trained eye — and immediately inherits measurement problems of its own. Raw watching isn't data. To turn behavior into numbers you need a coding scheme: an explicit rulebook that operationally defines each behavior of interest ("physical aggression = hitting, kicking, or shoving with apparent force; verbal protest ≠ aggression"), so that what counts is decided before observation begins, not improvised during it. A coding scheme is Lesson 06's operational definition, walking around in the field.
The test of a coding scheme is interrater reliability: two coders, independently applying the rulebook to the same behavior, should produce the same codes. In practice you check this by double-coding a portion of the data and computing agreement — and note that raw percent agreement flatters you, because two coders flipping coins agree a surprising amount by luck. The standard practice is a chance-corrected statistic such as Cohen's kappa, which asks how much better than coin-flipping your coders actually did. Low agreement doesn't mean firing a coder; it usually means the rulebook is ambiguous. You sharpen the definitions, retrain, and re-check.
But reliable coders can still be reliably biased — agreeing with each other while both drifting in the direction the hypothesis wants. This is observer bias, and its most theatrical demonstration has hooves. In early-1900s Berlin, a horse named Clever Hans appeared to do arithmetic, tapping out answers with his hoof to the astonishment of crowds and more than one scientific committee. The psychologist Oskar Pfungst (1907/1911) ran the tests that solved it: Hans succeeded only when his questioner knew the answer and was visible. The horse had learned to read microscopic, entirely unconscious cues — a relaxing of posture, a fractional head movement — that questioners emitted as his tapping approached the right count. Nobody was cheating. That's the point: expectancy leaks out of honest people through channels they cannot monitor.
Humans emit the same cues to human subjects. Rosenthal and Fode (1963) told student experimenters that their lab rats were bred "maze-bright" or "maze-dull." The labels were assigned at random — yet the "bright" rats learned faster, because students handled and scored them in subtly different ways. If expectations can move rat data, they can certainly move a coder's judgment of whether that playground shove "counted." The fix is the same one you'll meet again in the experiment lessons: masked (blind) coding — observers who don't know the hypothesis, the condition, or the group membership of who they're watching. Clever Hans went silent the moment his questioners were blinded; so does observer expectancy.
Being watchedReactivity — and how to measure without being noticed
Observation's second great problem runs the other direction: the observed change because they're observed. Reactivity is one of the deepest problems in behavioral science — the act of measuring alters the thing measured — and it has three working solutions. Habituation: stay long enough that you become furniture; field researchers from primatology to classroom studies rely on the fact that vigilance is expensive and fades. Blend the measurement into the setting so no discrete "being studied" moment exists. And, most elegantly, don't be there at all.
In the 1920s–30s studies at Western Electric's Hawthorne plant, researchers raised the lighting and productivity rose — then lowered the lighting, and productivity rose again. The workers weren't responding to lights; by the standard telling, they were responding to being studied. The Hawthorne effect became reactivity's most famous name, and it's why a promising intervention can shine in a study and evaporate in ordinary life: sometimes the effect was the observation, not the treatment. (Modern reanalyses debate how much the original plant data really show — a nice reminder that even methodological legends deserve interrogation.)
Webb, Campbell, Schwartz, and Sechrest (1966) wrote the classic manifesto for nonreactive research: measure the traces behavior leaves rather than the behavior itself. Their catalog is a delight — worn floor tiles around a museum exhibit as a popularity index; grime and dog-eared pages revealing which library chapters actually get read; garbage archaeology revealing what a household really consumes versus what it reports. Because no one knows they're being measured, there is nothing to react to — and nothing to socially-desirably edit. Unobtrusive measures solve reactivity and self-presentation in one stroke, which is why Webb's little book is still on methodologists' shelves sixty years later.
Two styles of watchingNaturalistic vs. structured observation
Observational designs sit on a continuum. Naturalistic observation records behavior in its ordinary habitat with minimal interference — high ecological validity, but you wait for the behavior to happen and control nothing. Structured observation stages a standardized situation and codes responses to it — every participant faces the same probe, at the cost of some artificiality. The developmental psychologist's "strange situation" for assessing infant attachment is the textbook structured case: same sequence of separations and reunions for every baby, coded with the same scheme.
| Naturalistic | Structured | |
|---|---|---|
| Setting | The behavior's real environment | Standardized, researcher-arranged situation |
| Strength | Ecological validity; behavior as it actually occurs | Comparability; every participant gets the same probe |
| Weakness | No control; rare behaviors mean long waits | Artificiality; the situation may not generalize |
| Shared requirements | A coding scheme, interrater reliability, masked coders, and a plan for reactivity | |
That factoid circulated for years in bestsellers and magazines, usually with no citable source. Then naturalistic observation did its job. Mehl, Vazire, Ramírez-Esparza, Slatcher, and Pennebaker (2007) equipped nearly 400 students with the Electronically Activated Recorder (EAR) — a device that periodically samples snippets of ambient sound as people live their lives — and simply counted. Women averaged about 16,215 words a day; men about 15,669. Statistically indistinguishable, and nowhere near a 3-to-1 ratio. Notice the method lesson stacked inside the gender lesson: no self-report of talkativeness (which stereotypes would contaminate at the judgment stage), minimal reactivity (you forget the recorder), and a behavior counted rather than estimated.
The modern updateData quality online, and observation in your pocket
The classic rules assume a human is honestly trying to answer. Now that most survey data arrives through platforms like Prolific and MTurk, that assumption needs enforcement. A fraction of online "participants" are bots, click-farms, or humans speed-running for the payment, and the toolkit has grown a gatekeeping layer: attention checks ("select 'strongly disagree' for this item"), seriousness checks, response-time flags for impossibly fast completion, CAPTCHAs and honeypot fields for scripts, and deduplication of suspiciously identical respondents. Notice these are the digital descendants of the classic tools — reverse-coded items and lie scales were always about catching respondents who weren't really answering; the "respondent" just wasn't at risk of being a script in 1985.
Ecological momentary assessment (EMA), also called experience sampling, pings participants at random moments to report what they're doing and feeling right now. It's a hybrid method: still self-report, but it amputates the biggest self-report failure point — retrospective estimation at the judgment stage. Instead of "how happy were you last week?" (a reconstruction), you catch the state in the moment, dozens of times, in the person's real environment. High ecological validity, minimal memory distortion; it's why so much of modern affect and clinical research lives on the phone.
SourcesCited in APA 7
Baumeister, R. F., Vohs, K. D., & Funder, D. C. (2007). Psychology as the science of self-reports and finger movements: Whatever happened to actual behavior? Perspectives on Psychological Science, 2(4), 396–403. https://doi.org/10.1111/j.1745-6916.2007.00051.x
Mehl, M. R., Vazire, S., Ramírez-Esparza, N., Slatcher, R. B., & Pennebaker, J. W. (2007). Are women really more talkative than men? Science, 317(5834), 82. https://doi.org/10.1126/science.1139940
Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259. https://doi.org/10.1037/0033-295X.84.3.231
Paulhus, D. L. (1984). Two-component models of socially desirable responding. Journal of Personality and Social Psychology, 46(3), 598–609. https://doi.org/10.1037/0022-3514.46.3.598
Payne, B. K., Cheng, C. M., Govorun, O., & Stewart, B. D. (2005). An inkblot for attitudes: Affect misattribution as implicit measurement. Journal of Personality and Social Psychology, 89(3), 277–293. https://doi.org/10.1037/0022-3514.89.3.277
Pfungst, O. (1911). Clever Hans (the horse of Mr. von Osten): A contribution to experimental animal and human psychology (C. L. Rahn, Trans.). Henry Holt. (Original work published 1907)
Rosenthal, R., & Fode, K. L. (1963). The effect of experimenter bias on the performance of the albino rat. Behavioral Science, 8(3), 183–189. https://doi.org/10.1002/bs.3830080302
Schwarz, N. (1999). Self-reports: How the questions shape the answers. American Psychologist, 54(2), 93–105. https://doi.org/10.1037/0003-066X.54.2.93
Schwarz, N., Hippler, H.-J., Deutsch, B., & Strack, F. (1985). Response scales: Effects of category range on reported behavior and comparative judgments. Public Opinion Quarterly, 49(3), 388–395. https://doi.org/10.1086/268936
Strack, F., Martin, L. L., & Schwarz, N. (1988). Priming and communication: Social determinants of information use in judgments of life satisfaction. European Journal of Social Psychology, 18(5), 429–442. https://doi.org/10.1002/ejsp.2420180505
Tourangeau, R., Rips, L. J., & Rasinski, K. (2000). The psychology of survey response. Cambridge University Press.
Webb, E. J., Campbell, D. T., Schwartz, R. D., & Sechrest, L. (1966). Unobtrusive measures: Nonreactive research in the social sciences. Rand McNally.