psySC › Modules › Do Interventions Work?
Module · read before session 14

Do Interventions Work?

The bill comes due: the MSC trial, the brief doses, the meta-analytic diamond — and the critiques that shrink it without erasing it. Plus the mirror this module holds up to your own project.

Why this is the week the translation project pays off — or doesn't

Recall the pipeline from the roots modules: practice → clinic → construct → scale → RCT. Everything before the arrow's final step was preparation. A construct with a scale can only show you correlations, and correlations — as you've now said in six different weeks — cannot adjudicate arrows. To claim self-compassion causes wellbeing, someone must give people self-compassion, at random, and watch. That is what this module's studies do, and reading them well requires exactly one skill above all others: interrogating the control condition. In intervention science, the control group is where the argument lives. Everything else is arithmetic.

The study, up close

Neff, K. D., & Germer, C. K. (2013). A pilot study and randomized controlled trial of the Mindful Self-Compassion program. Journal of Clinical Psychology, 69, 28–44.

Structure: first a pilot (does the program run — attendance, completion, do scores move?), then the real thing: community adults randomly assigned to the eight-week MSC program or a waitlist. Findings: the MSC group showed significant gains in self-compassion, mindfulness, and life satisfaction, and reductions in depression, anxiety, and stress; the waitlist stayed flat. Compassion-for-others measures moved less than self-directed ones — an echo of week ten's 'two separate systems' finding. And the result the field still cites: gains held at six-month and one-year follow-ups. Whatever moved, stayed moved — which matters because feel-good workshop effects that evaporate by finals week are the industry norm this trial had to beat.

Design anatomy — recite this: a waitlist controls for the passage of time, retesting effects, and regression to the mean (extreme scorers drifting back toward average on remeasurement). It does NOT control for expectancy (believing you're receiving something helpful), attention and social contact, group belonging, or the generic benefits of doing anything structured and hopeful with other humans for eight weeks. MSC beat nothing. That is a real achievement — most things don't — but it is not the same as beating something, and the difference is the seam every serious critique runs through.

The brief-dose question

Eight weeks is a real commitment. The obvious applied question: how little is enough? Two studies frame the answer, and one of them quietly outclasses the flagship on the design dimension that matters most.

The study, up close

Smeets, E., Neff, K., Alberts, H., & Peters, M. (2014). Meeting suffering with kindness: Effects of a brief self-compassion intervention for female college students. Journal of Clinical Psychology, 70, 794–807.

Design: three weeks, three group meetings — psychoeducation, the core practices, phrase work — for female college students. The control group received time-management training: same schedule, same group format, same sense of receiving something useful. That is an active control, and it changes what a win means: both groups got attention, structure, expectancy, and belonging — so a difference between them isolates something closer to the self-compassion content itself. Findings: the self-compassion group showed greater gains in self-compassion, optimism, and self-efficacy, with decreased rumination. The sentence this buys: not 'better than nothing' but 'better than a credible something.' Hold every future trial you read to this standard and you will be ahead of a distressing fraction of the published literature.

Dundas and colleagues (2017) pushed the dose lower still: a two-week, three-session course for students produced gains in healthy self-regulation and reductions in self-judgment. The pattern across the dose-response literature: the signal keeps showing at shrinking doses, with (unsurprisingly) smaller and less durable effects than full programs. For your purposes, note what this implies about the fourteen single-subject trials running in this room: a thirteen-week semester of guided practice is, by this literature's standards, a long, high-dose intervention. If the literature is right, something in your data should move. Whether you can trust what moves — hold that for the mirror.

Test your knowledge

Checkpoint 1 of 2 · Answer from memory first — every option gives feedback.

A waitlist control CANNOT rule out:

The MSC RCT's most impressive feature:

Smeets et al. (2014) beat a time-management control. Why is that a 'stronger sentence'?

The view from orbit

The study, up close

Ferrari, M., et al. (2019). Self-compassion interventions and psychosocial outcomes: A meta-analysis of RCTs. Mindfulness, 10, 1455–1473.

What a meta-analysis is, in one line: convert every eligible trial's result onto a common ruler (the standardized effect size), pool the results weighting by precision, and read the diamond at the bottom of the forest plot — the field's best single estimate. What this one found, across dozens of randomized trials of self-compassion-based interventions: large pooled effects on self-compassion itself, and moderate pooled effects on depression, anxiety, stress, rumination, and self-criticism, with benefits across clinical and non-clinical samples. How to read a forest plot when you meet one: each row is a study — a square (the estimate) with whiskers (its confidence interval); the diamond at the bottom is the pooled answer; if the diamond doesn't cross zero, the field's answer isn't nothing. The standard moderator finding: effects tend to be larger against passive controls than active ones — which is not a scandal but a measurement of exactly the expectancy-and-attention gap the Smeets design closed.

The critiques, at full strength

Waitlist inflation. Most trials in the pool used waitlists; you now know precisely what those fail to control. Effects against active comparators run smaller. The honest reading treats the diamond as an upper bound. The self-report problem, sharpened. Nearly all outcomes are questionnaires — completed by people who just spent weeks learning what self-compassion is and that they were supposed to gain some. This is not fraud; it's demand characteristics plus genuine conceptual drift (after training, you interpret the items differently — a subtle validity problem called response-shift). The strongest answer would be behavioral and observer-rated outcomes; the field has some (recall Sbarra's coders, Berry's scanner) but not enough. Volunteer samples. People who sign up for compassion training are not random humans — they're motivated, curious, and disproportionately open. Generalization to the skeptical, the coerced, and the merely-enrolled-for-credit is an open question this classroom is, awkwardly, currently testing. Publication bias. Null trials are harder to publish; meta-analyses correct for this imperfectly. None of these critiques erases the diamond. All of them shrink and blur it. The calibrated conclusion, exam-ready: self-compassion interventions reliably increase self-compassion and meaningfully improve distress outcomes; true effects are probably somewhat smaller than headline numbers, and the strongest designs still find them.

The mirror

Now the module's real assignment. Take every critique above and aim it at your own semester project. N = 1. No control condition of any kind — not even a waitlist (what would that be? a version of you that didn't take the course? that person doesn't exist to measure). Outcomes: self-report, from a reporter who has read the literature and knows the hypotheses. Researcher allegiance: total — the investigator hopes the intervention works, IS the participant, and is being graded on engagement. Expectancy: maximal. Every flaw the field's weakest trials have, your study has more of.

And now the move that separates December's strongest talks from the rest: this is not a reason for despair — it's a reason for calibration. Your journal contains something no RCT contains: dense, longitudinal, first-person process data — the texture of a practice landing or failing to land, mechanisms felt from the inside, the exact Tuesday the phrase started sounding like your own voice. That is real signal. It simply isn't proof, and it can't be. The students who say this out loud, unprompted, with specifics — "here are the three confounds my design cannot exclude" — are doing the thing this course was secretly always about: holding warmth and rigor in the same hand.

Objections and complications

One more objection, because a sharp student always raises it: "if expectancy explains part of the effect, isn't that fine? Placebo relief is still relief." Two answers. Pragmatically: partly yes — if believing-in-the-practice helps people suffer less, that's not nothing, and clinical medicine quietly banks placebo components all the time. Scientifically: no — because theories make mechanistic claims (threat-system downregulation, reduced rumination), and a field that can't separate its mechanism from its marketing can't improve its interventions, only its advertising. Both answers can be true at once; keeping them distinct is precisely the skill this module is for.

🌱
Bodhi says: Before class, do the mirror exercise on paper: list your project's design flaws in one column and what your journal can still legitimately claim in the other. Bring it — session 14's discussion is built on exactly that table, and it becomes your talk's caveat slide almost verbatim.

Key terms

TermWorking definition
Waitlist controlComparison group receiving nothing yet; controls time, retesting, regression to the mean — not expectancy, attention, or belonging
Active controlComparison group receiving a credible alternative activity; isolates content-specific effects from generic ones
Effect sizeA standardized ruler (e.g., group difference in SD units) letting different studies' results be pooled
Forest plot / diamondMeta-analysis graphic: one row per study, pooled estimate as the diamond at the bottom
Demand characteristicsCues telling participants what response is expected — inflating self-reported change
Response shiftPost-training reinterpretation of what scale items mean — a subtle threat to pre/post comparison
Researcher allegianceThe investigator's investment in one outcome — in your project, absolute, and worth naming aloud

Test your knowledge

Checkpoint 2 of 2 · Answer from memory first — every option gives feedback.

A meta-analytic effect size is:

Ferrari et al. (2019), pooled: ___ effects on self-compassion, ___ on distress outcomes.

Which critique applies to YOUR semester project?

Check yourself

Reflective questions — no clicking, just thinking. These are the kinds of questions that show up in class discussion and, in mutated form, on exams.

What does a waitlist rule out, what survives it, and why does Smeets matter?

Rules out time, retesting, regression; expectancy, attention, and belonging survive. Smeets's active control (time management) matched those, so its win isolates something closer to the content.

Give the calibrated conclusion on the intervention literature in one sentence.

Interventions reliably raise self-compassion and moderately improve distress; true effects likely run somewhat below headline estimates once waitlist inflation, self-report, and volunteer sampling are priced in.

What can your journal legitimately claim that an RCT cannot — and vice versa?

The journal: dense first-person process data — texture, timing, felt mechanism. The RCT: causal license via randomization and comparison. Signal vs. proof; a strong talk holds both.