Test the things every doctor believes: two-fifths fail. Audit the field everyone mocks: most of it stands. Tonight, nobody gets to round to zero or one hundred.
OC ACXLW Meetup #118 — The 40% Problem & The OC ACXLW Meetup #118 — The 40% Problem & The Field That Refused to DieField That Refused to Die
Date: Saturday, August 15, 2026
Time: 2:00 PM
Location: 1970 Port Laurent Place, Newport Beach, CA 92660
Hello folks! This week is a two-part exercise in calibration. First: what happens when someone actually tests the things all the experts agree on — using medicine, the field with the most money, the most regulation, and the best trials, as the test case. Second: what happens after a field gets debunked — using psychology, where the replication-crisis pile-on is now itself being rolled back, and the rollback is due for some scrutiny of its own. Watch for the thread running under both: in almost every case, what propped the practice up was a story so plausible nobody thought to check it — but that same kind of story also produced germ theory and cracked the thymus scandal open, so the live question is when a mechanistic story is doing science and when it's doing damage. Come ready to say how you weighed expert consensus in each field before the readings, and how you weigh it after — the two topics should pull you in opposite directions, and noticing whether they did is the evening's real exercise.
Conversation Starter 1: The 40% Problem — Medical Reversal as a Base Rate
Everyone knows a story where medicine got it wrong. The interesting question is the rate. One research group counted.
The number: Prasad and colleagues reviewed every original article across ten years (2001–2010) of one high-impact medical journal. Of 363 articles that tested an established practice, 146 (40.2%) reversed it and 138 (38.0%) reaffirmed it. When someone bothers to rigorously test standard of care, it is close to a coin flip (Prasad et al., Mayo Clinic Proceedings).
The gallery behind the number — a sampler of practices that were confident, standard, and wrong, sorted by how they failed: internal mammary artery ligation for angina, killed by sham-surgery trials in 1959–60 when sham incisions relieved chest pain just as well (subjective endpoints); infant thymus irradiation, built on cadaver reference norms drawn from the chronically ill poor, whose stress-shrunken thymuses made healthy glands look enlarged — survivors carried ~5.5× thyroid cancer risk decades later (corrupted baseline); stenting for stable angina, which survived COURAGE (2007) only to meet the sham-controlled ORBITA trial (2017) — the same trap and same solution as mammary ligation, 58 years apart (mechanistic plausibility over outcomes); and the era when cocaine, morphine, amphetamine diet pills, barbiturates, and "non-addictive" Valium were routine prescriptions (deferred harm the endpoints weren't watching).
The one that reversed 180° within living memory: after a heart attack, patients in the 1930s–50s were ordered onto 4–6 weeks of strict bed rest — in some regimens forbidden even the radio or a newspaper. Levine warned about the harms of recumbency in JAMA in 1944; he and Lown published "armchair treatment" in 1952 and were met with strong opposition; today early mobilization and cardiac rehab are the standard, and it's the bed rest that's considered dangerous. The evidence preceded the practice change by decades (Mampuya, Cardiovasc Diagn Ther; Levine & Lown, JAMA 1952).
Why this matters right now: we are mid-adoption of the fastest-spreading drug class in history, and the question the gallery teaches you to ask isn't "were the GLP-1 agonists tested?" but "tested against which failure mode?" The registration trials are three-to-five-year cardiovascular outcome studies in high-risk populations — which genuinely closes the subjective-endpoint trap that killed mammary ligation and the surrogate-endpoint trap that caught stenting. They say nothing about the mode that killed Valium and the amphetamine diet pills: deferred harm on endpoints nobody was watching — and note that Valium's addiction profile would have sailed through a three-year cardiac-events trial untouched. The early post-marketing signals (amotivation case reports, a tail of patients losing more than 15% of their lean mass) surfaced outside the trials precisely because the trials weren't instrumented to see them, and phase 4 surveillance is only a few years old. Meanwhile prescribing has already expanded far past the trial populations — the same indication creep that carried bypass surgery from left-main disease, where it earned its evidence, to everyone with chest pain. None of this says the peptides are 1960 redux; it says "safe and effective" is a claim about the failure modes we checked, and every reversal in tonight's gallery lived in one we didn't (SELECT and related outcome trials; 2026 post-marketing body-composition and case-series reports).
Questions for discussion:
The denominator problem, both directions. Prasad's 40.2% comes from practices someone chose to test — presumably the suspicious ones — so it overstates the reversal rate of medicine as a whole. But most practices never get tested at all, so the untested mass hides its own failures indefinitely. Net those against each other: what's your actual estimate for the fraction of today's standard-of-care that would reverse under rigorous testing, and what evidence sets it?
Sort the gallery, find the open flank. The cases fail in distinguishable ways: subjective endpoints (mammary ligation, lobotomy), corrupted baselines (thymus), surrogate endpoints and mechanism-worship (stenting), deferred harm (Valium, steroids), and sheer inertia against existing evidence (bed rest, 1944→1970s). Which failure mode have modern methods — sham controls, preregistration, hard endpoints — genuinely closed, and which is still wide open? Then the harder version: nearly every case above was propped up by a compelling mechanistic story — unblock the pipe, rest the wounded heart, the swollen gland is choking the baby. But mechanistic stories also did the debunking: the thymus scandal broke when someone told a better story about where the reference cadavers had come from, and H. pylori causing ulcers was dismissed as a just-so story for years before it turned out true. So how do these stories mislead us, and when do they genuinely aid understanding and generate the hypotheses that overturn the error? If you can't name a test you could have applied at the time, you're conceding that causal plausibility carries no evidential weight at all in domains this tangled — a bigger claim than it sounds, and you have to live with it.
Who pays to test what everyone already believes? Reversal requires funding a trial of something no one doubts, randomizing patients away from standard care, and publishing against the field's consensus — three things nobody has an incentive to do. Forced tradeoff: a system of mandatory periodic re-testing of established practices (expensive, ethically awkward, slows adoption) versus the status quo (40% of tested practices wrong, discovered by accident). Which error do you buy, and at what price?
Conversation Starter 2: The Debunking Gets Debunked — and the Rebunking Needs Scrutiny Too
Psychology spent a decade as everyone's favorite example of failed science. Scott Alexander just published a defense — then walked part of it back within a week, in public, crediting his commenters. Both moves deserve a hard look, and so does the walk-back.
Scott's defense: the replication crisis was concentrated in social psychology, and most centrally in social priming — by his count, the 0.05–0.2% of psychologists doing priming studies plus some fraction of the 10–20% doing social psych. Mocking "psychology" for it indicts the psychophysicists, psychometricians, behaviorists, and clinicians who had nothing to do with it; social psych is one chapter of sixteen in the intro textbook (Scott Alexander, ACX).
Scott's revision, one week later: commenters pointed to failed replications outside social psychology, and Scott conceded the subfield partition was the wrong frame — the real line runs between settled science and novel hyped findings, in any field. Note what sits on his own settled list: benzodiazepines "work but produce addiction and tolerance." That is 2026's settled science. In 1975, the settled science was that Valium wasn't addictive. The item that connects tonight's two topics is sitting right there (Scott Alexander, ACX).
The counter-case we'll argue: Scott weights contamination by share of the discipline — textbook chapters, practitioner headcount. Weight instead by contact surface: the public meets psychology through policy (social psych's export) and the clinic (personality theories, therapy schools), not through the psychophysics lab. And the clinic's crown jewel has soft spots the field's own auditors report: adjusting for publication bias drops CBT's effect for major depression from g = 0.75 to 0.65 and for generalized anxiety from 0.80 to 0.59 (Cuijpers et al., World Psychiatry), a 2025 review of recent trials finds adjusted effects fall to small (g ≈ 0.27) and notes outcomes are centered almost entirely on symptom scales — the absence of measured pathology being quietly equated with recovery of actual functioning (J Affect Disord 2025). If your depression measure samples negative verbal self-accounts, and your therapy trains more positive verbal self-accounts, Goodhart's law predicts improvement on the measure regardless of what happened to the life. And underneath the measurement problem sits a story problem: CBT rests on a causal narrative — distorted cognitions cause depression, correct them and the depression lifts — while Scott's own settled list concedes that psychotherapy works mostly through the relationship with the provider rather than the specifics of the modality. If that concession holds, the mechanism story is wrong even where the treatment works, which is a harder indictment than any effect size because it survives the effect sizes holding up. Though it also raises the question the medical half raises: is a wrong story attached to a working treatment a failure of science, or an ordinary stage of it?
A complication that cuts both ways: whether CBT's effects are declining over time is genuinely contested — Johnsen & Friborg found decline; Cristea and Cuijpers reanalyzed and found it unstable and mostly absent. Meanwhile a meta-meta-analysis finds effect declines run ~2:1 across psychology generally, including in the "trustworthy" subfields like intelligence research. Declining effects don't uniquely indict CBT — but ubiquitous decline mildly indicts Scott's original "the other chapters are fine" partition, while supporting his revised settled-vs-novel one. Even the criticisms need criticism.
Questions for discussion:
Denominator wars. Scott counts by chapters and headcount; the counter-case counts by exposure — clinic and policy. Notice this is the mirror image of Prasad's problem (his denominator overstates; Scott's understates). For the practical question "how much should I trust psychology?" — meaning your therapist, a parenting claim, a policy nudge — which weighting is the right one, and what number does it give you?
Does "evidence-based therapy" survive the common-factors concession? Scott's own settled list says psychotherapy works at moderate effect size but mostly via the relationship with the provider rather than the modality. If the mechanism-specific claims are weak, what is the "evidence-based" label actually certifying? Name the outcome measure — functioning, informant report, employment, anything — that would convince you CBT works beyond teaching the test, and say whether anyone is collecting it.
Five updates deep. The sequence so far: replication crisis → "psychology is fake" → Scott's defense → Scott's revision → tonight's counter-case. Each step corrected the last and introduced its own selection effects. State where your probability on "psychology research is mostly fine" sits right now, and name the single observation that would move it most. If your answer is 0 or 100, you've misunderstood the last decade.
The Calibration Discussion: Before and After
No worksheets — this one runs on honest self-report, in three rounds:
Before. Think back to before you read tonight's articles. How much did you trust medical consensus — the thing your doctor tells you is standard of care? And how much did you trust psychology — or had the replication crisis already written it off for you? Say where you actually stood, not where you'd like to have stood.
After. Did the readings move you, and in which direction? Note that the two topics push opposite ways — the medical material should lower trust in settled consensus, the psychology material should raise trust in a "debunked" field. If you moved in the same direction on both, or didn't move at all, that's worth examining out loud.
Forward. Name a current expert trend — in medicine, psychology, or adjacent — that you think is mid-arc right now: somewhere between confident adoption and eventual correction. Where does it land in ten years, and what's your reasoning? (The peptide boom, the anti-CBT correction, and the replication-reform movement itself are all fair game — including the possibility that today's debunkings are tomorrow's reversals.)
Bridge Prompt
The two topics teach the same lesson from opposite ends. Medicine is the field where full expert consensus should not have taken you to 100 — test the consensus and 40% of it fails. Psychology is the field where scandal should not have taken you to 0 — audit the wreckage and most of the field stands. And both correction processes were themselves flawed: Prasad's reversals come from a suspicious sample, Scott's defense needed a one-week retraction, and the retraction needs tonight's counter-arguments. So: is there any stopping point in this regress where you're entitled to just believe the current answer — or is "never 0, never 100, always ready to update" not a slogan but the literal steady state of an honest epistemic life? And if it is: how do you act decisively — take the drug, pick the therapy, fund the study — while living there?
After the meeting
Walk & Talk: we usually do an hour-long walk-and-talk after the meeting starts. Two mini-malls with hot takeout food are readily accessible nearby — search for Gelson's or Pavilions in the zip code 92660.
Share a Surprise: tell the group about something that unexpectedly changed your perspective on the universe.
Future Directions: contribute ideas for the group's future — topics, meeting formats, activities, etc.
Test the things every doctor believes: two-fifths fail. Audit the field everyone mocks: most of it stands. Tonight, nobody gets to round to zero or one hundred.
OC ACXLW Meetup #118 — The 40% Problem & The OC ACXLW Meetup #118 — The 40% Problem & The Field That Refused to DieField That Refused to Die
OC ACXLW Meetup #118 Announcement https://docs.google.com/document/d/1aAV-cwiVjL70LRjc7QReHeTd89MIFqycpg21Lw5fg14/edit?usp=sharing
Hello folks! This week is a two-part exercise in calibration. First: what happens when someone actually tests the things all the experts agree on — using medicine, the field with the most money, the most regulation, and the best trials, as the test case. Second: what happens after a field gets debunked — using psychology, where the replication-crisis pile-on is now itself being rolled back, and the rollback is due for some scrutiny of its own. Watch for the thread running under both: in almost every case, what propped the practice up was a story so plausible nobody thought to check it — but that same kind of story also produced germ theory and cracked the thymus scandal open, so the live question is when a mechanistic story is doing science and when it's doing damage. Come ready to say how you weighed expert consensus in each field before the readings, and how you weigh it after — the two topics should pull you in opposite directions, and noticing whether they did is the evening's real exercise.
Conversation Starter 1: The 40% Problem — Medical Reversal as a Base Rate
Everyone knows a story where medicine got it wrong. The interesting question is the rate. One research group counted.
Text (read, ~9 pages): A Decade of Reversal: An Analysis of 146 Contradicted Medical Practices (Prasad, Vandross, Cifu et al., Mayo Clinic Proceedings, 2013) https://www.mayoclinicproceedings.org/article/S0025-6196(13)00405-9/fulltext
Audio (listen, ~1 hr): Adam Cifu on Ending Medical Reversal (EconTalk with Russ Roberts) https://www.econtalk.org/adam-cifu-on-ending-medical-reversal/
Optional companion (same journal issue): Ioannidis, "How many contemporary medical practices are worse than doing nothing?" (Mayo Clinic Proceedings, 2013) https://www.mayoclinicproceedings.org/article/s0025-6196(13)00403-5/fulltext
Summary:
Questions for discussion:
Conversation Starter 2: The Debunking Gets Debunked — and the Rebunking Needs Scrutiny Too
Psychology spent a decade as everyone's favorite example of failed science. Scott Alexander just published a defense — then walked part of it back within a week, in public, crediting his commenters. Both moves deserve a hard look, and so does the walk-back.
Text (read, short): Psychology Research Is Mostly Fine (Scott Alexander, Astral Codex Ten) https://www.astralcodexten.com/p/psychology-research-is-mostly-fine
Text (read, short): The Beauty Of Settled Science — Scott's follow-up and partial retraction https://www.astralcodexten.com/p/the-beauty-of-settled-science
Text (read) — the field auditing itself: How effective are cognitive behavior therapies for major depression and anxiety disorders? A meta-analytic update (Cuijpers et al., World Psychiatry, 2016) https://onlinelibrary.wiley.com/doi/full/10.1002/wps.20346
Audio (listen): both ACX posts have Substack's built-in narration — same links as above, press play on the page or in the app.
Optional deeper reading: Cognitive behavioral therapies are evidence-based — based on what? (Journal of Affective Disorders, 2025) https://www.sciencedirect.com/science/article/pii/S0165032725012339
Optional — the decline-effect dispute: Battle of the meta-analyses: is CBT becoming less effective over time? (The Mental Elf, on Johnsen & Friborg vs. Cristea) https://www.nationalelfservice.net/treatment/cbt/battle-of-the-meta-analyses-is-cbt-becoming-less-effective-over-time/
Summary:
Questions for discussion:
The Calibration Discussion: Before and After
No worksheets — this one runs on honest self-report, in three rounds:
Bridge Prompt
The two topics teach the same lesson from opposite ends. Medicine is the field where full expert consensus should not have taken you to 100 — test the consensus and 40% of it fails. Psychology is the field where scandal should not have taken you to 0 — audit the wreckage and most of the field stands. And both correction processes were themselves flawed: Prasad's reversals come from a suspicious sample, Scott's defense needed a one-week retraction, and the retraction needs tonight's counter-arguments. So: is there any stopping point in this regress where you're entitled to just believe the current answer — or is "never 0, never 100, always ready to update" not a slogan but the literal steady state of an honest epistemic life? And if it is: how do you act decisively — take the drug, pick the therapy, fund the study — while living there?
After the meeting
Posted on: