Contents/ Part XIII · Inference Case Studies/ Chapter 88

Case Study: A Clinical Trial

A randomized trial of a cholesterol drug raises three distinct questions, each needing a different test on the same patients: did LDL fall on the drug, did it fall more than placebo, and were more patients responders? We answer all three, and show why the placebo arm is what makes the verdict trustworthy.

⏱️ ~15 min read
🐍 Notebook included
📊 Chapter 88

Medicine is where careful inference matters most. A randomized controlled trial is the gold standard precisely because it pins an effect on the treatment and nothing else. This case study takes one trial and shows how a single dataset answers three different questions, each with the right tool.

Rx
The scenario. Patients were randomly assigned to a new LDL-lowering treatment or a placebo. LDL cholesterol was measured at baseline and again at 12 weeks. A patient is a responder if their LDL dropped by at least 15%.
🎯
The three questions

(1) Within the treatment arm, did LDL fall from baseline to week 12? (2) Did LDL fall more on treatment than on placebo? (3) Was the responder rate higher on treatment than placebo?

1

The Data & The Questions

One row per patient: trial arm, ldl_before, ldl_after, and a responder flag. The within-patient change (before − after) is positive when LDL drops.

📂 Dataset · case-study-a-clinical-trial--clinical_trial.xlsx

One row per patient with arm (treatment/placebo), age, ldl_before, ldl_after (mg/dL), and responder (0/1).

At a glance, treatment patients dropped about 29 mg/dL on average versus about 10 mg/dL on placebo, and more of them were responders. But a trial proves nothing until each claim is tested against the right null, and crucially, the placebo arm is what separates the drug's effect from everything else that makes numbers move in a trial.

2

Three Questions, Three Tests

The design of each question dictates its test. The same patients, structured three different ways.

QuestionStructureTestH₀ / H₁
Did LDL fall on treatment?paired (same patient, before vs after)paired t-testno change / LDL decreased
More than placebo?2 independent arms, numeric changetwo-sample (Welch) tequal change / treatment change larger
Higher responder rate?2 independent arms, yes/notwo-proportion zpₜ = pₚ / pₜ > pₚ

The first question is paired, each post value has a matching baseline on the same person, so a paired test isolates the within-patient change. Treating those columns as two independent samples (a classic mistake) would throw away the pairing and weaken the test. The second and third questions compare the two independent arms, one on a numeric change, one on a yes/no outcome.

3

The Analysis & Results

Each test comes with an effect size and a confidence interval, and the placebo comparison reveals how much of the raw drop is really the drug.

Average LDL drop: placebo vs treatment 10 mg/dL placebo 29 mg/dL treatment +18 from the drug ~10 happens anyway
QuestionResultVerdict
Q1 paired (treatment)drop ≈ 29 mg/dL, t = −19.7, p ≈ 10⁻³⁴LDL fell, decisively
Q2 two-arm changeextra drop ≈ 18 mg/dL, t = 8.45, p ≈ 10⁻¹⁴, CI [14, 23]more than placebo
Q3 responder rate52% vs 17%, z = 4.80, p ≈ 10⁻₆, gap CI [+22, +48] ptsfar more responders

All three agree, and the placebo arm is the hero of the story. Treatment patients dropped 29 mg/dL, but placebo patients dropped about 10 mg/dL on their own (diet, measurement timing, the trial effect). So only the ~18 mg/dL difference can be credited to the drug. Without the control arm, the full 29 would have looked like the treatment's effect, a textbook reminder of why randomized comparisons matter.

4

The Statistician's Report

How to summarize for a clinical team or a non-statistician reviewer.

📋 Statistician's report · to the clinical team

Conclusion: the treatment works, on all three measures

What we found. Patients on the treatment lowered their LDL cholesterol by about 29 mg/dL over 12 weeks. The part we can credit to the drug, beyond what placebo patients dropped on their own, is about 18 mg/dL. And 52% of treated patients were responders (a 15%+ reduction) versus only 17% on placebo.

How confident are we? Extremely. Every comparison is far beyond the threshold for chance. Our 95% range for the drug's extra benefit is about 14 to 23 mg/dL, a clinically meaningful reduction even at the low end.

Why the placebo arm mattered. Placebo patients also improved by about 10 mg/dL. Comparing only before-versus-after on the treatment would have over-credited the drug by that amount. The randomized, controlled comparison is exactly what isolates the true effect, this is why single-arm before/after studies can be misleading.

Caveats. This measures a 12-week surrogate marker (LDL), not long-term cardiovascular outcomes. Safety, side effects, and durability are separate questions that a full trial program must answer before any clinical claim.

🤖
In the field

The paired-design lesson carries straight into machine learning: comparing two models on the same test cases is paired, and a paired test on the per-case differences is far more sensitive than treating the two score lists as independent, just as the paired before/after test is here.

🐍

Run all three trial tests in Python

The companion notebook explores the data with a baseline-balance check (confirming randomization), then runs the paired t-test, the two-sample Welch t-test, and the responder-rate test, using statsmodels (CompareMeans, proportions_ztest) for the intervals and the proportion test, with plots by arm.

📓 View Notebook (code & outputs) ▶ Open in Colab ⬇ View / Download on GitHub

View opens the rendered notebook instantly (no setup). Open in Colab runs & edits it live in your browser. To run locally, install numpy, pandas, scipy, matplotlib, statsmodels, and openpyxl and launch jupyter notebook.

🎓 Key Takeaways

  • One dataset, three designs: paired (within patient), two-sample (between arms), and a proportion (responder rate).
  • Paired before/after needs a paired test; treating it as independent wastes the design.
  • The placebo arm isolates the drug: 29 mg/dL total, but only ~18 mg/dL beyond placebo is the treatment effect.
  • Results: all three highly significant; responders 52% vs 17%; 95% CI for extra LDL drop [14, 23] mg/dL.
  • Caveat: a 12-week surrogate marker, not long-term outcomes; safety and durability are separate questions.
5

Quiz: Test Yourself

Eight quick questions on this case study. Answer them, hit Check Answers, and keep refining until you score 100%. Your progress is saved, so you can hop back to the chapter and return anytime.

➡️
Up next

Numbers were the easy part. A Customer Satisfaction Survey handles messier data, categories, ratings, and a top-box rate, with chi-square, a margin of error, and an ordinal test on one survey.