Statistical analysis

Psychosocial Treatments for Cocaine Dependence

Summary

487 patients were randomized across four arms. At month 3, 292 had a usable outcome and baseline; at month 6, 329. Adjusting for baseline, the arms differ at month 3 (p = .0110) and do not at month 6 (p = .2580).

Two things qualify that. Only 60% and 68% of randomized patients reach the two analyses, though the loss is even across arms. And the normality violation reported at month 6 does not survive being tested on the residuals of the model that was actually reported.

1. Who is in each analysis

This is a six-month trial in cocaine dependence, so retention is the first thing to establish. Roughly a third of randomized patients are missing from each cross-section.

How many randomized patients reach each analysis Of 487 randomized. Bars show the share with a usable score at that month. Month 3 Month 6 IDC+GDC 56% (68/121) 63% (76/121) CT+GDC 62% (74/119) 72% (86/119) SE+GDC 57% (71/124) 68% (84/124) GDC 64% (79/123) 68% (83/123)
Figure 1. Share of randomized patients with a usable score at each month.
Retention by arm. Analysed means a non-missing outcome and baseline.
ArmRandomizedMonth 3 %Month 6%
IDC+GDC1216856%7663%
CT+GDC1197462%8672%
SE+GDC1247157%8468%
GDC1237964%8368%

Retention does not differ significantly by arm at either timepoint (chi-squared p = .52 at month 3, p = .48 at month 6). That is the reassuring version: the arms are not losing patients at different rates, so the comparison is not obviously distorted.

It is not the whole story. Patients who stay in an addiction trial differ from those who leave in ways a balance test cannot see, and the month 6 group is not the month 3 group. A treatment effect that appears at one timepoint and not the other could reflect the treatment, or it could reflect who was still there to measure. These data cannot separate those.

2. What the data looks like before modelling

Before any model, it is worth seeing the distributions. Scores are wide and heavily overlapping in every arm, which is the honest starting picture: whatever treatment does here, it does not separate the groups cleanly.

Behaviour score by arm, at both timepoints Box shows the middle half, line the median, diamond the mean, whiskers the range. Month 3 Month 6 20 30 40 50 60 70 IDC+GDC n = 68 / 76 CT+GDC n = 74 / 86 SE+GDC n = 71 / 84 GDC n = 79 / 83 Behaviour score
Figure 2. Behaviour score by arm at both timepoints. Diamond is the mean, line the median.
Descriptive statistics for the outcome, by arm and timepoint.
MonthArmnMean SDMedianIQRRange
3CT+GDC7448.788.964943 to 5528 to 70
3GDC7949.499.784942 to 5725 to 67
3IDC+GDC6852.067.465046 to 5737 to 66
3SE+GDC7148.698.624744 to 5429 to 68
6CT+GDC8648.179.394840 to 5527 to 66
6GDC8349.369.365044 to 5827 to 68
6IDC+GDC7650.048.604943 to 5725 to 68
6SE+GDC8448.048.284642 to 5433 to 66

Two things stand out. IDC+GDC has the highest mean at month 3 and the ordering is visible, but the boxes overlap almost entirely, so the effect that follows is a shift in the middle of wide distributions rather than a separation of groups. And by month 6 even that ordering has largely gone.

3. The primary comparison

Each timepoint is analysed by ANCOVA: outcome on treatment arm, adjusted for the patient's own baseline score. Differences below are against group drug counselling alone.

ANCOVA at each timepoint. Arm rows show the adjusted difference from GDC with a 95% interval.
Model / contrastn F95% CIp Baseline p
Month 32920.34337.45<.0001.0110<.0001
  IDC+GDC vs GDC+2.500.15 to 4.86.0368
  CT+GDC vs GDC-1.06-3.36 to 1.24.3648
  SE+GDC vs GDC-1.07-3.40 to 1.25.3640
Month 63290.25227.34<.0001.2580<.0001
  IDC+GDC vs GDC+0.35-2.07 to 2.78.7747
  CT+GDC vs GDC-1.43-3.78 to 0.92.2312
  SE+GDC vs GDC-1.63-4.00 to 0.73.1754

At month 3 the model explains 0.343 of the variance and treatment matters (p = .0110). At month 6 it explains 0.252 and treatment does not (p = .2580). Baseline is overwhelmingly significant at both. Note what that means: most of what the model explains is the patient, not the therapy.

4. Assumptions, and a check that was run on the wrong model

The original analysis tested residual normality and found a violation at month 6, then investigated a Box-Cox transformation and concluded none was needed. Reproducing that turned up something worth reporting.

The normality check was run on residuals from a model containing treatment only. The model actually reported includes the baseline covariate. Those are different models with different residuals, and the assumption belongs to the model you report.

Which residuals were tested for normality Shapiro-Wilk p. Below .05 the test rejects normality. Higher is better. Month 3, model without the covariate p = 0.2668 Month 3, the reported ANCOVA p = 0.1238 Month 6, model without the covariate p = 0.0047 rejects Month 6, the reported ANCOVA p = 0.4285 .05
Figure 3. Shapiro-Wilk p for each set of residuals. Only the model without the covariate rejects.

At month 6 the residuals of the treatment-only model reject normality (p = .0047), while the residuals of the reported ANCOVA do not (p = .4285). At month 3 neither rejects (.2668 and .1238). So the violation that prompted the transformation work was a property of a model that was not being reported. The reported models satisfy the assumption.

Variances are homogeneous across arms at both timepoints (Levene p = .30 and p = .37).

This does not overturn anything. The Box-Cox investigation concluded no transformation was needed, which is the same place this ends up. It is worth recording because the reasoning is what a reader would rely on, and the reasoning tested the wrong residuals.

5. Subgroups, and how much weight they hold

The analysis also tested whether treatment effects depend on gender and race, through two-way and three-way interaction models. A striking difference appears at month 6, and it needs its denominators attached.

Every subgroup at month 6, with the patients behind it Adjusted means for all 16 arm x gender x race cells. Intervals widen as cells shrink. 40 45 50 55 60 CT+GDC F Non-Cauc. n=10 43.43 SE+GDC M Cauc. n=29 47.08 SE+GDC M Non-Cauc. n=36 47.08 CT+GDC M Cauc. n=33 48.22 IDC+GDC F Cauc. n=9 48.33 IDC+GDC M Cauc. n=25 48.52 CT+GDC M Non-Cauc. n=35 49.00 GDC M Non-Cauc. n=33 49.02 GDC F Cauc. n=10 49.05 IDC+GDC M Non-Cauc. n=28 49.08 GDC F Non-Cauc. n=11 49.39 CT+GDC F Cauc. n=8 49.95 SE+GDC F Cauc. n=8 50.30 GDC M Cauc. n=29 50.45 SE+GDC F Non-Cauc. n=11 51.33 IDC+GDC F Non-Cauc. n=14 55.19 Adjusted behaviour score
Figure 4. All 16 subgroup cells at month 6, with the number of patients in each.

The highest adjusted mean is IDC+GDC at 55.19, the lowest is CT+GDC at 43.43, both among non-Caucasian women. That is a gap of 11.8 points, which would be a large treatment effect if it were real.

It rests on 14 and 10 patients. Splitting 329 patients across four arms, two genders and two race categories leaves 16 cells, the smallest holding 8. At month 3 the picture is worse: 6 of 16 cells hold fewer than ten patients, the smallest 7.

Adjusted means for every cell at month 6, ordered high to low. Shaded rows hold fewer than 12 patients.
ArmGenderRacen Adjusted mean95% CI
IDC+GDCFemaleNon-Caucasian1455.1951.13 to 59.24
SE+GDCFemaleNon-Caucasian1151.3346.76 to 55.91
GDCMaleCaucasian2950.4547.64 to 53.27
SE+GDCFemaleCaucasian850.3044.94 to 55.66
CT+GDCFemaleCaucasian849.9544.59 to 55.31
GDCFemaleNon-Caucasian1149.3944.82 to 53.96
IDC+GDCMaleNon-Caucasian2849.0846.21 to 51.95
GDCFemaleCaucasian1049.0544.25 to 53.84
GDCMaleNon-Caucasian3349.0246.38 to 51.67
CT+GDCMaleNon-Caucasian3549.0046.44 to 51.57
IDC+GDCMaleCaucasian2548.5245.48 to 51.56
IDC+GDCFemaleCaucasian948.3343.27 to 53.39
CT+GDCMaleCaucasian3348.2245.58 to 50.87
SE+GDCMaleNon-Caucasian3647.0844.55 to 49.61
SE+GDCMaleCaucasian2947.0844.26 to 49.90
CT+GDCFemaleNon-Caucasian1043.4338.63 to 48.23

Read the intervals rather than the ordering. The cells at the extremes have intervals five points wide in each direction, and they overlap most of the middle of the table. The ordering is not stable information.

There is also a multiplicity problem that no single p-value reflects. Two timepoints, three model specifications, sixteen cells and a set of pairwise contrasts is a large family of tests, and the subgroup result is the most extreme cell in that family. Selecting the largest gap from sixteen cells and then quoting its p-value overstates the evidence by an amount the p-value does not show.

The honest reading is that this is hypothesis-generating. If the interaction between treatment and patient characteristics matters, it needs a study designed to test it, with subgroups sized in advance.

6. The recommendation, and how it was reached

The original analysis closed with a specific recommendation: assign non-Caucasian women to IDC+GDC. Because that is a clinical recommendation about a named group of patients, the chain of reasoning behind it deserves to be laid out in full rather than summarised.

The chain

1
The overall month 6 model showed nothing

Treatment had no detectable effect on its own (p = .2580).

2
A three-way model split the sample by gender and race

Sixteen cells from 329 patients.

3
Two cells sat at the extremes

Among non-Caucasian women, IDC+GDC gave the highest adjusted mean in the whole table (55.19) and CT+GDC the lowest (43.43).

4
The gap was read as a treatment effect

A 11.8-point difference, which became the recommendation.

The evidence, in full

Here is every arm for that subgroup, which is the comparison a treatment decision actually needs.

Non-Caucasian women at month 6, all four arms. Adjusted means with 95% intervals.
ArmnAdjusted mean 95% CI
IDC+GDC1455.1951.13 to 59.24
SE+GDC1151.3346.76 to 55.91
GDC (standard care)1149.3944.82 to 53.96
CT+GDC1043.4338.63 to 48.23

The IDC+GDC versus CT+GDC difference is real within this subgroup. Those two intervals do not overlap, so the comparison that generated the recommendation is not noise.

The comparison a treatment decision turns on is different. Nobody chooses between IDC+GDC and CT+GDC in a vacuum; the question is whether adding individual drug counselling beats the standard care every patient already receives. Against group drug counselling alone, IDC+GDC gives 55.19 against 49.39, and those intervals overlap substantially. On this evidence, IDC+GDC is not distinguishable from standard care for this subgroup.

So the finding supports a narrower statement than the recommendation makes: cognitive therapy appears to do worse than the alternatives for non-Caucasian women in this sample. That is worth investigating. It is not the same as establishing that individual drug counselling should be assigned to them.

Three further reasons for caution

The numbers behind it are small. 14 patients in the IDC+GDC cell, 10 in CT+GDC, 11 in the reference. Losing or gaining two or three patients would move these means noticeably.

The cells were chosen after seeing them. The highest and lowest of sixteen cells will differ substantially even when nothing is going on, simply because they are the extremes. A p-value computed on a comparison selected for being extreme does not mean what it appears to mean, and no correction was applied.

It did not replicate across timepoints. The same subgroup at month 3 had the highest adjusted mean too, and the original analysis records that it was not statistically significant there. A subgroup effect that appears at one timepoint and not the other, inside a null overall result, is the pattern one expects from chance.

What the recommendation is, and is not. It is a reasonable hypothesis generated by an exploratory analysis, and worth a study designed to test it, with subgroups sized in advance and the comparison specified before the data is seen.

It is not a basis for assigning treatment by race and gender. That would need replication in an independent sample, a prespecified comparison against standard care rather than against the worst-performing arm, and cells large enough to estimate the effect rather than merely rank it.

7. What this analysis supports

Adding individual drug counselling to group counselling was associated with better behaviour scores at three months. That association is not present at six months. Baseline severity predicts outcome more strongly than treatment assignment at both timepoints.

What it does not support: assigning treatment by gender or race. The subgroup signal in section 6 is worth following up and is not strong enough to act on.

What it cannot address: whether the month 3 advantage faded, or whether the patients measured at month 6 were simply different. Answering that needs the longitudinal model applied to a defined analysis population with the missingness handled explicitly, rather than two independent cross-sections.

8. How this was computed

Every figure is computed from the trial data with pandas and statsmodels, and the published models reproduce exactly: R² of 0.343 and 0.252, F of 37.45 and 27.34, treatment p of .0110 and .2580. The adjusted means in section 4 reproduce the values in the original write-up to two decimal places.

The trial data is restricted and is not in this repository. Aggregate results are written to tools/derived/, which is committed, so every page rebuilds without it. See the code page.