Summary
487 patients were randomized across four arms. At month 3, 292 had a usable outcome and baseline; at month 6, 329. Adjusting for baseline, the arms differ at month 3 (p = .0110) and do not at month 6 (p = .2580).
Two things qualify that. Only 60% and 68% of randomized patients reach the two analyses, though the loss is even across arms. And the normality violation reported at month 6 does not survive being tested on the residuals of the model that was actually reported.
1. Who is in each analysis
This is a six-month trial in cocaine dependence, so retention is the first thing to establish. Roughly a third of randomized patients are missing from each cross-section.
| Arm | Randomized | Month 3 | % | Month 6 | % |
|---|---|---|---|---|---|
| IDC+GDC | 121 | 68 | 56% | 76 | 63% |
| CT+GDC | 119 | 74 | 62% | 86 | 72% |
| SE+GDC | 124 | 71 | 57% | 84 | 68% |
| GDC | 123 | 79 | 64% | 83 | 68% |
Retention does not differ significantly by arm at either timepoint (chi-squared p = .52 at month 3, p = .48 at month 6). That is the reassuring version: the arms are not losing patients at different rates, so the comparison is not obviously distorted.
It is not the whole story. Patients who stay in an addiction trial differ from those who leave in ways a balance test cannot see, and the month 6 group is not the month 3 group. A treatment effect that appears at one timepoint and not the other could reflect the treatment, or it could reflect who was still there to measure. These data cannot separate those.
2. What the data looks like before modelling
Before any model, it is worth seeing the distributions. Scores are wide and heavily overlapping in every arm, which is the honest starting picture: whatever treatment does here, it does not separate the groups cleanly.
| Month | Arm | n | Mean | SD | Median | IQR | Range |
|---|---|---|---|---|---|---|---|
| 3 | CT+GDC | 74 | 48.78 | 8.96 | 49 | 43 to 55 | 28 to 70 |
| 3 | GDC | 79 | 49.49 | 9.78 | 49 | 42 to 57 | 25 to 67 |
| 3 | IDC+GDC | 68 | 52.06 | 7.46 | 50 | 46 to 57 | 37 to 66 |
| 3 | SE+GDC | 71 | 48.69 | 8.62 | 47 | 44 to 54 | 29 to 68 |
| 6 | CT+GDC | 86 | 48.17 | 9.39 | 48 | 40 to 55 | 27 to 66 |
| 6 | GDC | 83 | 49.36 | 9.36 | 50 | 44 to 58 | 27 to 68 |
| 6 | IDC+GDC | 76 | 50.04 | 8.60 | 49 | 43 to 57 | 25 to 68 |
| 6 | SE+GDC | 84 | 48.04 | 8.28 | 46 | 42 to 54 | 33 to 66 |
Two things stand out. IDC+GDC has the highest mean at month 3 and the ordering is visible, but the boxes overlap almost entirely, so the effect that follows is a shift in the middle of wide distributions rather than a separation of groups. And by month 6 even that ordering has largely gone.
3. The primary comparison
Each timepoint is analysed by ANCOVA: outcome on treatment arm, adjusted for the patient's own baseline score. Differences below are against group drug counselling alone.
| Model / contrast | n | R² | F | 95% CI | p | Baseline p |
|---|---|---|---|---|---|---|
| Month 3 | 292 | 0.343 | 37.45 | <.0001 | .0110 | <.0001 |
| IDC+GDC vs GDC | +2.50 | 0.15 to 4.86 | .0368 | |||
| CT+GDC vs GDC | -1.06 | -3.36 to 1.24 | .3648 | |||
| SE+GDC vs GDC | -1.07 | -3.40 to 1.25 | .3640 | |||
| Month 6 | 329 | 0.252 | 27.34 | <.0001 | .2580 | <.0001 |
| IDC+GDC vs GDC | +0.35 | -2.07 to 2.78 | .7747 | |||
| CT+GDC vs GDC | -1.43 | -3.78 to 0.92 | .2312 | |||
| SE+GDC vs GDC | -1.63 | -4.00 to 0.73 | .1754 | |||
At month 3 the model explains 0.343 of the variance and treatment matters (p = .0110). At month 6 it explains 0.252 and treatment does not (p = .2580). Baseline is overwhelmingly significant at both. Note what that means: most of what the model explains is the patient, not the therapy.
4. Assumptions, and a check that was run on the wrong model
The original analysis tested residual normality and found a violation at month 6, then investigated a Box-Cox transformation and concluded none was needed. Reproducing that turned up something worth reporting.
The normality check was run on residuals from a model containing treatment only. The model actually reported includes the baseline covariate. Those are different models with different residuals, and the assumption belongs to the model you report.
At month 6 the residuals of the treatment-only model reject normality (p = .0047), while the residuals of the reported ANCOVA do not (p = .4285). At month 3 neither rejects (.2668 and .1238). So the violation that prompted the transformation work was a property of a model that was not being reported. The reported models satisfy the assumption.
Variances are homogeneous across arms at both timepoints (Levene p = .30 and p = .37).
This does not overturn anything. The Box-Cox investigation concluded no transformation was needed, which is the same place this ends up. It is worth recording because the reasoning is what a reader would rely on, and the reasoning tested the wrong residuals.
5. Subgroups, and how much weight they hold
The analysis also tested whether treatment effects depend on gender and race, through two-way and three-way interaction models. A striking difference appears at month 6, and it needs its denominators attached.
The highest adjusted mean is IDC+GDC at 55.19, the lowest is CT+GDC at 43.43, both among non-Caucasian women. That is a gap of 11.8 points, which would be a large treatment effect if it were real.
It rests on 14 and 10 patients. Splitting 329 patients across four arms, two genders and two race categories leaves 16 cells, the smallest holding 8. At month 3 the picture is worse: 6 of 16 cells hold fewer than ten patients, the smallest 7.
| Arm | Gender | Race | n | Adjusted mean | 95% CI |
|---|---|---|---|---|---|
| IDC+GDC | Female | Non-Caucasian | 14 | 55.19 | 51.13 to 59.24 |
| SE+GDC | Female | Non-Caucasian | 11 | 51.33 | 46.76 to 55.91 |
| GDC | Male | Caucasian | 29 | 50.45 | 47.64 to 53.27 |
| SE+GDC | Female | Caucasian | 8 | 50.30 | 44.94 to 55.66 |
| CT+GDC | Female | Caucasian | 8 | 49.95 | 44.59 to 55.31 |
| GDC | Female | Non-Caucasian | 11 | 49.39 | 44.82 to 53.96 |
| IDC+GDC | Male | Non-Caucasian | 28 | 49.08 | 46.21 to 51.95 |
| GDC | Female | Caucasian | 10 | 49.05 | 44.25 to 53.84 |
| GDC | Male | Non-Caucasian | 33 | 49.02 | 46.38 to 51.67 |
| CT+GDC | Male | Non-Caucasian | 35 | 49.00 | 46.44 to 51.57 |
| IDC+GDC | Male | Caucasian | 25 | 48.52 | 45.48 to 51.56 |
| IDC+GDC | Female | Caucasian | 9 | 48.33 | 43.27 to 53.39 |
| CT+GDC | Male | Caucasian | 33 | 48.22 | 45.58 to 50.87 |
| SE+GDC | Male | Non-Caucasian | 36 | 47.08 | 44.55 to 49.61 |
| SE+GDC | Male | Caucasian | 29 | 47.08 | 44.26 to 49.90 |
| CT+GDC | Female | Non-Caucasian | 10 | 43.43 | 38.63 to 48.23 |
Read the intervals rather than the ordering. The cells at the extremes have intervals five points wide in each direction, and they overlap most of the middle of the table. The ordering is not stable information.
There is also a multiplicity problem that no single p-value reflects. Two timepoints, three model specifications, sixteen cells and a set of pairwise contrasts is a large family of tests, and the subgroup result is the most extreme cell in that family. Selecting the largest gap from sixteen cells and then quoting its p-value overstates the evidence by an amount the p-value does not show.
The honest reading is that this is hypothesis-generating. If the interaction between treatment and patient characteristics matters, it needs a study designed to test it, with subgroups sized in advance.
6. The recommendation, and how it was reached
The original analysis closed with a specific recommendation: assign non-Caucasian women to IDC+GDC. Because that is a clinical recommendation about a named group of patients, the chain of reasoning behind it deserves to be laid out in full rather than summarised.
The chain
Treatment had no detectable effect on its own (p = .2580).
Sixteen cells from 329 patients.
Among non-Caucasian women, IDC+GDC gave the highest adjusted mean in the whole table (55.19) and CT+GDC the lowest (43.43).
A 11.8-point difference, which became the recommendation.
The evidence, in full
Here is every arm for that subgroup, which is the comparison a treatment decision actually needs.
| Arm | n | Adjusted mean | 95% CI |
|---|---|---|---|
| IDC+GDC | 14 | 55.19 | 51.13 to 59.24 |
| SE+GDC | 11 | 51.33 | 46.76 to 55.91 |
| GDC (standard care) | 11 | 49.39 | 44.82 to 53.96 |
| CT+GDC | 10 | 43.43 | 38.63 to 48.23 |
The IDC+GDC versus CT+GDC difference is real within this subgroup. Those two intervals do not overlap, so the comparison that generated the recommendation is not noise.
The comparison a treatment decision turns on is different. Nobody chooses between IDC+GDC and CT+GDC in a vacuum; the question is whether adding individual drug counselling beats the standard care every patient already receives. Against group drug counselling alone, IDC+GDC gives 55.19 against 49.39, and those intervals overlap substantially. On this evidence, IDC+GDC is not distinguishable from standard care for this subgroup.
So the finding supports a narrower statement than the recommendation makes: cognitive therapy appears to do worse than the alternatives for non-Caucasian women in this sample. That is worth investigating. It is not the same as establishing that individual drug counselling should be assigned to them.
Three further reasons for caution
The numbers behind it are small. 14 patients in the IDC+GDC cell, 10 in CT+GDC, 11 in the reference. Losing or gaining two or three patients would move these means noticeably.
The cells were chosen after seeing them. The highest and lowest of sixteen cells will differ substantially even when nothing is going on, simply because they are the extremes. A p-value computed on a comparison selected for being extreme does not mean what it appears to mean, and no correction was applied.
It did not replicate across timepoints. The same subgroup at month 3 had the highest adjusted mean too, and the original analysis records that it was not statistically significant there. A subgroup effect that appears at one timepoint and not the other, inside a null overall result, is the pattern one expects from chance.
What the recommendation is, and is not. It is a reasonable hypothesis generated by an exploratory analysis, and worth a study designed to test it, with subgroups sized in advance and the comparison specified before the data is seen.
It is not a basis for assigning treatment by race and gender. That would need replication in an independent sample, a prespecified comparison against standard care rather than against the worst-performing arm, and cells large enough to estimate the effect rather than merely rank it.
7. What this analysis supports
Adding individual drug counselling to group counselling was associated with better behaviour scores at three months. That association is not present at six months. Baseline severity predicts outcome more strongly than treatment assignment at both timepoints.
What it does not support: assigning treatment by gender or race. The subgroup signal in section 6 is worth following up and is not strong enough to act on.
What it cannot address: whether the month 3 advantage faded, or whether the patients measured at month 6 were simply different. Answering that needs the longitudinal model applied to a defined analysis population with the missingness handled explicitly, rather than two independent cross-sections.
8. How this was computed
Every figure is computed from the trial data with pandas and statsmodels, and the published models reproduce exactly: R² of 0.343 and 0.252, F of 37.45 and 27.34, treatment p of .0110 and .2580. The adjusted means in section 4 reproduce the values in the original write-up to two decimal places.
The trial data is restricted and is not in this repository. Aggregate results are
written to tools/derived/, which is committed, so every page rebuilds without
it. See the code page.