Contents/ Part XXIX · Capstone Projects: Design & Causal Inference/ Chapter 182

Designing a Factorial Experiment

Capstone 22. An eight-run design named line speed as the second-biggest lever on coating strength. Line speed does nothing at all. Its column in the design matrix is identical to another column, and the estimate belonged to that one.

⏱️ ~21 min read
🎯 Factorial design
📊 Chapter 182

Capstone 7 analyzed a factorial experiment somebody else had designed. This one is about the design itself: how many runs you can afford, what a cheaper design stops being able to tell you, and how to find that out before the runs rather than after.

The brief
Setting
A contract coating line. Five process settings can be changed: oven temperature, line speed, primer viscosity, cure time and nozzle pressure. The response is peel strength in newtons, higher is better.
The question
Which settings raise coating peel strength, and what combination should the line actually run at?
Why it matters
Every run costs a line changeover, so the experiment has to be small. A design that is too small does not merely lose precision; it makes some effects impossible to tell apart, and the output gives no warning that this has happened.
What we do
Work out how many runs the full factorial would take, look at the eight-run design the team already ran and the recommendation it produced, show the alias that made that recommendation wrong, run a sixteen-run design that separates the effects, and price what blocking and randomization each contributed.
Two effects are aliased when the experiment's design gives them the same column of plus and minus signs. The analysis then returns one number for both, and no amount of care in the analysis can separate them. A design's resolution says which kinds of effect are aliased with which.
The finding, up front

The eight-run design put line speed at +6.22 N, the second-largest effect in the study. The sixteen-run design puts it at +0.17 N, p = 0.65. The 6.22 belonged to temperature crossed with cure time, which the eight-run design had no column for. Slowing the line changed nothing; raising both temperature and cure time together gained 9.46 N.

1

Five Factors, and the Arithmetic of Runs

Five factors at two levels each is 25 = 32 distinct settings. Replicated twice, so that noise can be estimated without assuming the model is right, that is 64 runs. Each run needs a line changeover, which is why this is a design problem before it is an analysis problem.

DesignSettingsRuns, 2 replicatesWhat it can separate
Full 253264Everything, up to the five-factor interaction
Half 25−1, resolution V1632All 5 main effects and all 10 two-factor interactions
Quarter 25−2, resolution III816Five effects, each of which is a sum of a main effect and interactions

The budget allowed 32 runs. The interesting question is what the quarter fraction, at half that cost, stops being able to tell you, and the honest answer is that it stops being able to tell you what it is telling you.

FULL FACTORIAL · 32 SETTINGS · 64 RUNS every effect has its own column HALF FRACTION, RESOLUTION V · 16 SETTINGS · 32 RUNS · the design we ran 5 main + 10 two-factor aliased only with 3- and 4-factor terms QUARTER FRACTION, RESOLUTION III · 8 SETTINGS · 16 RUNS · run last quarter 5 estimates each one is a main effect PLUS a two-factor interaction cost falls, and so does what you can learn
What each halving costs. The run count is the visible price and the resolution is the invisible one. A resolution III design returns five confident numbers and does not mention that each of them is a sum.
2

The Eight-Run Design That Was Run First

Last quarter the team ran a quarter fraction. With three factors you get eight settings for free; the other two have to be built out of the first three, and the team chose D = AB and E = AC. Eight settings, two replicates, sixteen runs. Five factors, five effects, all estimated. It looks complete.

FactorEffect (N)What the design actually estimates
Oven temperature+10.31A + BD + CE
Line speed+6.22B + AD
Primer viscosity+1.57C + AE
Cure time+1.28D + AB
Nozzle pressure−0.98E + AC

The standard error of an effect here is about 0.86 N, so line speed at +6.22 is more than seven standard errors from zero. By any conventional reading that is a solid, actionable finding, and the recommendation wrote itself: slow the line down.

3

Resolution, and Seeing an Alias for Yourself

The right-hand column of that table is not a theoretical caution. It is a statement about arithmetic that you can check in a few lines: build the column of signs for the interaction A×D, and compare it element by element with the column for B.

THE EIGHT RUNS, IN SIGNS run A D B A×D 1+ 2+ 3++ 4++ 5+ 6+ 7++++ 8++++ same The design cannot tell them apart, because within these eight runs they are the same eight numbers. Whatever the analysis returns for B, it is really returning B + A×D.
An alias, in full. Because the design was built with D = AB, the product A×D equals A×A×B, and A×A is a column of plus signs. So A×D is B. This is not a subtlety of inference; it is multiplication.

The same check across all five factors gives the whole picture: A is indistinguishable from B×D and C×E, B from A×D, C from A×E, D from A×B, and E from A×C. That is what resolution III means, and it is a property of the design, decided before any panel was coated. No analysis recovers information the design never collected.

4

Acting On It, and the Confirmation That Followed

The line was slowed from 12 to 8 meters per minute. Forty panels were measured, twenty before and twenty after. The design had promised about six newtons.

Before, 12 m/min
60.74 N
n = 20
After, 8 m/min
60.58 N
n = 20
Difference
−0.16 N
95% CI [−0.96, +0.63]
p-value
0.68
nothing happened

A confirmation run is the cheapest insurance in experimental work, and this is what it buys. The interval runs from about one newton down to two thirds of a newton up, and excludes the promised +6 many times over. Something was wrong with the design, not with the process, and forty panels were enough to find out.

5

The Sixteen-Run Design

The half fraction uses a single generator, E = ABCD, giving the defining relation I = ABCDE. Main effects are now aliased only with four-factor interactions, and two-factor interactions only with three-factor ones. Under the usual assumption that high-order interactions in a physical process are negligible, all five main effects and all ten two-factor interactions are separately estimable. That is resolution V, and the notebook confirms it the same way as before: among those fifteen columns, no two are identical.

StepRunsWhat happened
As delivered33The logger wrote one row twice
After deduplication32Two full replicates of 16 settings
After a failed gauge31One panel has no reading, so 15 of 16 settings have both replicates

The lost run leaves the design very slightly unbalanced. Least squares handles that without comment; the hand-computed contrast of the previous section would not, which is a good practical reason to fit a model rather than average columns once a design stops being perfectly balanced.

6

Line Speed Was Never There

TermEffect (N)SEp
Oven temperature+8.470.38<0.001
Temperature × cure time+6.470.38<0.001
Viscosity × pressure+2.450.38<0.001
Primer viscosity+2.010.38<0.001
Cure time+2.010.38<0.001
Nozzle pressure−1.040.380.015
Line speed+0.170.380.65

Line speed is +0.17 newtons on a standard error of 0.38. It was never there. The +6.22 the eight-run design reported belonged to oven temperature crossed with cure time, which appears here at +6.47 and could not have appeared at all in a design where its column and line speed's column were the same numbers.

Left: grouped bars of the effect on peel strength for five factors plus the temperature-by-cure-time interaction, comparing the 8-run resolution III design against the 16-run resolution V design. Line speed falls from 6.22 to 0.17 between the two, annotated with an arrow, while the temperature-by-cure interaction appears only in the 16-run design at 6.47 and is labeled the effect it was standing in for. The 8-run design has no bar there, labeled not estimable in 8 runs. Right: an interaction plot with mean peel strength against oven temperature, one line for 45 second cure rising gently from 60.9 to 63.2 and one for 75 second cure rising steeply from 56.7 to 71.7.
Left: the same five factors measured two ways. Line speed collapses from the second-largest effect to nothing, and the interaction the eight-run design had no column for takes its place. Right: the interaction itself. Raising the temperature buys 2.3 N at short cure and 14.9 N at long cure, so neither factor on its own is the recommendation.
Why the eight-run answer was so convincing

The aliased estimate was not noise. It was a real effect of about the right magnitude, attached to the wrong label. That is what makes resolution III dangerous rather than merely imprecise: a low-resolution design does not produce vague results that invite caution, it produces sharp results that invite action. The standard error was 0.86 N and it was correct. It was measuring the precision of a quantity nobody wanted.

7

Blocking: the Batch You Should Not Randomize

Thirty-two runs need two batches of primer, and batches differ. The plan put replicate 1 entirely on batch 1 and replicate 2 on batch 2, so batch is the block. It enters the model as a term, which takes the batch difference out of the error rather than leaving it there to inflate every standard error.

Residual SD, blocked
1.03 N
batch in the model
Residual SD, ignored
1.56 N
batch left in the error
SE of an effect
0.38 vs 0.57
34 percent narrower
Batch 2 offset
+1.69 N
real, and not of interest

Note what blocking did not do. The batch difference is still there in the process; it has not been removed from the world, only from the yardstick. And note the choice it represents. Batch was known in advance, so it became a block. Had the batches instead been randomized across runs, the same variation would have gone into the error term and every interval would have been a third wider for no benefit. The rule is short: block what you can predict, randomize what you cannot.

8

Randomization: the Drift Nobody Measured

The oven creeps upward across a shift. Nobody measured that and nobody had to, because run order within each batch was randomized. The value of randomizing is easiest to see by simulating the alternative: running the settings in a convenient order, all the low-pressure runs first, which on a real line saves genuine time.

True effect
−0.80 N
nozzle pressure
Randomized order
−0.805
bias −0.005
Sorted by pressure
−0.408
bias +0.392
Consequence
Halved
across 2,000 simulations

Four tenths of a newton is a five percent distortion of the temperature effect, which nobody would ever notice. On nozzle pressure it is half the effect, and it changes the conclusion about whether pressure matters at all. You cannot know in advance which of your effects are small, which is the entire argument for randomizing: it converts a bias you cannot see into noise you can measure.

Left: two overlaid histograms of the estimated effect of nozzle pressure across 2000 simulated experiments. The randomized-order distribution is centered on the dashed truth line at minus 0.8, while the distribution from runs sorted by pressure is shifted right, centered near minus 0.4. Right: two bars of the standard error of an effect, 0.377 newtons when blocked on batch and 0.567 when batch is ignored, titled blocking narrowed every interval by 34 percent.
Left: two thousand simulated experiments each way. Randomizing centers the estimate on the truth; a convenient run order shifts the whole distribution. Right: what the block was worth, on the same runs and the same data.
9

The Recommendation, and Confirming It

The recommendation is not a factor. It is a combination, and the interaction is the reason.

Mean peel strength (N)Cure 45 sCure 75 s
Temperature 170 °C60.8956.73
Temperature 190 °C63.1971.67

Read the bottom row against the top. Raising the temperature at short cure gains 2.3 N. Raising it at long cure gains 14.9 N. And at the low temperature, extending the cure actually makes the coating worse, from 60.89 down to 56.73. A recommendation that named either factor alone would be wrong in a way that a recommendation naming both is not.

Baseline
60.54 N
170 °C, 45 s
Recommended
70.00 N
190 °C, 75 s
Gain
+9.46 N
95% CI [+8.61, +10.31]
Confirmed
40 panels
fresh, after the fact

The design that cost twice as much found the lever the cheap design pointed away from, and the confirmation run is what turns an estimate into a change somebody is willing to sign.

10

What to Watch

11

Design of Experiments in Data Science & AI

Factorial thinking is older than computing and it keeps being rediscovered wherever runs are expensive.

Where it appearsThe same idea, in a different costume
Hyperparameter searchGrid search is a full factorial; random and Latin-hypercube search are fractional designs chosen because the grid does not fit in the budget
Ablation studiesRemoving components one at a time is a one-factor-at-a-time design, which cannot see interactions between components at all
Multivariate web testingSeveral page elements varied at once, usually in a fractional design, usually without the alias structure being reported
Prompt and pipeline tuningInstruction, examples, temperature and retrieval depth crossed together, where the interactions are the interesting part
Simulation and A/B infrastructureBlocking on day, region or cohort, for exactly the reason batch was blocked here
Where the research went

Fractional designs come from Finney in 1945 and were developed for industry by Box, Hunter and Hunter, whose book remains the standard reference. Two directions are worth knowing. Response surface methodology adds center points and axial runs to fit curvature, which a two-level design cannot see at all: every factor here was tested at exactly two settings, so the analysis assumes the response between them is a straight line. Optimal design, the D-optimal and I-optimal families, drops the requirement for a regular fraction entirely and instead searches for the run set that minimizes a stated criterion, which is what most modern software does when the factors will not cooperate with a textbook layout.

🐍

The full project, step by step

The companion notebook works out the run arithmetic, builds the eight-run design and finds its aliases by comparing columns rather than by quoting a rule, reproduces the screening result and the confirmation run that contradicted it, verifies that the sixteen-run design has no duplicated columns, fits all fifteen effects with batch as a block, prices the block and the randomization in newtons, and closes with the interaction table and the final confirmation.

📓 View Notebook (code & outputs) ▶ Open in Colab ⬇ View / Download on GitHub
Read the reports & get the data

The dataset (capstone-designing-a-factorial-experiment.xlsx) holds both designed experiments with their coded and decoded factor columns, the two confirmation studies, the design plan written before any panel was coated, and the model that generated the data so every estimate can be checked. Two written reports accompany it: a plain-language brief for the plant manager, and a technical report covering the alias structure, the blocked model and the confirmation.

🎓 Key Takeaways

  • An alias is two effects sharing one column of signs. Because D was built as AB, the A×D column is the B column, and the design returns one number for both.
  • Low resolution produces sharp wrong answers, not vague ones. Line speed came out at +6.22 N with a correct standard error of 0.86, and it is truly +0.17.
  • Block what you can predict, randomize what you cannot. Blocking on batch cut every standard error by 34 percent; randomizing run order kept a 0.4 N oven drift out of the effects.
  • The recommendation was a combination. Raising temperature gains 2.3 N at short cure and 14.9 N at long cure, so neither factor alone is the answer.
  • Confirm on fresh material. Forty panels showed the first recommendation did nothing, and forty more confirmed the second was worth +9.46 N.

Quiz: Test Yourself