Contents/ Part XXVIII · Capstone Projects: Sampling & Data Collection/ Chapter 176

The Sampling Plan

Every capstone so far began with a file. These four begin with nothing: a question, a population you cannot reach all of, and a budget. What you decide before collecting anything determines what the analysis at the end is allowed to say.

⏱️ ~15 min read
🎯 Part opener
📊 Chapter 176

The sixteen capstones in the previous part all started the same way: someone handed you a spreadsheet. That is the normal situation, and it is also the situation in which the most important decisions have already been made by someone else, usually without a record of why.

🎯
A sampling plan is the written answer to five questions, settled before collection: who counts as the population, what list you will actually draw from, how you will draw, how many you need, and how you will measure them. Everything downstream inherits those five answers.
The claim this part rests on

No analysis can repair a bad sample. A confidence interval describes uncertainty from drawing a sample; it says nothing about the people your frame never contained or the ones who declined to answer. Those errors do not shrink as n grows, and the arithmetic never announces them.

1

The Chain From Population to Estimate

Between the group you care about and the number you finally report there are four links, and each one is a place error gets in. Only one of the four is the kind a p-value knows about.

FOUR PLACES ERROR ENTERS, AND ONLY ONE OF THEM IS IN YOUR p-VALUE target population sampling frame drawn sample respondents estimate coverage error sampling error non-response error measurement error green: shrinks with n red: does NOT shrink with n A confidence interval quantifies the green arrow only. The three red ones are design problems, and a larger sample makes them no smaller, only more precisely wrong.
This is the whole argument for the part. The interval you report at the end covers one of four error sources. The other three are decided before collection begins, and the only defense against them is a plan written down in advance and a limitation section that admits what the plan could not fix.
ErrorWhere it entersA concrete exampleDoes more data help?
CoverageThe frame does not match the populationSurveying customers by email when a fifth of them never gave an addressNo
SamplingYou drew some members and not othersThe 200 people you happened to select differ from the 200 you might haveYes
Non-responseSelected members do not answerThe dissatisfied stop replying, so satisfaction looks higher than it isNo
MeasurementThe instrument distorts the answerA leading question, or a scale with no neutral optionNo
2

Choosing How to Draw

The single most consequential line in the plan is whether every member of the population has a known, non-zero chance of selection. If they do, the sample is a probability sample and the machinery of the previous twenty-seven parts applies. If they do not, that machinery does not apply, and pretending otherwise is the most common error in applied research.

WHICH SAMPLING METHOD? Does every member have a known, non-zero chance of selection? YES NO PROBABILITY · inference is valid NON-PROBABILITY · inference needs repair Simple random a complete list, no subgroups Stratified subgroups you must compare Cluster / multistage no list, but natural groups Systematic an ordered stream, every kth Quota, convenience, panel, snowball cheap and fast, and the selection rule is unknown Repair: rake to known population margins helps for what you weighted on, nothing else The left branch buys you a standard error. The right branch buys you speed, and owes you a caveat.
Cost drives most real decisions. A simple random sample of a national population is usually impossible, because no complete list exists and reaching a scattered sample is expensive. That is why stratified and multistage designs dominate serious survey work, and why the second project in this part spends its time on the price clustering charges in precision.
3

How Big, and How Precise

Sample size is where intuition fails most reliably. Two facts do most of the work.

The two facts

Precision improves with the square root of n. To halve a margin of error you need four times the sample, not twice. That single relationship sets the budget for most studies.

The fraction of the population barely matters. A sample of 1,000 estimates a national population of 60 million about as precisely as it estimates a town of 60,000. What matters is the absolute size of the sample, not the share it represents, which is why national polls of 1,000 people are not the absurdity they look.

Sample sizeMargin of error on a 50% proportionCost relative to n = 400
100± 9.8 points0.25×
400± 4.9 points
1,000± 3.1 points2.5×
1,600± 2.5 points
6,400± 1.2 points16×

Read the last two rows together. Quadrupling the sample from 1,600 to 6,400 buys 1.3 percentage points of precision for four times the money. Whether that is worth it depends entirely on what decision the estimate feeds, which is why the plan states the required precision first and derives n from it, rather than picking a round number and reporting whatever precision falls out.

Two adjustments the textbook formula leaves out

Clustering costs you. If you sample schools and then students within them, students in a school resemble each other, so each additional student adds less information than a fresh random one would. The design effect is the multiplier: a design effect of 2 means your 1,000 clustered responses carry the information of 500 independent ones. Plans that ignore it are systematically under-powered.

Non-response costs you twice. Once in sample size, which is easy to fix by inviting more people, and once in bias, which is not fixable that way at all. Plan the invitation count from the expected response rate, then plan the analysis for the bias that remains.

4

The Instrument Is Part of the Design

A questionnaire is a measuring device, and like any measuring device it can be miscalibrated. The data type of every answer is fixed the moment the question is written, and with it the entire analysis that will be possible later.

Question typeProducesWhat you can compute
Single choice from a listNominalProportions, chi-square, Cramer's V
Choose all that applySeveral binary variablesProportions per option, but the options are not mutually exclusive and must not be treated as one variable
Likert agreement scaleOrdinalMedians, top-box shares, rank correlations. Not means
Numeric entryContinuousMeans, standard deviations, the full parametric toolkit
Ranking taskOrdinal, dependent within respondentRank correlation, Kendall's W
Open textUnstructuredCoding into categories, then whatever the coding produces

Four failures account for most bad instruments, and all four are cheap to avoid at the writing stage and impossible to repair afterwards.

5

After Fielding: Who Did Not Answer

Every survey in this part ends with a section on the people who are not in it. That is not a formality, and it is where the most commonly repeated mistake in survey reporting lives.

A low response rate is not the problem. Non-response bias is

A 20 percent response rate is harmless if the 80 percent who declined are like the 20 percent who answered. A 90 percent response rate is dangerous if the missing 10 percent are all the people who had a terrible experience. The response rate tells you nothing on its own. What matters is whether responding is related to what you are trying to measure, and the response rate cannot tell you that.

Three moves are available, in increasing order of usefulness. Compare respondents to the frame on any variable you hold for everyone, since a frame usually carries age, region, tenure or purchase history whether or not the person replied. Weight so the respondent profile matches known population margins, which is what post-stratification and raking do. Follow up a subsample of non-respondents aggressively and compare them to the easy responders, which is the only one of the three that measures the bias rather than assuming it away.

Weighting deserves one warning. It corrects the variables you weighted on and nothing else, and it makes the estimate less precise, because unequal weights waste information. Weights that vary wildly are a sign the sample was badly out of shape to begin with, and no amount of weighting turns a broken sample into a good one.

6

How the Framework Changes

The twelve-step framework from the previous part still runs, unchanged, from the moment the data exists. What these capstones add is five steps in front of it.

StepDecisionWhat it determines
D1 · DefineWho exactly is the target population?What the estimate is an estimate of
D2 · FrameWhat list will you draw from?Coverage error, and who is invisible from the start
D3 · MethodHow will you draw?Whether standard errors are valid at all
D4 · SizeHow precise must the answer be?Budget, and what differences you can detect
D5 · InstrumentHow will you ask?The data type of every variable, and so every test
1 to 12The analysis framework, unchangedEverything from here is the previous part

Notice that D5 determines the data type, and the data type was step 2 of the analysis framework. These capstones simply push the chain back one link: instead of discovering that a variable is ordinal, you chose to make it ordinal, and you could have chosen otherwise.

7

Sampling in Data Science & AI

Nothing in this part is confined to questionnaires. Every training set is a sample, and almost none of them come with a sampling plan.

Where it appearsThe sampling question nobody asked
Training data for a modelWhat population does this sample represent, and does the deployment population match it?
Scraped corporaThe frame is "whatever was reachable and not blocked". Who is systematically absent from that?
Benchmark suitesA benchmark is a sample of tasks. Scoring well on it generalizes only as far as the sample does
Human labels and preferencesWhich annotators, recruited how, and are their judgments the population's judgments?
Logged behaviorYou observe only the users the current system already serves well. That is non-response by another name
Practice note

Dataset shift, the standard explanation for a model that degrades after deployment, is usually a sampling failure that was invisible at training time: the training frame did not cover the population the model eventually met. The discipline that prevents it is the one in this chapter, written down before collection rather than reconstructed afterwards from a post-mortem. A dataset shipped without a statement of what population it represents is a dataset whose limits nobody can check.

8

The Four Projects Ahead

Four projects, each starting before the data exists. They become clickable as each one is published; the Contents always shows what is live.

Probability designs · when you can draw properly
1
Stratified Sampling
A National Customer Study
Write the plan, field the questionnaire, then weight for the people who never replied.
2
Cluster & Multistage
Schools and Students
When you sample groups rather than people, what does that cost you in precision?
When the frame is the problem
3
Web Scraping
Job Postings as a Sampling Frame
You built the dataset yourself. Which population does it actually represent?
4
Non-Probability
Repairing an Online Panel
A fast, cheap, biased sample. How much can weighting rescue, and how would you know?

🎓 Key Takeaways

  • A confidence interval covers sampling error only. Coverage, non-response and measurement error are design problems, and they do not shrink as n grows.
  • Probability or not is the decisive line. Known non-zero selection probabilities are what make a standard error mean anything.
  • Precision goes as the square root of n, and barely depends on the fraction of the population sampled. Derive n from the precision you need.
  • The instrument fixes the data type, and therefore fixes which tests are available. That decision is made when the question is written, not when the analysis starts.
  • The response rate is not the problem; whether responding relates to the answer is. Weighting fixes the variables you weighted on and nothing else.
9

Quiz: Test Yourself

Eight questions on the plan that comes before the data. Answer them, hit Check Answers, and keep refining until you score 100%. Your progress is saved.