Contents/ Part XXVIII · Capstone Projects: Sampling & Data Collection/ Chapter 180

Repairing an Online Panel

Capstone 20. Five thousand two hundred people, matched to the census on age, sex and region. It is ten points wrong, the standard repair does not move it at all, and a probability sample one twelfth the size beats it comfortably.

⏱️ ~18 min read
🎯 Non-probability samples
📊 Chapter 180

This panel passes every check a client knows how to ask for. Age matches the census, sex matches, region matches. It is ten percentage points wrong, and the weighting procedure designed to fix exactly this problem leaves it precisely where it was.

The brief
Setting
An opt-in online panel of 5,197 adults recruited to quotas that match the census on age, sex and region, alongside a 420-person probability sample of the same population.
The question
What share of adults have used a particular public service?
Why it matters
The panel is cheap, fast, and passes every check a client knows how to run. Someone has to decide whether that is enough to publish the number.
What we do
Rake the panel to census margins and watch it change nothing, find the variable that actually selected people, rake again on that, and finally price the whole panel in equivalent random-sample terms.
∑w
Quota sampling recruits until the sample matches known population shares. Raking then adjusts weights until every margin matches. Both operate on margins, and neither controls who was selected within a cell, which is where a volunteer panel does its damage.
The finding, up front

The panel says 44.2%; the truth is 34.1%. Raking to the census changes the estimate by 0.001 of a percentage point. Priced honestly, 5,197 opt-in respondents are worth a random sample of 22 people, and the 420-person probability sample is worth 769.

1

A Sample With No Selection Probabilities

The target population is 240,000 adults. There is no frame. Panel members opted in through advertising and referrals, so the probability that any given adult is in this sample is unknown and cannot be recovered. That one sentence rules out every standard error in the book.

Quota sampling is the industry's answer. Recruitment continues until the panel matches known population shares on a few demographics: fast, cheap, and the resulting file looks convincing.

Matching a margin is not random selection

A quota tells the recruiter how many 35-to-54-year-old women in the Midwest to sign up. It says nothing about which ones. Inside every quota cell, membership was decided by willingness to join an online panel, and willingness is not a coin flip.

2

Every Check a Client Would Run, Passed

Cleaning removed two duplicate submissions and three blank outcomes, leaving 5,197 usable respondents. Then the representativeness checks.

VariablePanelPopulationGap
Age 18-3430.1%30.1%0.0
Age 35-5434.9%35.0%0.0
Age 55+34.9%34.9%0.0
Female51.1%51.1%0.0
Northeast / Midwest27.8% / 22.2%27.8% / 22.2%0.0
South / West34.9% / 15.0%34.9% / 15.0%0.0

Every quota was met to within a rounding error. If a client asked whether the sample was representative, every check they know how to run would say yes.

3

And It Is Ten Points Wrong

Opt-in panel
44.2%
n = 5,197
Probability sample
32.4%
n = 420, CI 27.9 to 36.9
True value
34.1%
knowable only here
Panel error
+10.2 pp
6× the small sample's

The probability sample of 420 is out by 1.7 points and its interval contains the truth. The panel's error is six times larger on twelve times the sample. That is the entire argument for probability sampling in two lines, and it is why survey organizations still pay for expensive fieldwork when a panel costs a fraction as much.

4

The Standard Repair Does Nothing

Raking is the accepted fix: adjust the weights until every margin matches the census. Applied here to age by sex and to region, it produces weights between 0.999 and 1.002 and moves the estimate by 0.001 of a percentage point.

The trap, in one sentence

Raking corrects a sample whose margins are wrong, and these margins were already right, because the panel was recruited to make them right. There is nothing left for the procedure to adjust.

So the panel was built to pass exactly the check that the standard repair performs. The repair certifies it, changes nothing, and the ten-point error survives untouched. A report that said "weighted to census margins" would be true and would mean nothing.

5

Find What Actually Selected People

If demographics are not what makes this panel unusual, something else is. The questionnaire also carried a self-reported internet-use band, and an official communications survey publishes the same bands for the whole population.

Left: grouped bars of region shares for panel and population, matching exactly with gaps of zero. Right: grouped bars of internet-use bands, where the panel has 8.9 percent low users against a population 29.4 percent, and 45.3 percent high users against 26.3 percent.
Left: every quota variable matches. Right: the variable nobody set a quota on. The panel holds 8.9% light internet users against a population 29.4%, and 45.3% heavy users against 26.3%.
Internet usePanelPopulationGapUsed the service
Low8.9%29.4%−20.5 pp13.3%
Medium45.9%44.3%+1.6 pp32.2%
High45.3%26.3%+19.0 pp62.5%

There it is. That is who volunteers for an online panel, and it was invisible in every demographic check because it cuts across age, sex and region rather than lining up with them. It also matters enormously here: heavy internet users are nearly five times likelier to have used an online service. A variable that predicts both joining and the outcome is the definition of the thing that biases an estimate.

WHY WEIGHTING ON DEMOGRAPHICS CHANGED NOTHING internet confidence nobody set a quota on this joined the panel used the service age, sex, region quota-ed, and already balanced Weighting fixes the gray box. The red arrows were never touched.
The selector has to be on the weighting list. Demographics were balanced by recruitment, so adjusting them accomplished nothing. The variable driving both membership and the outcome sat outside the procedure entirely, which is why the estimate did not move.
6

Rake Again, on the Variable That Matters

EstimateValueErrorRecovered
Panel, unweighted44.24%+10.15 pp
Panel, raked on demographics44.24%+10.15 pp0%
Panel, raked + internet use35.59%+1.50 pp85%
Probability sample32.38%−1.71 pp
True value34.09%

Eighty-five percent of the bias, gone. Two warnings come with it.

The weight range is 34 to 1, far beyond the four-to-one in Capstone 17. A few hundred light internet users are each standing in for dozens of people, and if those few happen to be unusual the correction inherits their oddity. The effective sample size falls from 5,197 to about 3,000, a design effect of 1.74.

The repair depended on luck

This worked because someone thought to ask about internet use and an official source published the matching margin. Neither was guaranteed. Had the panel carried nothing but demographics, the ten-point error would have been not merely unfixable but undetectable, and the study would have reported 44 percent with total confidence.

7

What Is This Panel Actually Worth?

Sample size is the wrong currency for a biased sample. The useful question is how large a random sample would have to be to be this far off. Setting the standard error of a simple random sample equal to the observed error and solving for n answers it.

Left: four bars of estimates against a dashed truth line at 34.1 percent, with the unweighted and demographically raked bars identical at 44.2. Right: horizontal bars of equivalent random sample size, 22 for the unweighted panel, 22 for the raked panel, 1003 with internet use, and 769 for the 420-person probability sample.
Left: the first two bars are identical, which is the finding. Right: the same four samples priced in random-sample equivalents. The 420-person probability sample towers over an opt-in panel twelve times its size.
SamplenErrorEquivalent random sample
Opt-in panel, unweighted5,19710.15 pp22
Opt-in panel, raked demographics5,19710.15 pp22
Opt-in panel, + internet use5,1971.50 pp1,003
Probability sample4201.71 pp769

The 5,197-person panel is worth about twenty-two people. That is not rhetoric, it is the arithmetic. Raking on demographics leaves it at twenty-two, because it changed nothing. Raking on internet use lifts it to around a thousand, a genuine achievement that still means 80 percent of the fieldwork bought nothing.

WHAT GROWING THE SAMPLE DOES, AND DOES NOT DO OPT-IN PANEL RANDOM SAMPLE sample size → sample size → sampling error bias: flat forever total error never falls below the bias no bias to stop at total error goes to zero as n grows
Sample size fights variance, and only variance. On the left the gray curve is the part that shrinks and the red band is the part that does not, so the total error flattens out at the bias and stays there no matter how many people are recruited. On the right there is no floor to hit. This is why 420 randomly chosen adults beat 5,197 volunteers.
And bigger panels do not help

Doubling the panel to 10,000 would halve the sampling error, which was never the problem, and leave the ten-point bias exactly where it is. Sample size fights variance, and this was never a variance problem. The 420 expensive interviews beat the 5,197 cheap ones because they were drawn at random, and nothing else about them was better.

8

What to Watch

9

Non-Probability Samples in Data Science & AI

Almost every dataset in modern practice is a convenience sample, and the panel's failure mode is the standard one.

Where it appearsThe unmeasured selector
App telemetryWillingness to keep the app installed, which tracks satisfaction
Opt-in feedback widgetsStrength of feeling, in both directions
Crowdsourced labelsWhoever finds the task worth the fee at that hour
Beta program usersEnthusiasm and technical confidence, exactly as here
Balanced benchmark suitesBalanced on the axes someone thought of, and no others
Practice note

Rebalancing a training set on demographic attributes is this chapter's raking step, and it inherits the same limitation. If the mechanism that decided which examples were collected is not one of the attributes you balanced on, balancing changes the composition and not the bias. The uncomfortable corollary is the arithmetic in section 7: a very large biased dataset can carry less information about a population than a small carefully drawn one, and no amount of scale reverses it. Meng's work on this shows the effect worsens as the population grows, which is precisely the regime modern datasets operate in.

🐍

The full project, step by step

The companion notebook verifies that every quota was met, compares the panel against a probability sample and against the truth, implements iterative proportional fitting from scratch and shows it changing nothing, locates the variable that actually selected respondents, rakes again to recover 85 percent of the bias, and closes by pricing all four estimates in equivalent-random-sample terms.

📓 View Notebook (code & outputs) ▶ Open in Colab ⬇ View / Download on GitHub
Read the reports & get the data

The dataset (capstone-repairing-an-online-panel.xlsx) holds the 5,200-member opt-in panel, an independent 420-person probability sample for comparison, the census margins for raking, an official internet-use margin, and the true population value. Two written reports accompany it: a plain-language brief for a research buyer, and a technical report covering the raking, the weight diagnostics and the equivalent-sample-size argument.

🎓 Key Takeaways

  • Matching the census is not being representative. Every quota matched to a rounding error and the estimate was 10 points out.
  • Raking cannot fix margins that already match. Weights came out at 1.00 and the estimate moved by 0.001 points.
  • Weight on the selector. Internet use drove both joining and the outcome; adding it recovered 85% of the bias.
  • Price the sample honestly. 5,197 opt-in respondents were worth a random sample of 22; the 420-person probability sample was worth 769.
  • Scale fights variance, not bias. Doubling the panel would halve a standard error that was never the problem.
10

Quiz: Test Yourself

Eight questions on quota samples, raking and what a biased sample is worth. Answer them, hit Check Answers, and keep refining until you score 100%. Your progress is saved.