Contents/ Part X Β· Sampling & Data Collection/ Chapter 69

Survey & Questionnaire Design

A survey is only as good as its questions. This chapter is a practical guide to how researchers decide what to ask, choose response scales for each kind of question, word items to avoid bias, structure a questionnaire, and meet the industry standards that make survey data trustworthy.

⏱️ ~20 min read
πŸ“‹ Reference chapter
πŸ“Š Chapter 69

The previous chapters showed how sampling and collection introduce bias. Surveys add one more, earlier, source: the questions themselves. A poorly worded item biases every answer before a single response is recorded. Good questionnaire design is the craft of measuring what you actually intend to measure.

?
A questionnaire is a measurement instrument. Designing one means turning an abstract construct (satisfaction, intent, health) into concrete items with appropriate response scales, worded so that the answer reflects the respondent's true state, not the wording, the order, or social pressure.
πŸ“‹
How to read this chapter

This is a reference chapter: no code, just principles, examples, and standards. We move from deciding what to askchoosing scaleswording itemsstructuring the questionnairefour real-world examplesthe standards bodies that govern professional practice.

1

From Research Goal to Question

Good surveys are not written question-first. They start from the research objective and work down through a deliberate process called operationalization: defining the abstract concept, breaking it into measurable dimensions, and only then writing items.

Operationalization: from a concept down to concrete items Research objective: "Are employees satisfied?" Construct: job satisfaction Pay & benefits Workload Recognition Item: "My workload is manageable." Strongly disagree · Disagree · Neutral · Agree · Strongly agree

Each step answers a question. What do we want to learn? sets the objective. What exactly do we mean? defines the construct, often multidimensional. What observable things reveal it? gives the dimensions. How do we ask? produces the items. Skipping to items is the most common failure: you end up with a pile of questions that do not add up to the thing you meant to measure.

Validity and reliability

Two properties decide whether items work. Validity asks "does it measure the right thing?" (content, construct, and criterion validity). Reliability asks "does it measure consistently?" Multi-item scales are often checked with Cronbach's alpha (a value of 0.70 or higher is the common threshold for acceptable internal consistency). A measure can be reliable yet invalid, so you need both.

2

Question Types & Response Scales

The response format must match what you are measuring. Choosing the wrong scale, a yes/no for an attitude, a 5-point agreement for a behavior count, throws away information or fabricates precision.

Common response scales Dichotomous Yes No Likert (5-pt) SDDNASA odd # of points = a neutral midpoint Semantic diff. Cheap Expensive bipolar adjective pairs Numeric (0–10) 0510 rating / NPS Ranking 1. Price 2. Quality 3. Service forced order of preference Match the scale to the construct: attitudes → Likert; preferences → ranking; magnitude → numeric
TypeBest forMeasurement level & notes
Open-endedExploration, unanticipated answers, "why"rich but costly to code; use sparingly
DichotomousClear yes/no facts, screeningnominal; fast but low information
Multiple choiceDiscrete categoriesnominal; options must be mutually exclusive & exhaustive
LikertAttitudes, agreement, satisfactionordinal; usually 5 or 7 points, labeled
Semantic differentialImage, perception, brand feelbipolar adjectives; treated as interval-like
Numeric / ratingMagnitude, likelihood (e.g., NPS 0–10)interval-like; anchor the endpoints clearly
FrequencyHow often a behavior occursuse specific ranges, not vague "often/rarely"
RankingRelative priorities among a few itemsordinal; hard beyond ~5–7 items

Scale design choices matter as much as the type. An odd number of points offers a neutral midpoint; an even number forces a lean (a "forced choice"). Label every point where you can, not just the ends, so respondents interpret them the same way. Keep scales balanced (equal positive and negative options) and decide deliberately whether to offer "Don't know" or "Not applicable", omitting it pushes people to guess, while including it can invite satisficing. The original 5-point agreement format traces to Rensis Likert's 1932 monograph; the bipolar semantic differential to Osgood and colleagues in 1957.

3

Writing Questions That Don't Bias the Answer

Wording is where most surveys fail. Respondents work through four cognitive steps for every question, they comprehend it, retrieve the relevant memory, form a judgment, and map it to a response (Tourangeau, Rips & Rasinski, 2000). A flaw at any step distorts the answer.

✗ Leading
"How much did you enjoy our excellent new app?"
presumes enjoyment
✓ Neutral
"How would you rate your experience with our app?"
Very poor … Very good
✗ Double-barreled
"Was the staff friendly and knowledgeable?"
two questions, one answer
✓ Split in two
"Was the staff friendly?" & "Was the staff knowledgeable?"
one idea per item
✗ Loaded / double negative
"Don't you agree the policy shouldn't be repealed?"
confusing, pressuring
✓ Plain & balanced
"Do you favor or oppose the policy?"
offers both sides equally
✗ Vague frequency
"Do you exercise often?"
"often" means different things
✓ Specific ranges
"In the past 7 days, how many days did you exercise?"
0 · 1–2 · 3–4 · 5+
PitfallWhy it biasesFix
Leading / loaded wordingsteers toward an answerneutral, value-free language
Double-barreledtwo ideas, one answer slotone concept per question
Double negativeshard to parse, errorsstate it positively
Ambiguous termseveryone reads them differentlydefine terms; use concrete ranges
Jargon / reading levelexcludes or confusesplain language, ~grade 6–8
Acquiescence biastendency to agreemix positively/negatively keyed items
Social desirabilityflattering self-reportsanonymity, indirect or list techniques
Recall errormemory fades and distortsshort, bounded reference periods
Non-exhaustive optionsrespondent has no valid answercover all cases; add "Other"

A famous demonstration of how much wording matters: in repeated experiments the U.S. public expressed far more support for spending on "assistance to the poor" than on "welfare", though the two refer to the same programs. The label, not the policy, moved the answer, which is why professional pollsters test alternative wordings before fielding.

4

Questionnaire Structure & Best Practices

Individual questions are necessary but not sufficient; their order and flow shape responses too. The standard arc moves from easy and engaging, through the substance, to sensitive items and demographics at the very end.

A well-sequenced questionnaire Welcomepurpose, consent Easy & warm-upbuild rapport Main topicgrouped, logical Sensitive itemslater, once trust built Demographics& thank you Funnel from general to specific; keep related items together; minimize burden Always pretest before fielding
Best practiceWhy
Open with easy, non-threatening itemsbuilds rapport and commitment, lowers drop-off
Group related questions; logical flowreduces cognitive load and confusion
Sensitive questions & demographics lasttrust is established; refusals do not sink the whole survey
Watch order & context effectsan earlier question can prime later answers; randomize when apt
Randomize answer-option ordercounters primacy/recency response-order bias
Keep it short; minimize burdenfatigue causes satisficing and break-offs
Design mobile-firstmost surveys are now taken on phones; avoid wide grids
Pretest: pilot & cognitive interviewscatches misreadings before they cost you the whole sample
!
Pretesting is non-negotiable

Before fielding, run a pilot and ideally cognitive interviews, where a few respondents think aloud as they answer, so you hear how they actually interpret each item. It is far cheaper to fix a confusing question with five testers than to discover, after 2,000 responses, that half of them misunderstood it. Dillman's Tailored Design Method is the standard practitioner reference for the full process.

5

Four Real-World Examples

Theory becomes concrete in widely used, well-documented instruments. Each of these makes deliberate design choices, and illustrates a different lesson. For every example below you can view a sample questionnaire (PDF) of 10–15 questions in varied formats and download the matching data sheet (Excel), laid out the analysis-ready way a statistician would receive it: one row per respondent, coded columns, plus a codebook and analyst notes.

Example 1 Β· Business / customer experience

Net Promoter Score (NPS)

Introduced by Fred Reichheld (Bain & Company) in the Harvard Business Review, 2003, NPS asks a single question: "How likely are you to recommend us to a friend or colleague?" on an 0–10 numeric scale. Respondents are bucketed into Promoters (9–10), Passives (7–8), and Detractors (0–6), and the score is %Promoters − %Detractors (ranging −100 to +100). Lesson: a carefully chosen single-item scale can be a powerful, comparable metric, but the 0–10 anchoring and the bucketing rules are essential and must be applied consistently to compare across time or companies.

Example 2 Β· Clinical / health screening

PHQ-9 Depression Screener

The Patient Health Questionnaire-9 (Kroenke, Spitzer & Williams, 2001) is a validated nine-item instrument asking how often, over the last two weeks, a person was bothered by each symptom, on a four-point frequency scale: "Not at all" (0), "Several days" (1), "More than half the days" (2), "Nearly every day" (3). Item scores sum to 0–27, with cutoffs at 5, 10, 15, and 20 marking mild, moderate, moderately severe, and severe depression. Lesson: a validated, reliable instrument with specific frequency anchors and a published scoring rule yields a clinically actionable, comparable measure, the payoff of rigorous design and testing.

Example 3 Β· Public opinion / polling

Pew Research Center wording experiments

Pew, a gold standard in survey research, routinely runs split-sample wording experiments: half the respondents see one phrasing, half another. Their work documented how "assistance to the poor" versus "welfare" can swing apparent support by tens of points, and they deliberately offer balanced response options ("favor or oppose") and rotate answer order. Lesson: question wording is itself a measurable source of error; professional pollsters test alternatives and report methodology transparently rather than trusting a single phrasing.

Example 4 Β· Government / official statistics

U.S. Census Bureau & the American Community Survey

Federal surveys are designed and revised through years of cognitive testing and field experiments. The Census Bureau has repeatedly tested its race and Hispanic-origin questions, formatting, examples, and combined-versus-separate items, because small wording and layout changes measurably shift how millions of people respond. Lesson: at scale, format and pretesting are as consequential as wording; even the order of examples or whether two questions are merged can change national statistics, so changes are studied exhaustively before adoption.

6

Standards & Industry References

Survey research is a governed profession with formal standards. Designing to them is what separates defensible research from a casual poll.

Total Survey Error: question design sits in the measurement branch Total Survey Error Representation (coverage, sampling, nonresponse) Measurement (the questions & answers) wording · scales · order · respondent & interviewer effects (Chapters 61–66)
Standard / bodyWhat it covers
AAPORAmerican Association for Public Opinion Research, code of ethics and disclosure standards; response-rate definitions (RR1–RR6)
ESOMARglobal market, opinion and social research, the ICC/ESOMAR International Code
ISO 20252:2019international quality standard for market, opinion and social research processes
Total Survey Error (TSE)the framework (Groves et al.) organizing every error source, design choices map onto it
The Psychology of Survey ResponseTourangeau, Rips & Rasinski (2000), the cognitive model of answering
Tailored Design MethodDillman, the practitioner standard for mail, web, and mixed-mode surveys
πŸ“
The designer's checklist

Before a questionnaire goes live: every item maps to a research objective; each question measures one thing in neutral language; scales are balanced, labeled, and matched to the construct; answer options are mutually exclusive and exhaustive; sensitive items and demographics come last; the instrument has been pretested; and the methodology is documented to AAPOR/ESOMAR/ISO 20252 standards. Good questions are designed, tested, and standardized, never improvised.

πŸŽ“ Key Takeaways

  • βœ“Operationalize first: objective → construct → dimensions → items; aim for validity (right thing) and reliability (consistently).
  • βœ“Match the scale to the construct: dichotomous, Likert, semantic differential, numeric/NPS, frequency, or ranking, and design points, labels, and balance deliberately.
  • βœ“Word to avoid bias: no leading, loaded, double-barreled, or vague items; keep options mutually exclusive and exhaustive.
  • βœ“Structure and pretest: easy → topic → sensitive → demographics; mind order effects; always pilot and run cognitive interviews.
  • βœ“Real instruments and standards: NPS, PHQ-9, Pew experiments, and Census testing show the craft; AAPOR, ESOMAR, ISO 20252, and Total Survey Error govern it.
7

Quiz: Test Yourself

Eight quick questions on survey and questionnaire design. Answer them, hit Check Answers, and keep refining until you score 100%. Your progress is saved, so you can hop back to the chapter and return anytime.

🏁
That completes Sampling & Data Collection

You can now gather data worth analyzing: make the case for sampling, draw representative samples, size a study, design for causation, recognize collection bias, and write survey questions that measure what you intend. With trustworthy data in hand, Estimation & Confidence Intervals turns to drawing rigorous conclusions from it, beginning with point versus interval estimation.