The previous chapters showed how sampling and collection introduce bias. Surveys add one more, earlier, source: the questions themselves. A poorly worded item biases every answer before a single response is recorded. Good questionnaire design is the craft of measuring what you actually intend to measure.
This is a reference chapter: no code, just principles, examples, and standards. We move from deciding what to ask → choosing scales → wording items → structuring the questionnaire → four real-world examples → the standards bodies that govern professional practice.
From Research Goal to Question
Good surveys are not written question-first. They start from the research objective and work down through a deliberate process called operationalization: defining the abstract concept, breaking it into measurable dimensions, and only then writing items.
Each step answers a question. What do we want to learn? sets the objective. What exactly do we mean? defines the construct, often multidimensional. What observable things reveal it? gives the dimensions. How do we ask? produces the items. Skipping to items is the most common failure: you end up with a pile of questions that do not add up to the thing you meant to measure.
Two properties decide whether items work. Validity asks "does it measure the right thing?" (content, construct, and criterion validity). Reliability asks "does it measure consistently?" Multi-item scales are often checked with Cronbach's alpha (a value of 0.70 or higher is the common threshold for acceptable internal consistency). A measure can be reliable yet invalid, so you need both.
Question Types & Response Scales
The response format must match what you are measuring. Choosing the wrong scale, a yes/no for an attitude, a 5-point agreement for a behavior count, throws away information or fabricates precision.
| Type | Best for | Measurement level & notes |
|---|---|---|
| Open-ended | Exploration, unanticipated answers, "why" | rich but costly to code; use sparingly |
| Dichotomous | Clear yes/no facts, screening | nominal; fast but low information |
| Multiple choice | Discrete categories | nominal; options must be mutually exclusive & exhaustive |
| Likert | Attitudes, agreement, satisfaction | ordinal; usually 5 or 7 points, labeled |
| Semantic differential | Image, perception, brand feel | bipolar adjectives; treated as interval-like |
| Numeric / rating | Magnitude, likelihood (e.g., NPS 0–10) | interval-like; anchor the endpoints clearly |
| Frequency | How often a behavior occurs | use specific ranges, not vague "often/rarely" |
| Ranking | Relative priorities among a few items | ordinal; hard beyond ~5–7 items |
Scale design choices matter as much as the type. An odd number of points offers a neutral midpoint; an even number forces a lean (a "forced choice"). Label every point where you can, not just the ends, so respondents interpret them the same way. Keep scales balanced (equal positive and negative options) and decide deliberately whether to offer "Don't know" or "Not applicable", omitting it pushes people to guess, while including it can invite satisficing. The original 5-point agreement format traces to Rensis Likert's 1932 monograph; the bipolar semantic differential to Osgood and colleagues in 1957.
Writing Questions That Don't Bias the Answer
Wording is where most surveys fail. Respondents work through four cognitive steps for every question, they comprehend it, retrieve the relevant memory, form a judgment, and map it to a response (Tourangeau, Rips & Rasinski, 2000). A flaw at any step distorts the answer.
"How much did you enjoy our excellent new app?"
presumes enjoyment
"How would you rate your experience with our app?"
Very poor … Very good
"Was the staff friendly and knowledgeable?"
two questions, one answer
"Was the staff friendly?" & "Was the staff knowledgeable?"
one idea per item
"Don't you agree the policy shouldn't be repealed?"
confusing, pressuring
"Do you favor or oppose the policy?"
offers both sides equally
"Do you exercise often?"
"often" means different things
"In the past 7 days, how many days did you exercise?"
0 · 1–2 · 3–4 · 5+
| Pitfall | Why it biases | Fix |
|---|---|---|
| Leading / loaded wording | steers toward an answer | neutral, value-free language |
| Double-barreled | two ideas, one answer slot | one concept per question |
| Double negatives | hard to parse, errors | state it positively |
| Ambiguous terms | everyone reads them differently | define terms; use concrete ranges |
| Jargon / reading level | excludes or confuses | plain language, ~grade 6–8 |
| Acquiescence bias | tendency to agree | mix positively/negatively keyed items |
| Social desirability | flattering self-reports | anonymity, indirect or list techniques |
| Recall error | memory fades and distorts | short, bounded reference periods |
| Non-exhaustive options | respondent has no valid answer | cover all cases; add "Other" |
A famous demonstration of how much wording matters: in repeated experiments the U.S. public expressed far more support for spending on "assistance to the poor" than on "welfare", though the two refer to the same programs. The label, not the policy, moved the answer, which is why professional pollsters test alternative wordings before fielding.
Questionnaire Structure & Best Practices
Individual questions are necessary but not sufficient; their order and flow shape responses too. The standard arc moves from easy and engaging, through the substance, to sensitive items and demographics at the very end.
| Best practice | Why |
|---|---|
| Open with easy, non-threatening items | builds rapport and commitment, lowers drop-off |
| Group related questions; logical flow | reduces cognitive load and confusion |
| Sensitive questions & demographics last | trust is established; refusals do not sink the whole survey |
| Watch order & context effects | an earlier question can prime later answers; randomize when apt |
| Randomize answer-option order | counters primacy/recency response-order bias |
| Keep it short; minimize burden | fatigue causes satisficing and break-offs |
| Design mobile-first | most surveys are now taken on phones; avoid wide grids |
| Pretest: pilot & cognitive interviews | catches misreadings before they cost you the whole sample |
Before fielding, run a pilot and ideally cognitive interviews, where a few respondents think aloud as they answer, so you hear how they actually interpret each item. It is far cheaper to fix a confusing question with five testers than to discover, after 2,000 responses, that half of them misunderstood it. Dillman's Tailored Design Method is the standard practitioner reference for the full process.
Four Real-World Examples
Theory becomes concrete in widely used, well-documented instruments. Each of these makes deliberate design choices, and illustrates a different lesson. For every example below you can view a sample questionnaire (PDF) of 10–15 questions in varied formats and download the matching data sheet (Excel), laid out the analysis-ready way a statistician would receive it: one row per respondent, coded columns, plus a codebook and analyst notes.
Net Promoter Score (NPS)
Introduced by Fred Reichheld (Bain & Company) in the Harvard Business Review, 2003, NPS asks a single question: "How likely are you to recommend us to a friend or colleague?" on an 0–10 numeric scale. Respondents are bucketed into Promoters (9–10), Passives (7–8), and Detractors (0–6), and the score is %Promoters − %Detractors (ranging −100 to +100). Lesson: a carefully chosen single-item scale can be a powerful, comparable metric, but the 0–10 anchoring and the bucketing rules are essential and must be applied consistently to compare across time or companies.
PHQ-9 Depression Screener
The Patient Health Questionnaire-9 (Kroenke, Spitzer & Williams, 2001) is a validated nine-item instrument asking how often, over the last two weeks, a person was bothered by each symptom, on a four-point frequency scale: "Not at all" (0), "Several days" (1), "More than half the days" (2), "Nearly every day" (3). Item scores sum to 0–27, with cutoffs at 5, 10, 15, and 20 marking mild, moderate, moderately severe, and severe depression. Lesson: a validated, reliable instrument with specific frequency anchors and a published scoring rule yields a clinically actionable, comparable measure, the payoff of rigorous design and testing.
Pew Research Center wording experiments
Pew, a gold standard in survey research, routinely runs split-sample wording experiments: half the respondents see one phrasing, half another. Their work documented how "assistance to the poor" versus "welfare" can swing apparent support by tens of points, and they deliberately offer balanced response options ("favor or oppose") and rotate answer order. Lesson: question wording is itself a measurable source of error; professional pollsters test alternatives and report methodology transparently rather than trusting a single phrasing.
U.S. Census Bureau & the American Community Survey
Federal surveys are designed and revised through years of cognitive testing and field experiments. The Census Bureau has repeatedly tested its race and Hispanic-origin questions, formatting, examples, and combined-versus-separate items, because small wording and layout changes measurably shift how millions of people respond. Lesson: at scale, format and pretesting are as consequential as wording; even the order of examples or whether two questions are merged can change national statistics, so changes are studied exhaustively before adoption.
Standards & Industry References
Survey research is a governed profession with formal standards. Designing to them is what separates defensible research from a casual poll.
| Standard / body | What it covers |
|---|---|
| AAPOR | American Association for Public Opinion Research, code of ethics and disclosure standards; response-rate definitions (RR1–RR6) |
| ESOMAR | global market, opinion and social research, the ICC/ESOMAR International Code |
| ISO 20252:2019 | international quality standard for market, opinion and social research processes |
| Total Survey Error (TSE) | the framework (Groves et al.) organizing every error source, design choices map onto it |
| The Psychology of Survey Response | Tourangeau, Rips & Rasinski (2000), the cognitive model of answering |
| Tailored Design Method | Dillman, the practitioner standard for mail, web, and mixed-mode surveys |
Before a questionnaire goes live: every item maps to a research objective; each question measures one thing in neutral language; scales are balanced, labeled, and matched to the construct; answer options are mutually exclusive and exhaustive; sensitive items and demographics come last; the instrument has been pretested; and the methodology is documented to AAPOR/ESOMAR/ISO 20252 standards. Good questions are designed, tested, and standardized, never improvised.
π Key Takeaways
- βOperationalize first: objective → construct → dimensions → items; aim for validity (right thing) and reliability (consistently).
- βMatch the scale to the construct: dichotomous, Likert, semantic differential, numeric/NPS, frequency, or ranking, and design points, labels, and balance deliberately.
- βWord to avoid bias: no leading, loaded, double-barreled, or vague items; keep options mutually exclusive and exhaustive.
- βStructure and pretest: easy → topic → sensitive → demographics; mind order effects; always pilot and run cognitive interviews.
- βReal instruments and standards: NPS, PHQ-9, Pew experiments, and Census testing show the craft; AAPOR, ESOMAR, ISO 20252, and Total Survey Error govern it.
Quiz: Test Yourself
Eight quick questions on survey and questionnaire design. Answer them, hit Check Answers, and keep refining until you score 100%. Your progress is saved, so you can hop back to the chapter and return anytime.
You can now gather data worth analyzing: make the case for sampling, draw representative samples, size a study, design for causation, recognize collection bias, and write survey questions that measure what you intend. With trustworthy data in hand, Estimation & Confidence Intervals turns to drawing rigorous conclusions from it, beginning with point versus interval estimation.