Introduction to Social Research Methodology

Questionnaire design and survey fieldwork

Ben Stanley

Department of Social Sciences, SWPS University

December 15, 2026

Today’s lecture

  • Asking questions — the rules of question wording, and the many ways a single item can go wrong
  • Response formats — Likert items, category design, and the “don’t know” dilemma
  • Structure, flow and layout — question order, sensitive questions, and how a questionnaire hangs together
  • From items to composite measures — the essentials of indexes, scales, and Likert scaling in practice
  • Pretesting and fieldwork — piloting the instrument and running a PAPI or CAWI survey in the field
  • By the end, you should be able to write, assemble, test, and field a questionnaire — and diagnose someone else’s bad one

Asking questions

The questionnaire as measurement instrument

  • A questionnaire is an instrument designed to elicit information that will be useful for analysis (Babbie)
  • It is where operationalisation becomes concrete: every abstract concept from session 8 — engagement, satisfaction, commitment — ends up as specific wording on a page or a screen
  • Standardisation is the whole point: every respondent gets the same questions in the same form, so answers can be aggregated and compared
  • That strength is also the risk: a flawed question is asked flawlessly of everyone, and no amount of analysis afterwards can repair it
  • Questions may take the form of direct questions or of statements respondents agree or disagree with

Running example: Meridian’s board wants an all-staff survey on engagement and turnover intentions. Every slide today feeds into that instrument.

Open-ended and closed-ended questions

  • Open-ended questions — the respondent provides their own answer in their own words
    • “What would most improve your day-to-day work at Meridian?” followed by a text box
    • Rich, unanticipated answers; but costly to code, and quality varies with respondents’ willingness to write
  • Closed-ended questions — the respondent selects from a list the researcher provides
    • “How satisfied are you with your current job?” on a 0–10 scale
    • Uniform, instantly analysable; but only as good as the list of options offered
  • Most questionnaires are predominantly closed-ended, with one or two open questions where discovery matters

Rule of thumb: close what you understand well enough to list; open what you are still exploring.

Make items clear, short, and relevant

  • Clear: items must be unambiguous to the respondent, not just to you
    • “How much do you trust management?” — the store manager? Regional management? The board? Specify the referent
  • Short: assume respondents read quickly and answer quickly; an item that must be studied will be misread
  • Relevant: ask only what respondents have thought about and can meaningfully answer
    • Asking warehouse staff detailed questions about the e-commerce checkout flow yields noise, not data
  • Avoid jargon: “How satisfied are you with our omnichannel fulfilment KPIs?” measures exposure to management vocabulary, not satisfaction

Ask one thing at a time

  • A double-barrelled question asks for a single answer to what is really two (or more) questions
  • “Should Meridian cut store costs and invest more in e-commerce?” — some agree with both, some with neither, and some with exactly one half
    • Whatever they tick, you cannot tell which
  • The tell-tale sign is almost always the word and (sometimes or) joining two distinct objects or propositions
  • The fix is mechanical: split the item into one question per proposition
  • Double-barrelled items are among the most common flaws in real organisational surveys — satisfaction items are especially prone (“your pay and benefits”, “your manager and colleagues”)

Avoid leading and loaded wording

  • Bias is any property of a question that encourages respondents to answer in a particular way
  • Leading questions signal the expected answer: “Don’t you agree that our generous benefits package is a strength?”
  • Prestige bias: attaching a position to a respected person or body — “As the CEO has argued…” — pulls answers toward (or against) it
  • Loaded terms: wording alone shifts results — Americans support “assistance to the poor” far more than “welfare”, the same policy in different clothes
  • Social desirability: people answer through a filter of what makes them look good — respected behaviours are over-reported, disreputable ones under-reported
    • Neutral wording and guaranteed anonymity are the main defences

Avoid negations

  • A negation in an item paves the way for easy misinterpretation
  • Babbie’s classic example: asked to agree or disagree that “The United States should not recognize Cuba”, many respondents read over the not — some agree meaning recognition, others agree meaning the opposite
  • Double negatives are worse still: “I do not think Meridian is failing to support its staff — agree or disagree?” is a small logic puzzle, not a question
  • Disagreeing with a negative statement requires mental gymnastics that respondents performing at speed will not do
  • The fix: state items positively and let the response scale carry the direction

Can respondents answer?

  • Respondents must be competent to answer: able to recall the fact, or to hold a genuine attitude on the topic
  • Unanswerable recall demands produce fabricated precision: “How many customer interactions did you have in the past 12 months?” — nobody knows; everyone will still write a number
    • Bound the recall period (“in a typical week”) or offer banded categories
  • Attitude questions on topics respondents have never considered manufacture non-attitudes: answers that look like data but reflect nothing
  • Asking teenagers about K-pop yields real variation; asking 65-year-olds the same question mostly measures whether they know it exists — relevance depends on the population

Will respondents answer?

  • Competence is not enough — respondents must also be willing
  • Income is the classic case: “How much do you earn each month?” invariably produces refusals
    • The indirect fix: “On a scale of 1 to 7, where 1 is the least well-off and 7 the most well-off, where would you place yourself?”
  • In organisational surveys the sensitive topics are different but the logic is identical: intentions to quit, opinions of one’s direct manager, experiences of unfair treatment
  • Willingness depends on trust: anonymity, confidentiality, and who is seen to run the survey matter as much as wording
    • An employee survey visibly run by HR gets different answers from one run by an external agency

Response formats

Likert response items

  • The workhorse of attitude measurement: a statement plus a symmetric agree–disagree scale
    • Strongly disagree — Disagree — Neither agree nor disagree — Agree — Strongly agree
  • Typically 5 or 7 points; the response categories are the same for every item, so a whole battery can share one grid
  • Design decisions that matter:
    • Odd or even? An odd number offers a genuine midpoint; an even number forces a direction
    • Label every point, not just the ends — labelled points are interpreted more consistently
    • Keep the scale balanced: as many positive as negative options, symmetric wording
  • Alternatives exist for other jobs: frequency scales (never — always), satisfaction scales, 0–10 numeric ratings

Exhaustive and mutually exclusive categories

  • Closed-ended response lists must satisfy two requirements — the same two you met for variable categories in session 8
  • Exhaustive: every respondent must have a category that fits
    • Tenure options “0–2 years / 3–5 years / 6–10 years” strand anyone with 11+ years — add the top band or an “other”
  • Mutually exclusive: no respondent should fit two categories at once
    • “0–2 / 2–5 / 5–10” — where does exactly 2 years go? Overlapping boundaries force arbitrary choices
  • Both failures corrupt data silently: respondents still tick something, and the damage only surfaces (if ever) at analysis

Checklist habit: for every closed item ask — can everyone answer? can anyone answer twice?

To offer “don’t know” or not?

  • The hardest recurring judgement call in response design
  • Offering “don’t know” / “not applicable”:
    • respects respondents who genuinely cannot answer — new hires rating “career development over the past three years”, employees with no contact with a department they are asked to assess
    • but offers an easy exit that some respondents with real attitudes will take
  • Withholding it forces an answer — and forced answers from the genuinely ignorant are noise dressed as data
  • A practical rule: offer an explicit opt-out wherever a meaningful share of respondents truly may have no basis to answer; withhold it for attitudes everyone in the population plausibly holds
  • Never make an opinion item required in CAWI without an opt-out: you will harvest fabricated answers or abandoned questionnaires

Structure, flow and layout

Question order and context effects

  • The order of items affects the answers: the appearance of one question can change the answers given to later ones
  • Babbie’s managerial example: ask a series of questions about what makes a good manager, then ask respondents to rate their own manager — the ratings shift, positively or negatively, because you have just supplied the standard of comparison
  • General satisfaction items answered after a long list of specific complaints come out lower than when asked first
  • There is no neutral order — but there is a consistent one: within one survey, everyone must get the same context
  • Defences: ask general before specific; separate items that contaminate each other; randomise blocks where the mode allows it

Placing sensitive questions

  • Placement is a wording decision by other means
  • Never open with sensitive items: turnover intentions or manager ratings as question 1 trigger abandonment before trust is built
  • Place sensitive items late in the questionnaire, after routine items have established rhythm and demonstrated seriousness
  • Precede them with a short reassurance: “Your answers are anonymous and will only be reported in aggregate”
  • Demographics — age, tenure, role — go at the end: they feel like surveillance at the start of an anonymous employee survey, and they are the items least harmed by fatigue

Opening, grouping and instructions

  • Opening questions should be easy, interesting, and applicable to everyone — they set the tone and build commitment
    • A good Meridian opener: overall job satisfaction; a bad one: a four-row grid about payroll administration
  • Group items by topic, and introduce each subsection with a short statement of its content and purpose
    • “The next few questions are about your day-to-day work…” — these introductions make the questionnaire feel coherent rather than chaotic, and put respondents in the right frame of mind
  • Instructions must be explicit wherever the task changes: tick one vs tick all that apply; whether brief or extended open answers are wanted
  • Contingency (filter) questions route respondents past what does not apply: “Do you supervise other employees?” — only supervisors see the supervision block

Layout and length

  • The format of a questionnaire is just as important as the wording of its questions (Babbie)
  • Uncluttered layout, adequate spacing, one question at a time — a cramped page produces skipped items and mis-ticks; in CAWI, avoid oversized grids that collapse on a phone screen
  • Length is a wording budget: every marginal item costs response rate and answer quality at the end of the questionnaire
  • For a CAWI employee survey, 10–15 minutes is a realistic ceiling; state the expected length in the invitation and be honest about it
  • The discipline that enforces brevity: for every item, name the analysis it will feed — if you cannot say how you will use the answer, cut the question

From items to composite measures

Why one item is rarely enough

  • Most concepts worth measuring — engagement, commitment, satisfaction — are too broad for any single item to capture
  • A single item is hostage to its own wording quirks; several items on the same concept let the quirks average out
  • The solution is a composite measure: combine multiple indicators of the concept into one score per respondent
  • This is the direct continuation of operationalisation: session 8 gave one concept several indicators; today those indicators become questionnaire items, and the items are combined into a measurement
  • Two main families of composite measure: indexes and scales

Indexes vs scales

  • An index accumulates indicators: scores are summed or averaged, and every item counts equally toward the total
    • The Human Development Index combines life expectancy, schooling, and income per capita into one number per country
    • An engagement index might sum “yes” answers across six workplace behaviours
  • A scale exploits intensity structure among items: some items represent stronger degrees of the attitude than others, so response patterns, not just totals, carry information
    • Agreeing that “I would recommend Meridian to a friend” says more than agreeing that “my workplace is acceptable”
  • In practice the distinction blurs — and the term “Likert scale” is used loosely for what is, strictly, summated rating

Likert scaling in practice

  • The standard recipe for a composite attitude measure:
    1. Write a battery of Likert items (typically 4–8) all tapping the same concept
    2. Include some reversed items (“I often think about leaving Meridian”) to catch respondents who tick the same column all the way down
    3. Reverse-code those items after collection, so that a high number always means the same direction
    4. Sum or average across items into one score per respondent
  • Before trusting the score, check internal consistency — do the items hang together? (Cronbach’s alpha is the usual statistic; session 13 shows the mechanics)
  • The composite score is what gets analysed: compared across stores, tracked over time, correlated with turnover

Typologies, in one slide

  • A third tool combines variables not into a score but into categories: a typology classifies cases by the intersection of two or more dimensions
  • Crossing satisfaction (high/low) with intention to stay (high/low) yields four types of employee — including the two that should worry Meridian most: the satisfied leaver and the trapped stayer
  • Typologies are powerful for description and communication, but they are categories, not quantities — you cannot average a typology
  • Weber’s three types of authority and Merton’s modes of adaptation are the classic social-science examples
  • For this course: know that the option exists; indexes and scales will do most of your work

Pretesting and fieldwork

Pretesting and piloting

  • No questionnaire should meet real respondents untested — every defect from today’s first four sections is cheap to fix before fieldwork and expensive after
  • Pretesting (small scale): have a handful of people from the target population complete the questionnaire while thinking aloud
    • You are hunting misread items, ambiguous referents, missing categories, confusing instructions — not collecting data
  • Piloting (fuller dress rehearsal): run the whole procedure — invitation, questionnaire, data capture — on a small sample and inspect the resulting data
    • Check timings, completion rates, item non-response, and whether answers show usable variation
  • The iron law of pretesting: the pilot always finds something — a pilot that finds nothing was not looking

Survey modes: PAPI, CAWI, CATI, CAPI

  • The mode is how the questionnaire reaches the respondent — and it shapes cost, speed, reach, and data quality
  • PAPI — paper-and-pencil: printed questionnaires completed by hand; the traditional pre-computer mode
  • CAWI — computer-assisted web interviewing: an online questionnaire, increasingly completed on smartphones
  • CATI — computer-assisted telephone interviewing: an interviewer reads a scripted questionnaire over the phone and keys answers directly
  • CAPI — computer-assisted personal interviewing: face-to-face, with the interviewer recording answers on a tablet or laptop
  • This session’s practical pair is PAPI and CAWI — the two modes a junior team can realistically run itself

PAPI and CAWI compared

PAPI CAWI
Cost Printing, distribution, manual data entry Near-zero marginal cost per respondent
Speed Slow: collection plus keying-in Answers arrive on the server instantly
Reach Works without internet access or digital skills — e.g. on a shop floor Only reaches the online; skews younger and more digital
Routing & checks None: skips and validations depend on the respondent Automatic skip logic, validations, required formats
Errors Data-entry errors added at keying stage No input stage; errors easy to correct in the form
Social desirability Low if self-completed anonymously Low: no interviewer present

Running example: Meridian mixes modes — CAWI for office and e-commerce staff, PAPI on paper in stores for staff without company e-mail. Mixing modes means checking that the two forms are truly equivalent.

Organising CAWI fieldwork: invitations and reminders

  • CAWI fieldwork is won or lost in the contact strategy, not the questionnaire
  • The invitation e-mail must state: who is running the survey, its purpose, how long it takes (honestly), that participation is voluntary, how anonymity is protected, the deadline, and the link
    • Sender matters: an invitation endorsed by a trusted figure outperforms an anonymous system address
  • Reminders do heavy lifting: each wave typically recovers a substantial share of remaining non-respondents
    • A standard schedule: invitation → reminder after ~1 week → final “closing soon” reminder a few days before the deadline
    • Remind non-respondents only where the platform allows; a “thank you or reminder” wording covers anonymous designs
  • A fieldwork window of 2–3 weeks balances reach against losing momentum

Monitoring fieldwork

  • Fieldwork is not fire-and-forget: monitor while the field is open, because problems found on day 3 can be fixed — problems found at close cannot
  • Response rate over time: plot completes per day; the spikes should line up with the invitation and each reminder
  • Sample composition: compare respondents so far against the known population — are stores responding as well as head office? Are all regions represented?
  • Completion behaviour: where do people abandon the questionnaire? A consistent drop-off point marks a broken or exhausting item
  • Item nonresponse: an item that many respondents skip is telling you it is unclear, intrusive, or inapplicable
  • For PAPI: monitor returns per site and chase collection points that go quiet

Spotting and fixing problems mid-field

  • Some mid-field signals and their standard responses:
    • Low overall response → extend the window, add a reminder wave, re-engage endorsers — never quietly relax who counts as the target population
    • A skewed sample → targeted follow-up on the under-responding groups (the night shift, specific stores), not just more blanket reminders
    • Straight-lining — identical answers all the way down a grid → flag those cases for scrutiny at analysis; in future designs, shorten grids and keep reversed items
    • A broken skip or a garbled item discovered mid-field → fix it, record the date and nature of the change, and treat before/after answers to that item as potentially non-comparable
  • The discipline throughout: keep a fieldwork log — every decision, every change, every anomaly, with dates

Conclusion

Conclusion

  • A questionnaire is a measurement instrument: every rule today — one thing at a time, neutral wording, no negations, answerable and askable questions — protects the link between concept and datum
  • Response formats and category design fail silently; exhaustive, mutually exclusive, and honestly opt-out-able is the standard
  • Order, grouping, and layout are part of measurement, not decoration — context effects are real, and sensitive items need shelter
  • Single items are fragile; indexes and Likert scaling turn batteries of items into robust composite scores
  • Pretest always, pilot when you can, monitor while the field is open — the pilot always finds something
  • Next session: with the instrument fielded, the data arrive — and session 13 turns to analysing them
  • Questions and discussion are welcome

Exercise

Today’s exercise: Rescue this survey

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-11-exercise

Study guide

Full summary of this session, for revision: Questionnaire design and survey fieldwork

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-11-handout