Questionnaire design and survey fieldwork
Introduction to Social Research Methodology
Asking questions
A questionnaire is an instrument designed to elicit information that will be useful for analysis (Babbie). It is the point at which operationalisation, discussed in session 8, becomes concrete: every abstract concept a study cares about — engagement, satisfaction, commitment — must eventually be expressed as specific wording on a page or a screen. The defining feature of the questionnaire is standardisation: every respondent receives the same questions in the same form, so that answers can be aggregated and compared. This strength is also the central risk. A flawed question is asked, flawlessly, of every single respondent, and no amount of statistical sophistication afterwards can repair data that a bad item has corrupted at the source. Questionnaire items may take the form of direct questions or of statements with which respondents are asked to agree or disagree.
Open-ended and closed-ended questions
Researchers asking questions have two basic options. Open-ended questions ask the respondent to provide an answer in their own words — for example, “What would most improve your day-to-day work at Meridian?” followed by a text box. They yield rich and sometimes unanticipated answers, but they are costly to code, and the quality of what comes back depends heavily on respondents’ willingness to write. Closed-ended questions ask the respondent to select an answer from a list the researcher provides — for example, “How satisfied are you with your current job?” rated on a scale from 0 to 10. They produce uniform, immediately analysable data, but they are only as good as the list of options offered. Most questionnaires are predominantly closed-ended, with one or two open questions reserved for topics where discovery matters. A useful rule of thumb: close what you understand well enough to list; open what you are still exploring.
Clarity, brevity, and relevance
Three basic qualities apply to every item. First, items must be clear — unambiguous to the respondent, not merely to the researcher. “How much do you trust management?” may seem perfectly clear to its author, but a shop-floor employee cannot tell whether it refers to the store manager, regional management, or the board; the referent must be specified. Second, items must be short. Respondents read quickly and answer quickly; an item that has to be studied in order to be understood will instead be misread. Provide clear, short items that will not be misinterpreted under quick-reading conditions. Third, items must be relevant to the people being asked: questions on topics respondents have not thought about or do not care about produce answers that look like data but mean little. Relevance depends on the population — asking warehouse staff detailed questions about the e-commerce checkout flow yields noise, not measurement. A related failure is jargon: an item such as “How satisfied are you with our omnichannel fulfilment KPIs?” measures respondents’ exposure to management vocabulary rather than their satisfaction.
The rules of question wording
| Rule | The failure it prevents | Example and fix |
|---|---|---|
| Ask one thing at a time | Double-barrelled questions demand a single answer to two propositions | “Should Meridian cut store costs and invest more in e-commerce?” — some respondents endorse one half only, and their answer is uninterpretable. Split into two items. The word and is the tell-tale sign |
| Avoid leading wording | Leading questions signal the expected answer | “Don’t you agree that our generous benefits package is a strength?” invites agreement. State the item neutrally |
| Avoid loaded terms and prestige bias | Emotionally charged words, or association with a respected figure, pull answers | Americans support “assistance to the poor” far more than “welfare” — the same policy in different clothes. “As the CEO has argued…” biases responses toward (or against) the position |
| Beware social desirability | Respondents answer through a filter of what makes them look good | Respected behaviours are over-reported and disreputable ones under-reported. Use neutral wording and credible anonymity |
| Avoid negations | A negation invites misreading | Asked about “The United States should not recognize Cuba”, many respondents read past the not; some agree meaning recognition, others agree meaning the opposite. Double negatives are worse still. State items positively |
| Ask only what can be answered | Unanswerable recall demands produce fabricated precision | “How many customer interactions did you have in the past 12 months?” — nobody knows, but everybody writes a number. Bound the recall period (“in a typical week”) or offer banded categories |
| Ask only what will be answered | Unwilling respondents refuse or distort | “How much do you earn each month?” invariably produces refusals; the indirect alternative asks respondents to place themselves on a 1–7 scale from least to most well-off |
Competence and willingness deserve emphasis because they concern the respondent rather than the wording as such. Respondents must be able to answer: capable of recalling the fact requested, or of holding a genuine attitude on the topic. Attitude questions about topics respondents have never considered manufacture non-attitudes — answers that reflect nothing but the obligation to tick a box. And respondents must be willing to answer: in organisational surveys the sensitive topics are intentions to quit, opinions of one’s direct manager, and experiences of unfair treatment, and willingness depends on trust. Anonymity, confidentiality, and who is seen to run the survey matter as much as the wording itself — an employee survey visibly run by HR receives different answers from one run by an external agency.
Response formats
Likert response items
The workhorse format of attitude measurement is the Likert item: a statement followed by a symmetric agree–disagree response scale, typically Strongly disagree — Disagree — Neither agree nor disagree — Agree — Strongly agree. Likert scales usually have five or seven points, and because the response categories are identical for every item, a whole battery of statements can share a single grid. Several design decisions matter. An odd number of points offers a genuine midpoint, while an even number forces respondents to lean one way — a deliberate choice, not an accident. Every point should carry a verbal label, not just the endpoints, because fully labelled scales are interpreted more consistently. And the scale must be balanced: as many positive as negative options, with symmetric wording. Likert items are not the only format — frequency scales (never to always), satisfaction scales, and 0–10 numeric ratings serve other jobs — but they are the default for attitude statements.
Exhaustive and mutually exclusive categories
Closed-ended response lists must satisfy the same two requirements introduced for variable categories in session 8. They must be exhaustive: every respondent must find a category that fits. Tenure options of “0–2 years / 3–5 years / 6–10 years” strand every employee with more than ten years of service; the list needs a top band or an “other” option. And they must be mutually exclusive: no respondent should fit two categories at once. Options of “0–2 / 2–5 / 5–10” leave a respondent with exactly two years of service facing an arbitrary choice between two boxes. Both failures corrupt data silently — respondents still tick something — and the damage surfaces, if at all, only at analysis. The habit to build: for every closed item, ask whether everyone can answer, and whether anyone can answer twice.
The “don’t know” dilemma
Whether to offer an explicit “don’t know” or “not applicable” option is the hardest recurring judgement call in response design. Offering it respects respondents who genuinely cannot answer — a new hire asked to rate “career development over the past three years”, or an employee asked to assess a department they never deal with. But it also offers an easy exit that some respondents with real attitudes will take. Withholding it forces an answer — and forced answers from the genuinely ignorant are noise dressed as data. A practical rule: offer the opt-out wherever a meaningful share of respondents may truly have no basis to answer; withhold it for attitudes that everyone in the population plausibly holds. In CAWI, never make an opinion item required without an opt-out: the result is either fabricated answers or abandoned questionnaires.
Structure, flow and layout
Question order and context effects
The order in which items appear affects the answers given: the appearance of one question can change responses to later ones. Babbie’s managerial example makes the point directly — if respondents first answer a series of questions about what makes a good manager and are then asked to rate their own manager, the ratings shift, positively or negatively, because the earlier questions supplied the standard of comparison. Similarly, a general satisfaction item answered after a long list of specific complaints comes out lower than the same item asked first. There is no neutral order, but there is a consistent one: within a single survey, every respondent must encounter the same context. The practical defences are to ask general questions before specific ones, to separate items likely to contaminate each other, and to randomise blocks of items where the survey mode allows it.
Sensitive questions, openings, grouping, and instructions
Placement is a wording decision by other means. Sensitive items — turnover intentions, ratings of one’s manager, experiences of unfair treatment — should never open a questionnaire; asked first, they trigger abandonment before any trust has been built. They belong late in the instrument, after routine items have established rhythm, and they benefit from a short preceding reassurance that answers are anonymous and reported only in aggregate. Demographic items (age, tenure, role) go at the end: at the start of an anonymous employee survey they feel like surveillance, and they are the items least harmed by respondent fatigue.
Opening questions should be easy, interesting, and applicable to everyone, because they set the tone and build commitment to finishing. Items should be grouped by topic, with each subsection introduced by a short statement of its content and purpose (“The next few questions are about your day-to-day work…”). Such introductions make the questionnaire feel coherent rather than chaotic and put respondents in the proper frame of mind. Instructions must be explicit wherever the task changes — tick one versus tick all that apply, and whether brief or extended open answers are expected. Contingency (filter) questions route respondents past what does not apply to them: only those who answer “yes” to “Do you supervise other employees?” should see the block about supervision.
Layout and length
The format of a questionnaire is just as important as the nature and wording of its questions. A questionnaire should be adequately spaced, with an uncluttered layout and one question presented at a time: a cramped page produces skipped items and mis-ticks, and in CAWI, oversized answer grids collapse on a phone screen. Length is a budget: every marginal item costs response rate and degrades answer quality toward the end of the instrument. For a CAWI employee survey, ten to fifteen minutes is a realistic ceiling, and the invitation should state the expected length honestly. The discipline that enforces brevity is to name, for every item, the analysis it will feed — a question whose use cannot be stated should be cut.
From items to composite measures
Why one item is rarely enough
Most concepts worth measuring — engagement, commitment, satisfaction — are too broad for any single item to capture, and a single item is hostage to its own wording quirks. Several items on the same concept allow those quirks to average out. The solution is the composite measure: multiple indicators of one concept combined into a single score per respondent. This is the direct continuation of operationalisation — session 8 gave one concept several indicators; here those indicators become questionnaire items, and the items are combined into a measurement.
Indexes, scales, and typologies
| Tool | How it combines | Character | Example |
|---|---|---|---|
| Index | Accumulates indicators: scores are summed or averaged, every item counting equally | Composite quantity | The Human Development Index combines life expectancy, schooling, and income per capita into one number per country; an engagement index might sum “yes” answers across six workplace behaviours |
| Scale | Exploits intensity structure: some items represent stronger degrees of the attitude, so response patterns carry information beyond the total | Ordinal measure of degree | Agreeing that “I would recommend Meridian to a friend” indicates more engagement than agreeing that “my workplace is acceptable” |
| Typology | Classifies cases by the intersection of two or more dimensions | Categories, not quantities | Crossing satisfaction (high/low) with intention to stay (high/low) yields four employee types — including the satisfied leaver and the trapped stayer |
In practice the index–scale distinction blurs, and the term “Likert scale” is used loosely for what is, strictly speaking, a summated rating. Typologies are powerful for description and communication — Weber’s three types of authority and Merton’s modes of adaptation are the classic examples — but they are categories and cannot be averaged. For this course, indexes and Likert-type summated scales will do most of the work.
Likert scaling in practice
The standard recipe for building a composite attitude measure has four steps. First, write a battery of Likert items — typically four to eight — all tapping the same concept. Second, include some reversed items (for engagement, e.g. “I often think about leaving Meridian”), which catch respondents who tick the same column all the way down. Third, after collection, reverse-code those items so that a high number always means the same direction. Fourth, sum or average across the items to produce one score per respondent. Before the score is trusted, its internal consistency should be checked — do the items hang together as measures of one thing? Cronbach’s alpha is the usual statistic, and session 13 shows the mechanics. The composite score is then what gets analysed: compared across stores, tracked over time, and correlated with turnover.
Pretesting and fieldwork
Pretesting and piloting
No questionnaire should meet real respondents untested. Every defect catalogued above is cheap to fix before fieldwork and expensive — often irreparable — after it. Pretesting is the small-scale version: a handful of people from the target population complete the questionnaire while thinking aloud, and the researcher hunts for misread items, ambiguous referents, missing categories, and confusing instructions. This is not data collection; it is instrument inspection. Piloting is the fuller dress rehearsal: the whole procedure — invitation, questionnaire, and data capture — is run on a small sample, and the resulting data are inspected for timings, completion rates, item non-response, and whether answers show usable variation. The iron law of pretesting is that the pilot always finds something; a pilot that finds nothing was not looking.
Survey modes
The mode is the channel through which the questionnaire reaches the respondent, and it shapes cost, speed, reach, and data quality.
| Mode | Full name | Character |
|---|---|---|
| PAPI | Paper-and-pencil interviewing | Printed questionnaires completed by hand; the traditional pre-computer mode |
| CAWI | Computer-assisted web interviewing | An online questionnaire or web page provided to the respondent, increasingly completed on smartphones |
| CATI | Computer-assisted telephone interviewing | An interviewer reads a scripted questionnaire over the phone and keys answers directly into the system |
| CAPI | Computer-assisted personal interviewing | Face-to-face interviewing with answers recorded on a tablet, phone, or laptop |
The practical pair for a junior research team is PAPI and CAWI. Their trade-offs:
| PAPI | CAWI | |
|---|---|---|
| Cost | Printing, distribution, and manual data entry | Near-zero marginal cost per respondent; no printing, no surveyors |
| Speed | Slow: physical collection followed by keying-in | Answers arrive on the server instantly, allowing continuous tracking |
| Reach | Works without internet access or digital skills — for example, on a shop floor | Reaches only the online; respondent pools skew younger and more digitally fluent |
| Routing and checks | None: skips and validations depend on the respondent following instructions | Automatic skip logic, validations, and required formats; multimedia possible |
| Errors | Data-entry errors added at the keying stage | No input stage; errors in the form itself are easy to correct |
| Social desirability | Low if self-completed anonymously | Low: no interviewer is present to perform for |
Meridian’s fieldwork mixes modes: CAWI for office and e-commerce staff, and PAPI on paper in stores for staff without company e-mail. Mixing modes requires checking that the two versions of the instrument are genuinely equivalent.
Organising CAWI fieldwork
CAWI fieldwork is won or lost in the contact strategy rather than the questionnaire itself. The invitation e-mail must state who is running the survey, its purpose, how long it takes (honestly), that participation is voluntary, how anonymity is protected, the deadline, and the link. The sender matters: an invitation endorsed by a trusted figure outperforms one from an anonymous system address. Reminders do heavy lifting — each wave typically recovers a substantial share of remaining non-respondents. A standard schedule runs: invitation, a reminder after about one week, and a final “closing soon” reminder a few days before the deadline. Where the platform allows, reminders should go to non-respondents only; where the design is fully anonymous, a “thank you or reminder” wording covers everyone. A fieldwork window of two to three weeks balances reach against loss of momentum.
Monitoring fieldwork and spotting problems mid-field
Fieldwork is not fire-and-forget. Problems discovered on day three can be fixed; problems discovered at close cannot. Four things should be watched while the field is open. The response rate over time: completes per day should spike with the invitation and each reminder wave. The composition of the sample so far, compared against the known population: are stores responding as well as head office, and are all regions represented? Completion behaviour: a consistent abandonment point in the questionnaire marks a broken or exhausting item. And item nonresponse: an item that many respondents skip is signalling that it is unclear, intrusive, or inapplicable. For PAPI, the equivalent is monitoring returns per site and chasing collection points that go quiet.
| Mid-field signal | Standard response |
|---|---|
| Low overall response | Extend the window, add a reminder wave, re-engage endorsers — never quietly relax the definition of the target population |
| Skewed sample composition | Targeted follow-up on under-responding groups (the night shift, particular stores), not just more blanket reminders |
| Straight-lining (identical answers down a grid) | Flag the cases for scrutiny at analysis; in future designs, shorten grids and retain reversed items |
| A broken skip or garbled item found mid-field | Fix it, record the date and nature of the change, and treat before/after answers to that item as potentially non-comparable |
The discipline throughout is the fieldwork log: every decision, every change, and every anomaly, recorded with dates.
Conclusion
A questionnaire is a measurement instrument, and every rule in this session protects the link between concept and datum: one thing at a time, neutral wording, no negations, and questions that respondents are both able and willing to answer. Response formats fail silently, so categories must be exhaustive and mutually exclusive, with an honest opt-out wherever some respondents truly cannot answer. Order, grouping, and layout are part of measurement rather than decoration: context effects are real, and sensitive items need shelter late in the instrument. Because single items are fragile, batteries of Likert items combined into indexes and summated scales carry most serious attitude measurement. Finally, the instrument must be pretested before fieldwork, piloted when possible, and monitored while the field is open — the pilot always finds something. With the survey fielded, the data arrive; session 13 turns to analysing them.