The revision circuit: answer key
Introduction to Social Research Methodology
Station A: Error correction — full list of distortions
Credit any correctly identified distortion with a correct account of how it misleads; the lists below are complete for marking purposes. Strong groups find 4–5 per figure; the truncated axis and the extrapolation are the two no group should miss.
Figure 1 (bar chart) — distortions
- Truncated y-axis (the central error). The axis starts at 3.0, but a bar encodes its value by length, so the lengths no longer represent the values: a 0.3-point difference (3.9 vs 3.6 — about 8% on the measured range) is drawn as a 50% difference in bar height. A busy reader takes away a large divisional gap that the data do not contain.
- The title asserts a conclusion the data do not support. “Dramatically outperform” is editorial. Titles should carry the chart’s message, but the message must be true; here the drama is manufactured by the axis.
- No uncertainty and no sample sizes. With group means from n = 54–118 on a 1–5 scale, differences of 0.1 between adjacent divisions are almost certainly within the margin of error; without error bars or n, the reader cannot tell signal from noise. (Groups who note that 3.9 vs 3.6 might also be within uncertainty are reasoning correctly.)
- The scale is not shown. Nothing says the scores run from 1 to 5. On an unstated scale, 3.6 vs 3.9 cannot be interpreted at all — the reader cannot even tell whether these are good or bad scores.
- (Bonus-level) Ordering invites over-reading. The four means form a smooth staircase (3.9, 3.8, 3.7, 3.6) that visually suggests a strong systematic ranking; with the axis truncated, this staircase looks like a cliff.
The honest version. A bar chart with the y-axis running from 1 (or 0) to 5, all four bars labelled with their values and ns, error bars or a stated margin of error, and a neutral, accurate title such as “Engagement is similar across divisions; Logistics slightly lowest”. The correct impression: all four divisions cluster between 3.6 and 3.9 on a five-point scale — a small difference worth investigating, not a dramatic one. (Also acceptable: a dot plot, or a table, given how few and how close the values are.)
Figure 2 (line chart) — distortions
- Three data points are not a trend. Two months of decline (June→July→August) cannot establish a trajectory; any conclusion about direction is premature.
- Twelve-month linear extrapolation presented as forecast (the central error). The dashed line assumes the current slope continues unchanged for a year — an assumption, not a finding. Extrapolating four times beyond the observed range treats a guess as data.
- The projection is visually continuous with the data. A dashed line at the same slope, on the same chart, reads as “what will happen”; the eye does not separate observation from speculation, and the terminal value (44) anchors the reader.
- Seasonality ignored. June–August are summer months in retail; a summer dip may be an annual pattern, not a decline. Year-on-year comparison (vs last June–August) is the minimum check.
- Truncated and convenient y-axis (40–80). For a line chart zooming is not inherently dishonest, but here the range steepens the visual fall and — combined with the extrapolation — is chosen so the projected collapse fills the frame. The zoom is not acknowledged anywhere.
- Alarmist title. “Free fall” converts a 5-point dip over two months into a catastrophe; the title asserts the conclusion the distorted design manufactures.
The honest version. A line chart of the three observed points only (or, better, the same months plus the previous year for comparison), y-axis range acknowledged, no projection — or a projection clearly separated, labelled as an assumption-laden scenario, with alternatives (e.g. flat and recovering scenarios) shown alongside. Title along the lines of “Satisfaction dipped 5 points over the summer; too early to call a trend”. The correct impression: a modest recent decline that warrants monitoring and further data, not a predicted collapse.
Station B: Executive summary
The faults to elicit (task 1)
- Methodology first, findings buried — the answer arrives in the penultimate sentence.
- Jargon for a lay audience — “operationalised a convergent mixed-methods design”, “Cronbach’s alpha”, “CAWI-administered instrument” tell the board nothing.
- The recommendation is timid and mislabelled — “consideration be given to piloting” hides the action; the summary never says plainly what was found about pay.
- (Also creditable) Passive voice throughout; no statement of what the board should decide.
A model 60-word rewrite
Staff are leaving mainly because of unpredictable schedules and weak supervisor support — not pay. The problem is concentrated among store employees in their first year. Evidence: a survey of 312 employees (69% response) and twelve interviews with recent leavers. Recommendation: pilot fixed scheduling and supervisor training in five high-turnover stores, and re-measure turnover after six months.
(58 words.) Mark against the criteria, not this exact text: findings first; the null pay result stated plainly (it is the most valuable finding — money not wasted on raises); plain language; only the numbers that matter (312, 69%, twelve); the recommendation explicitly labelled; self-contained. Deduct for summaries that open with the method, exceed 60 words, drop the pay finding, or smuggle the recommendation in as if it were a finding.
Station C: Quick-fire answers with one-line justifications
- An experiment (randomised controlled trial). Randomly assign a sample of e-commerce customers to receive the voucher or not, then compare repeat-purchase rates — random assignment rules out confounding, so the difference is attributable to the voucher. (Credit “A/B test” — same design, managerial name.)
- Stratified random sampling (probability sampling). The complete staff list provides a sampling frame, so probability sampling is feasible and licenses generalisation; stratifying by division guarantees each division — including small ones — is properly represented.
- A confidentiality breach. The data were collected under an explicit promise of confidentiality; releasing named responses would break that promise, harm respondents, and destroy trust in future research. The team must refuse and offer only aggregated results in groups large enough to prevent identification.
- Qualitative — in-depth (semi-structured) interviews with recent leavers. The question is about lived experience and meaning (“what it felt like from the inside”), which is precisely what interviews capture and what closed survey items cannot.
- CAWI (computer-assisted web interviewing). E-mail addresses exist for the whole target group, so it is fast and cheap at n ≈ 2,000. Main coverage risk: the frame covers loyalty-programme members only — occasional customers, and those who avoid such programmes, are unrepresented (credit also: low response/self-selection bias).
- Snowball sampling. Participants recruit further participants through their networks — appropriate for hard-to-reach groups. Main limitation: it is non-probability sampling of connected, similar people, so findings cannot be generalised to all leavers (angry leavers who know each other may share one story).
- Poor operationalisation → a validity (and reliability) problem. A single yes/no item cannot capture a multidimensional concept like engagement, forces a crude binary, and invites socially desirable answering. Better: a multi-item index rated on scales (e.g. several statements on energy, commitment, and advocacy averaged into a score) — as in the actual Meridian instrument.
- No — the wording overclaims causality. The evidence is a cross-sectional association: reverse causation (churn depresses engagement of those who remain) and confounding (a weak store manager causes both) are live rivals. Defensible wording: “stores with lower engagement also have higher turnover”. Defensible next step: a longitudinal/panel design tracking stores over time, or an experiment piloting an intervention in randomly selected stores.
Marking note: the one-line reason is the point. In the wrap-up, take answers from different groups per question and press on the reasons; where groups disagree (5 and 8 are the usual candidates), let them argue before adjudicating.
Station D: Model bullet-plans
Full marks for any plan that makes a defensible choice at every required element and gives a reason for each. The plans below are models, not the only acceptable answers.
D1 — diagnosing the store–online satisfaction gap
- Research question: what drives the decline in customer satisfaction in Meridian’s stores, given that online satisfaction is stable?
- Design and approach: mixed methods, explanatory sequential — quantitative first to locate the problem, qualitative second to explain it. Reason: the board needs both the pattern (which stores, which aspects of service) and the mechanism (why), and neither strand alone supplies both.
- Quantitative strand: analyse existing satisfaction data by store, region, and service dimension (secondary data first — cheapest evidence available); then a short CAWI customer survey via loyalty-programme e-mails, with in-store PAPI cards to reach non-members (covers the frame gap).
- Sampling: for the survey, stratified random sample of customers by store format/region; for the qualitative strand, purposive sampling of stores — e.g. three stores with the steepest decline and three that are stable, for contrast.
- Qualitative strand: short interviews or two focus groups with customers, plus interviews with store staff (who often know exactly what has changed — staffing levels, queues, stock).
- Ethical safeguard: informed consent and confidentiality for all participants — especially staff, whose comments about their own stores could identify them to managers; report only aggregated or anonymised material.
- One limitation (any): satisfaction survey respondents are self-selected; loyalty members are not all customers; a cross-sectional design shows the gap but not its direction; store staff accounts may be defensive.
D2 — validity vs reliability
- Validity = accuracy: does the measure capture what it is intended to capture? Example: a “customer satisfaction” score built solely from complaint counts is invalid — complaints measure willingness to complain, not satisfaction.
- Reliability = consistency: does the same procedure yield the same result on repetition? Example: if two HR analysts coding the same exit interviews assign the same reasons for leaving, the coding is reliable; if the engagement survey gives similar scores for the same stable team two weeks apart, the instrument is reliable.
- Reliable but not valid: the archer whose arrows cluster tightly in the wrong corner of the target. A mis-aimed measure can be perfectly consistent — e.g. the complaints metric may be highly stable month to month while consistently measuring the wrong concept. Consistency is not truth.
- Therefore both must be established: reliability is necessary (an inconsistent measure can’t be trusted at all) but not sufficient (a consistent measure can still be wrong); validity presupposes reliability but must be demonstrated separately — e.g. by checking the measure against an established instrument or against behaviour.
- (Credit the archer/target image explicitly — it was the course’s canonical illustration.)
Timing and discussion guidance
- Keep the rotation strict — ten minutes per station with a visible timer. Incomplete answers are fine; the wrap-up fills gaps, and time pressure is part of the exam rehearsal.
- Station A is where groups linger; push them from “the axis is wrong” to how it misleads (length encodes value) and insist on the honest-version descriptions, which is where the learning consolidates.
- Station B: have two or three groups read their summaries aloud in the wrap-up and let the class vote; word-count discipline is half the exercise.
- Station C answers are deliberately one-per-session-cluster: an error pattern (e.g. every sampling question shaky) tells each group precisely which handouts to revisit — say this out loud.
- Close the session by connecting the circuit to the exam: Station C is the closed-question paper, Station D the open-question paper, and Stations A and B are what the course is actually for — being the person in the organisation who cannot be fooled by a chart and cannot be ignored on a page.