First look at the data: answer key

Introduction to Social Research Methodology

Author
Affiliation

Ben Stanley

Department of Social Sciences, SWPS University

Published

January 12, 2027

All values below are computed from meridian-staff-survey.csv (N = 84, no missing values). Small rounding differences (±0.1 of a percentage point, ±0.01 on a mean) are fine; anything further off usually means a filter was left on, a header row was included in a range, or the pivot table is counting a numeric column with SUM instead of COUNT.

Task 1: One variable at a time

1. Frequency table for intends_to_leave:

intends_to_leave n %
No 53 63.1
Yes 31 36.9
Total 84 100.0

2. Frequency table for division:

division n %
Retail 60 71.4
Ecommerce 15 17.9
HQ 9 10.7
Total 84 100.0

Credit students who note in passing that HQ (9 cases) and Ecommerce (15 cases) are small — the point becomes load-bearing in Task 3’s caveats.

3. Mean and median tenure: mean = 33.0 months (33.04); median = 22.0 months. (Range runs from 1 to 180 months.)

4. Which is more informative, and why? The median. The mean sits eleven months above the median, which signals a right-skewed distribution: most staff are relatively recent, while a small number of long-serving employees (up to 15 years) drag the mean upwards. “Half our staff have been here under two years” (median) describes the typical employee; “the average employee has almost three years’ service” (mean) is arithmetically true but misleading about the typical case. Award full marks for any answer that (a) picks the median, (b) explains the mean–median gap as evidence of skew / a long right tail / outliers pulling the mean up. An answer defending the mean can earn partial credit only if it explicitly acknowledges the skew and gives a reason to want the mean anyway (e.g. total staffing-cost arithmetic).

Task 2: Two variables together

1. Crosstab, percentaged within contract type:

Raw counts:

Leave: Yes Leave: No Total
Full-time 18 37 55
Part-time 13 16 29
Total 31 53 84

Row percentages:

Leave: Yes Leave: No Total (N)
Full-time 32.7% 67.3% 100% (55)
Part-time 44.8% 55.2% 100% (29)

Expected sentence: part-timers are the higher flight risk — 44.8% intend to leave against 32.7% of full-timers, a gap of about 12 percentage points. Watch for the classic error of reporting “58% of leavers are full-time” (18/31) — that is the wrong-way percentage rehearsed in the lecture and tested in Task 4; if a pair produces it, ask them what it is a percentage of.

2. Mean engagement of leavers vs stayers:

Group Mean engagement N
Intends to stay (No) 3.17 53
Intends to leave (Yes) 1.90 31

Expected sentence: the gap is about 1.3 points on a 5-point scale — very large; on a scale with only four points of travel, leavers and stayers look like different populations, and engagement is strongly associated with intention to leave. (For reference, overall mean engagement is 2.70, SD ≈ 1.1.)

Task 3: Three findings for the board

Mark against the three requirements — claim + number-with-N + honest caveat — not against matching these examples. At least one finding must state a relationship. Model answers:

  1. “A substantial minority of staff are at risk: 31 of our 84 respondents (37%) say they intend to leave within a year. With only 84 of roughly 450 staff surveyed, the company-wide figure could plausibly be some points higher or lower, but even the low end would be alarming at our current 27% turnover.”
  2. “Part-time staff are markedly more likely to be planning an exit: 44.8% of part-timers intend to leave, against 32.7% of full-timers (N = 29 and 55). The part-time group is the smaller of the two, so this gap comes with more uncertainty, but it points to reviewing part-time terms and rotas first.”
  3. “Intention to leave goes hand in hand with disengagement: staff intending to leave average 1.9 out of 5 on engagement, against 3.2 for those staying (N = 31 and 53). The survey cannot tell us which comes first — disengagement may drive the decision to leave, or a decision already taken may kill engagement — so we should not assume an engagement programme alone will retain these staff.”

Other creditable findings: the tenure story (“half our staff have been here under two years — median 22 months, N = 84 — so we are describing a largely new workforce; long-serving veterans pull the average to 33 months”), or a division comparison — but a division finding must carry the subgroup-size caveat (HQ N = 9: report “1 of 9 HQ respondents intends to leave”, not “11%”). Deduct for: percentages without an N, jargon (“statistically significant”, “correlation coefficient”), causal language without a caveat (“disengagement is driving staff out”), or three findings that are all univariate.

Common weaknesses to name in the debrief: caveats bolted on as ritual (“but the sample is small”) rather than the caveat that actually applies to that claim; and findings that restate the table instead of telling the board something (“32.7% of full-timers intend to leave” is a number, not yet a finding).

Task 4: The suspicious table

1. Row percentages (within each overtime group):

Leave: Yes Leave: No Total (N)
Works regular overtime 16/40 = 40% 24/40 = 60% 100% (40)
No regular overtime 4/20 = 20% 16/20 = 80% 100% (20)

2. Which way did the manager percentage? Within the columns — within the leavers: his 80% is 16/20, the share of intending leavers who work overtime. That percentage answers “what does the leaver group look like?” (composition), not “which group is more likely to leave?” (risk). It is inflated by group size: two-thirds of the pilot sample (40 of 60) work regular overtime, so leavers were always going to be mostly overtime-workers, whatever overtime does.

3. Corrected conclusion (model): “Employees working regular overtime are twice as likely to say they intend to leave — 40% (16 of 40) against 20% (4 of 20) of those without regular overtime. That is a real and sizeable gap, but this table alone cannot show that overtime causes the intention to leave — the busiest stores may differ in other ways (staffing levels, management, pay), and with only 60 respondents in two stores the figures are rough — so eliminating overtime cannot be assumed to ‘virtually eliminate the problem’.”

Award full marks for: correct four percentages; correctly identifying column-percentaging and naming the composition-vs-risk distinction in some wording; and a rewritten claim that (a) compares 40% with 20% (or “twice as likely”), (b) rejects the “overwhelmingly dominant factor / virtually eliminate” overreach, and (c) offers a genuine caveat (causation, third variables, small pilot, two stores only). Note for discussion: the manager’s arithmetic is entirely correct — 16/20 is 80% — which is precisely why this error survives in real reports; the numbers are right and the question they answer is wrong.

Timing and debrief guidance

If pairs finish Task 2 early, stretch them: ask for mean schedule_satisfaction by division (Retail 2.43 vs Ecommerce 3.33 vs HQ 3.44 — the Retail rota problem) or the leave rate by division (Retail 43.3% (26/60), Ecommerce 4 of 15, HQ 1 of 9 — insist on raw counts for the small divisions). In the plenary debrief, collect one Task 3 finding per pair on the board and grade it live against claim + number + caveat; then close on Task 4 by asking who in the room initially believed the manager’s 80% — someone always did — and let that be the lesson.