Weighing the evidence: answer key

Introduction to Social Research Methodology

Author
Affiliation

Ben Stanley

Department of Social Sciences, SWPS University

Published

November 17, 2026

Task 1: A defensible ranking

There is no single correct ordering, and the point of the task is the justification, not the sequence. The ranking below is defensible; the notes indicate where reasonable groups will diverge. Mark for whether students name the feature that raises or lowers credibility (peer review, disclosed methods, independence, quality control, currency), not for matching this order.

Figure 1 summarises the ranking below as it appears on the model-answer slide; the dashed brackets mark the three places where reasonable groups diverge.

Model answer to Task 1: a defensible ranking of the seven sources A defensible ranking, one among several; credit the feature named, not the sequence. 1, source B, journal article (Kowalczyk & Ferreira): peer-reviewed; a three-wave panel of 1,842; actual quitting, not intention; controls for rival explanations. 2, source F, scholarly book chapter (Novak): editorial review; maps four decades of research; 2019 is not current; others' evidence, not new findings. 3, source G, industry-association report (Retail Workforce Barometer): method disclosed (240 member firms); no independent review; a lobbying interest; over-represents firms willing to respond. 4, source A, consultancy white paper (Hartwell & Grey): current; 4,000 employees in 12 markets; sells its own framework; method undisclosed; the 68% is intention, not behaviour. 5, source E, newspaper article (The Business Courier): journalistic standards; anecdote plus authority: chase the ministry statistics, cite those. 6, source C, practitioner blog post (Lem): a hypothesis worth testing; no quality control; unverifiable personal experience. 7, source D, wikipedia entry (“Employee turnover”): a gateway: definitions, references; anyone can edit; flags missing citations: never cite it. Dashed yellow brackets mark where reasonable groups diverge: places 1 and 2, F first is defensible if argued, for its synthesis value; places 3 and 4, G against A is the classic dispute, and either order is fine if the sales motive is acknowledged; places 5 and 6, C against E, either order with reasoning. One defensible ranking among several: credit the feature named, not the sequence 1 B Journal article Kowalczyk & Ferreira + peer-reviewed; a three-wave panel of 1,842; actual quitting, not intention; controls for rival explanations 2 F Scholarly book chapter Novak + editorial review; maps four decades of research − 2019 is not current; others' evidence, not new findings 3 G Industry-association report Retail Workforce Barometer + method disclosed (240 member firms) − no independent review; a lobbying interest; over-represents firms willing to respond 4 A Consultancy white paper Hartwell & Grey + current; 4,000 employees in 12 markets − sells its own framework; method undisclosed; the 68% is intention, not behaviour 5 E Newspaper article The Business Courier + journalistic standards − anecdote plus authority: chase the ministry statistics, cite those 6 C Practitioner blog post Lem + a hypothesis worth testing − no quality control; unverifiable personal experience 7 D Wikipedia entry “Employee turnover” + a gateway: definitions, references − anyone can edit; flags missing citations: never cite it F first is defensible if argued: its synthesis value G vs A, the classic dispute: either order, if the sales motive is acknowledged C vs E: either order, with reasoning
Figure 1: Model answer to Task 1: a defensible ranking, with the feature that raises or lowers each source’s credibility and the three places where groups may diverge.

1. Source B — peer-reviewed journal article (Kowalczyk & Ferreira)

The only source vetted by independent experts before publication. Beyond peer review itself, the abstract signals methodological strength students should spot: a large sample (1,842), a three-wave panel design tracking the same people over time, measurement of actual quitting rather than intention, and statistical controls for rival explanations (pay, contract type, labour market). This is the anchor source for an academic review.

2. Source F — scholarly book chapter (Novak)

Published by an academic press with editorial review — real quality control, though generally less adversarial than journal peer review. Its distinctive value is synthesis: it maps four decades of research and the competing models, making it an ideal orientation source and a rich mine of references for citation searching. Some groups may rank it first for exactly that reason; that is defensible if argued — but note that a 2019 review is not current, and a review chapter reports others’ evidence rather than new findings.

3. Source G — industry-association report (Retail Workforce Barometer)

Grey literature, but comparatively transparent grey literature: the method (annual survey of 240 member firms) and breakdowns are disclosed, and the descriptive statistics (turnover by format, region, role) are exactly the sectoral context an academic review can legitimately cite as context. Two credibility deductions: no independent review, and a clear lobbying interest — the report ends by calling for lower employment costs, so its framing of the problem is motivated. Membership surveys also over-represent firms willing to respond.

4. Source A — consultancy white paper (Hartwell & Grey)

Also grey literature, ranked below G because the commercial motive is more direct: the paper exists to sell the consultancy’s trademarked retention framework (“deployed with over 40 retail clients”). The headline statistic (68% “considering leaving”) measures loosely defined intention, not behaviour, and the proprietary survey’s sampling and questions are undisclosed. Its strengths are currency and reach (4,000 employees, 12 markets — data academics rarely have). G vs A is the classic dispute: groups who rank A above G on grounds of scale and multi-country coverage are making a reasonable argument, provided they acknowledge the sales motive. Either order earns full credit with that acknowledgement.

5. Source E — newspaper article (The Business Courier)

A quality broadsheet applies journalistic standards — fact-checking, editing, named sources including official statistics and an economist. But it is written for news value on a deadline, not as cumulative knowledge: three executives and one economist is anecdote plus authority, not a study, and the causal claim (wage competition from warehouses) is the journalist’s synthesis, not tested evidence. Useful as a signpost to the labour-ministry statistics it cites — chase the primary source, cite that instead.

6. Source C — practitioner blog post (Lem)

No quality control beyond the author; evidence is fifteen years of personal experience — unsystematic, unverifiable, and possibly shaped by the author’s consulting interests. Yet it is not worthless: the claim that exit interviews mislead is a genuine hypothesis-generator a research team could go on to test, and practitioner insight into mechanisms can be sharp. C vs E is the other legitimate dispute: a group ranking C above E because it at least proposes a testable mechanism, while the newspaper recycles others’ claims, has understood the exercise. Either order with reasoning is acceptable.

7. Source D — Wikipedia entry

Anyone can edit it, content changes without notice, and this entry itself flags that it needs additional citations — an unstable, uncitable source for academic work. Its proper use is as a gateway: the definitions (voluntary/involuntary, functional/dysfunctional turnover) orient a newcomer, and the reference list leads to citable sources. Students should conclude “use it, mine its references, never cite it”.

Managerial-context question

Expect (and credit) answers along these lines: A and G are the sources a board would find most immediately persuasive — current sector benchmarks (G) and multi-market scale plus actionable framing (A); Meridian’s 27% only means something against a sector baseline. E establishes salience and points to the wage-competition threat from e-commerce warehouses — directly relevant to Meridian’s own e-commerce expansion. C suggests a cheap practice (stay interviews) worth piloting even before research concludes. The key insight to reward: credibility for an academic review and usefulness to a decision-maker are different tests — grey and popular sources supply context, currency, and hypotheses, but the review’s arguments must rest on B and F.

Figure 2 sets the two tests side by side: B and F carry the review’s arguments; A and G, then E and C, give the Meridian board its context; E, C and D keep other legitimate uses in the review.

Model answer to the managerial-context question: what each source is good for Credibility for an academic review and usefulness to a decision-maker are different tests. The review's arguments must rest on B and F, the journal article and the book chapter. Context for the Meridian board: A and G are what a board finds most immediately persuasive. G gives current sector benchmarks, since Meridian's 27% only means something against a sector baseline; A gives multi-market scale and an actionable framing; E gives salience and the wage-competition threat from e-commerce warehouses, relevant to Meridian's own e-commerce expansion; C suggests a cheap practice, stay interviews, worth piloting even before research concludes. Other legitimate uses in the review: E is a signpost, so chase the labour-ministry statistics it cites and cite those instead; C is a hypothesis-generator, since a research team could test whether exit interviews mislead employers; D is a gateway, to orient with its definitions and mine its references, but never to cite. Credibility for an academic review and usefulness to a decision-maker are different tests CARRIES THE REVIEW'S ARGUMENTS B and F the journal article and the book chapter: the review's arguments must rest on them CONTEXT FOR THE MERIDIAN BOARD A and G: what a board finds most immediately persuasive G current sector benchmarks: Meridian's 27% only means something against a sector baseline A multi-market scale and an actionable framing E salience, and the wage-competition threat from e-commerce warehouses — relevant to Meridian's own e-commerce expansion C a cheap practice, stay interviews, worth piloting even before research concludes Other legitimate uses in the review E A signpost Chase the labour-ministry statistics it cites, and cite those instead. C A hypothesis-generator Do exit interviews mislead employers? A research team could go on to test it. D A gateway Orient with its definitions and mine its references — never cite it.
Figure 2: Model answer to the managerial-context question: what each source is good for.

Task 2: Example search strings with commentary

Any coherent iteration deserves credit; what matters is that each refinement is motivated by inspection of results and the change is correctly explained. Approximate hit counts will vary — do not mark them against a target. A model log:

Version Search string Approx. hits Commentary
1 employee turnover retail Hundreds of thousands Unusable breadth. Words matched separately, so results include finance papers on asset turnover and anything mentioning retail in passing.
2 "employee turnover" AND retail Tens of thousands Quotation marks force the exact phrase — the single highest-value refinement. Results now on-topic but dominated by descriptive and non-European studies, and still missing work that says “attrition” or “quit”.
3 ("employee turnover" OR "staff attrition" OR "turnover intention") AND retail AND (antecedents OR causes OR "supervisory support") A few thousand OR bundles the field’s synonyms; the third AND block tilts results from describing turnover to explaining it. First page now dominated by studies that would survive abstract triage.

Figure 3 is the same log as it appears on the model-answer slide, with the highest-value refinement — quotation marks — picked out in version 2.

Model answer to Task 2: a search log in three versions A model search log; any coherent iteration earns credit if each refinement is motivated by inspecting the results, and hit counts vary, so do not mark them against a target. Version 1: employee turnover retail; approximately hundreds of thousands hits. Unusable breadth: the words are matched separately, so results include finance papers on asset turnover and anything mentioning retail in passing. Version 2: "employee turnover" AND retail; approximately tens of thousands hits. Quotation marks force the exact phrase: the single highest-value refinement. On-topic now, but dominated by descriptive and non-European studies, and still missing work that says “attrition” or “quit”. Version 3: ("employee turnover" OR "staff attrition" OR "turnover intention") AND retail AND (antecedents OR causes OR "supervisory support"); approximately a few thousand hits. OR bundles the field's synonyms; the third AND block tilts results from describing turnover to explaining it. The first page is now dominated by studies that would survive abstract triage. Any coherent iteration earns credit, if each refinement is motivated by inspecting the results. Hit counts vary: do not mark them against a target. Search string Approx. hits Commentary 1 employee turnover retail Hundreds of thousands Unusable breadth: the words are matched separately, so results include finance papers on asset turnover and anything mentioning retail in passing. 2 "employee turnover" AND retail Tens of thousands Quotation marks force the exact phrase: the single highest-value refinement. On-topic now, but dominated by descriptive and non-European studies, and still missing work that says “attrition” or “quit”. 3 ("employee turnover" OR "staff attrition" OR "turnover intention") AND retail AND (antecedents OR causes OR "supervisory support") A few thousand OR bundles the field's synonyms; the third AND block tilts results from describing turnover to explaining it. The first page is now dominated by studies that would survive abstract triage.
Figure 3: Model answer to Task 2: the search log in three versions.

Further refinements worth crediting: a date filter (e.g. since 2015) for currency; -financial or -asset to remove residual finance hits; restricting a term to titles (intitle:turnover); adding frontline OR "service sector" to sharpen the population.

Common faults to correct: treating hit counts as the goal (the test is first-page quality, not smallness); stacking AND terms until nothing survives; writing OR between concepts instead of within them (turnover OR retail broadens disastrously); refining without saying why — the log’s final column is where the learning is.

Figure 4 gathers the rule (OR within a concept, AND between concepts), the further refinements worth crediting and the common faults to correct.

Model answer to Task 2: refinements to credit and faults to correct At the top, the rule: OR within a concept, AND between concepts. Correct: open bracket, quote employee turnover, OR attrition, close bracket, AND retail, meaning either name, but in retail. Wrong: turnover OR retail, because OR between concepts broadens disastrously. Further refinements worth crediting: a date filter, for example since 2015, for currency; minus financial or minus asset, to remove residual finance hits; intitle colon turnover, to restrict a term to titles; frontline OR quote service sector, to sharpen the population. Common faults to correct: treating hit counts as the goal, when the test is first-page quality, not smallness; stacking AND terms until nothing survives; writing OR between concepts instead of within them; refining without saying why, when the log's final column is where the learning is. OR WITHIN A CONCEPT, AND BETWEEN CONCEPTS ✓ ("employee turnover" OR attrition) AND retail either name, but in retail ✗ turnover OR retail OR between concepts broadens disastrously FURTHER REFINEMENTS WORTH CREDITING A date filter (e.g. since 2015), for currency -financial or -asset to remove residual finance hits intitle:turnover to restrict a term to titles frontline OR "service sector" to sharpen the population COMMON FAULTS TO CORRECT Treating hit counts as the goal: the test is first-page quality, not smallness Stacking AND terms until nothing survives Writing OR between concepts instead of within them Refining without saying why: the log's final column is where the learning is
Figure 4: Model answer to Task 2: refinements to credit and faults to correct.

Timing and discussion guidance

  • If groups stall on Task 1, push them past “it’s biased” to how the bias operates — who paid, who benefits, what was left undisclosed.
  • In the wrap-up, stage the G-vs-A and C-vs-E disputes deliberately: they show credibility is a judgement with criteria, not a lookup table.
  • Close by connecting to the session: the ranking exercise is the hierarchy-of-credibility slide made concrete; the search log is the documented search a systematic review demands — and exactly what their project reports should contain.