Introduction to Social Research Methodology

Secondary data analysis

Ben Stanley

Department of Social Sciences, SWPS University

November 17, 2026

Today’s lecture

  • Foundations — what secondary data is, and how reusing data differs from collecting your own and from reviewing the literature
  • Sources — the data already inside your organisation, and the data waiting outside it
  • Benefits and risks — why secondary analysis is cheap, fast, and dangerous in equal measure
  • Fitness for purpose — a practical checklist for deciding whether someone else’s data can answer your question
  • Combining approaches — how secondary analysis and primary research work together in a real diagnosis

Foundations of secondary analysis

Primary and secondary data

  • Primary data are collected by the researcher, for the researcher’s own question, using an instrument the researcher designed
    • A survey you commission, interviews you conduct, an experiment you run
  • Secondary data are data that already exist — collected by someone else, usually for some other purpose
    • Government statistics, a rival consultancy’s survey, your own company’s HR records
  • The distinction is about origin and purpose, not about quality: excellent studies are built on each

The key consequence: with secondary data, every design decision — sampling, wording, definitions — was made by someone else, before you arrived.

What is secondary analysis?

  • Secondary analysis is a form of research in which data collected and processed by one researcher are reanalysed — often for a different purpose — by another (Babbie)
  • The analysis is new; the data are not
  • It is a genuine data collection method: “collection” here means locating, obtaining, and preparing existing data rather than generating fresh observations
  • It has grown enormously as governments, organisations, and large academic projects have opened their datasets to reuse
  • For a manager, it is usually the first method to consider: the answer may already be sitting in a filing cabinet or a public database

Not the same as a literature review

  • In the previous session you learned to review the literature; today you learn to reuse the data — and the two are different operations
Literature review Secondary analysis
What you reuse Other researchers’ findings and conclusions Other researchers’ raw data
What you do Read, evaluate, synthesise Reanalyse: new questions, new comparisons, new statistics
Output A synthesis of what is known A new empirical result
You depend on Authors’ interpretations The original data collection’s quality

In short: a review tells you what others concluded; secondary analysis lets you check, extend, or contradict them with the evidence itself.

Why managers should care

  • Organisational questions rarely arrive with a budget for a bespoke study attached
  • Before commissioning anything, a competent junior research team asks: what do we already have, and what already exists out there?
  • Data that already exist can establish context: is our problem unusual, or is the whole sector experiencing it?

Running example: Meridian’s turnover has risen from 14% to 27% in two years. Before surveying a single employee, we can ask what Meridian’s own records — and Poland’s official statistics — already reveal.

Sources of secondary data

Internal sources: the organisation’s own records

  • Most organisations are sitting on more data than they ever analyse
  • HR records — headcount, contracts, tenure, absence, pay grades, exits by store and month
  • Sales and CRM data — transactions, customer complaints, satisfaction scores, loyalty-card behaviour
  • Exit interviews — routinely conducted, rarely systematically analysed
  • Past employee surveys — engagement or satisfaction studies from previous years, often run as CAWI questionnaires
  • These were collected for administration, not research — which is both their weakness and their innocence: nobody shaped them to please a researcher

Internal sources at Meridian

  • What could Meridian’s own records tell the board before any new data are collected?
    • HR records: who leaves — which stores, which roles, which tenure bands, which managers?
    • Exit interviews: what reasons do leavers give — and were the interviews conducted consistently enough to compare?
    • Past engagement surveys: did warning signs appear before the turnover rise — and were the same questions asked each year?
    • Sales and rota data: do high-turnover stores also show falling sales or heavier weekend scheduling?

Notice: every one of these is analysis of existing data. The marginal cost is analyst time, not fieldwork.

External sources: official statistics

  • Statistics Poland (GUS) — the national statistical office: employment, wages, and labour demand by sector and region, much of it in the Local Data Bank
  • Eurostat — the EU’s statistical office: harmonised labour-market data (including the EU Labour Force Survey), comparable across member states
  • OECD — cross-national indicators on employment, job tenure, and earnings across rich democracies
  • Collected under statutory mandates, with published methodologies and large, carefully designed samples
  • Their strength is credibility and comparability; their limit is that they describe sectors and countries, not your firm

External sources: reports and academic surveys

  • Industry-association and consultancy reports — retail federations, HR consultancies, and job portals publish studies of turnover, pay, and engagement
    • Often timely and sector-specific — but ask who paid for them and what they are selling
  • Large academic surveys — designed for reuse from the outset:
    • European Social Survey (ESS) — attitudes and behaviour across Europe, every two years, with full documentation
    • European Values Study (EVS) — long-running study of values, including attitudes to work
  • Academic surveys publish their questionnaires, sampling designs, and data files openly — the gold standard of documentation

Digital trace and big data

  • A newer category: data generated as a by-product of behaviour rather than by asking anyone anything
    • Website clickstreams, till and loyalty-card transactions, staff rota and badge-swipe logs, social media posts, online job-board activity
  • Enormous in volume, fine-grained, and collected continuously — behaviour recorded as it happens, not recalled afterwards
  • But it is found data, not designed data: no sampling frame, no consent for research, and it covers only those who leave traces
  • Treat it as powerful raw material that needs more methodological caution, not less

One line to remember: big data tells you what people did at scale; it is usually silent on why.

Benefits and risks

Benefits: cost, speed, scale

  • Cost — the most expensive phases of research (sampling, fieldwork, data entry) have already been paid for by someone else
  • Speed — analysis can begin in days; a bespoke survey takes months to design, field, and clean
  • Scale — national statistical offices and academic surveys achieve sample sizes and coverage no single organisation could fund
  • For a student, a start-up, or a junior research team, secondary analysis makes otherwise impossible studies feasible

Benefits: reach and burden

  • Longitudinal reach — existing series stretch back years or decades, so you can study change: trends, turning points, before-and-after comparisons
    • No new study, however lavish, can go back in time and measure last decade
  • No participant burden — nobody is interrupted, surveyed, or observed again
    • Ethically lighter: no new consent to seek, no fatigue imposed on over-surveyed employees
    • Practically valuable at Meridian, where staff already distrust head office questionnaires

But none of these benefits answers the decisive question: were these data collected in a way that can answer your question?

Risks: fit and quality

  • Fit to your question — the data were collected for someone else’s purpose; the variable you need may be missing, or measured in a way that misses your concept
    • A national “job satisfaction” item is not a measure of Meridian employees’ satisfaction
  • Unknown quality — you did not oversee the collection: response rates, interviewer conduct, and data cleaning are outside your control and sometimes outside your knowledge
  • Validity and reliability must be assessed at second hand, through documentation — if the documentation exists

Risks: mismatch and missing documentation

  • Unit-of-analysis mismatch — the data describe one level (the sector, the country, the region); your question concerns another (one firm, one store, one team)
    • Retail turnover for Poland as a whole cannot tell you whether store 17’s manager is the problem
    • Inferring individual-level facts from aggregate data is the road to the ecological fallacy
  • Missing documentation — without the questionnaire, the sampling design, and the definitions, you cannot judge what the numbers mean
    • An undocumented dataset is a rumour with decimal places

Risks: definitions change over time

  • Long series are the great gift of secondary data — and definitions quietly shifting underneath them are the great trap
  • Statistical offices revise classifications and methods; what counts as “employed”, “vacancy”, or “retail” can change between years
  • Organisations do it too: if Meridian redefined “leaver” two years ago — say, from voluntary quits only to all departures including seasonal contract expiries — part of the rise from 14% to 27% may be an artefact of the definition
  • Rule: before interpreting any trend, ask whether the measure changed before believing the world did

Benefits and risks: a summary

Benefits Risks
Low cost — collection already paid for Fit — collected for someone else’s question
Speed — analysis can start at once Quality — collection outside your control
Scale — samples you could never fund Unit of analysis — wrong level for your question
Longitudinal reach — study change over time Documentation — may be thin or absent
No participant burden — ethically lighter Definitions — may shift over time

The pattern: the benefits are practical; the risks are methodological. Secondary analysis trades control for convenience — the checklist that follows is how you audit that trade.

Assessing fitness for purpose

The fitness-for-purpose checklist

  • Before using any secondary dataset, interrogate it like a witness:
    1. Who collected it — and are they competent and independent?
    2. Why — for what purpose, and with what interest in the result?
    3. When — is it current enough, and consistent over time?
    4. From whom — what population, what sample, what response rate?
    5. How — by what mode and procedure (PAPI, CAWI, CATI, administrative records, digital trace)?
    6. With what instrument — what exact questions or measures?
    7. Under what definitions — how are the key concepts defined, and have the definitions changed?
  • If the documentation cannot answer these questions, that silence is itself a finding

Who collected it, and why

  • Who: a national statistical office, a university consortium, an industry body, a consultancy, a vendor with a product to sell — each has different competence and different incentives
  • Why matters because purpose shapes design:
    • Data collected to administer (payroll, tills) record what the process needed, no more
    • Data collected to persuade (some consultancy reports) may sample and word questions conveniently
    • Data collected to research (GUS, Eurostat, ESS) are designed for inference — but for someone else’s inference

Quick test: would the collector be embarrassed by any particular result? If yes, read the methodology section twice.

When, and from whom

  • When: a 2019 study of retail employment describes a pre-pandemic labour market; treat vintage as a substantive issue, not a footnote
  • For trends, you need consistent measurement across waves — the same question, the same definition, the same population each time
  • From whom: identify the population, the sampling method, and the achieved response rate
    • A “national survey of employees” recruited from one job portal’s users is a survey of that portal’s users
    • Aggregates hide composition: sector-level wages mix hypermarkets, boutiques, and e-commerce warehouses

How, with what instrument, under what definitions

  • How: mode shapes answers — CAWI panels, CATI interviews, and administrative registers each have characteristic biases (face-to-face and telephone modes, for instance, invite socially desirable answers — more in session 8)
  • With what instrument: get the exact wording; “are you satisfied with your job?” and “would you recommend working here?” are different variables wearing similar labels
  • Under what definitions: the decisive question for comparability
    • Does “turnover” mean all separations, voluntary quits, or replacements hired?
    • Does “retail” include e-commerce fulfilment? Does “employee” include agency staff?
  • Only when definitions match can you compare Meridian’s 27% with anyone else’s number

Applying the checklist at Meridian

  • Suppose we find an industry report: “retail staff turnover in Poland averages 31%” — reassuring for the board?
  • Run the checklist before relaxing:
    • Who/why: a recruitment agency that profits from portraying turnover as universal
    • From whom: an online panel of 400 HR managers, self-selected, response rate unreported
    • Definitions: turnover includes planned seasonal contract expiries — Meridian’s figure excludes them
  • The comparison collapses: the 31% and the 27% are not measuring the same thing

The lesson: a number without its methodology is not evidence; fitness for purpose is established, never assumed.

Combining secondary and primary research

Secondary first, primary second

  • In organisational diagnosis, the efficient sequence is almost always: exhaust the existing data before collecting new data
  • Secondary analysis does three preparatory jobs:
    • Establishes context — is the problem unusual, or sector-wide?
    • Sharpens the question — internal records show where and among whom the problem sits, so primary research can ask why
    • Provides benchmarks — external figures give the yardstick against which your own results will be judged
  • It also disciplines spending: never pay to collect what already exists

What secondary data cannot do

  • It cannot answer questions nobody’s data address — no public dataset contains Meridian employees’ reasons for leaving
  • It cannot rescue a mismatched unit of analysis — sector aggregates will never identify which stores, teams, or managers drive the problem
  • It cannot tell you about meanings and motives — administrative records show that people left, not what leaving meant to them
  • And it inherits every flaw of the original collection, permanently: no reanalysis can repair a biased sample or a leading question

So the question is never secondary or primary? but what must the primary study still do, once the secondary work is done?

A combined design for Meridian

  • A defensible research plan for the board, in sequence:
    1. Internal secondary: analyse HR records — turnover by store, role, tenure, and manager; check whether the definition of “leaver” changed
    2. External secondary: benchmark against GUS/Eurostat retail employment and wage data — is Meridian an outlier, or riding a sector-wide wave?
    3. Primary qualitative: interviews with recent leavers from the worst-affected stores, targeting the why the records cannot reach
    4. Primary quantitative: a CAWI survey of current staff, designed around what stages 1–3 revealed
  • Each stage is cheaper than the next, and each makes the next one smarter

Conclusion

Conclusion

  • Secondary analysis reuses data, not findings — that is what separates it from the literature review, and what makes it a data collection method in its own right
  • The sources span your own organisation’s records and a rich external ecosystem: GUS, Eurostat, OECD, industry reports, and academic surveys such as the ESS and EVS
  • Its benefits — cost, speed, scale, longitudinal reach, no participant burden — are bought by surrendering control over design, and the risks all flow from that surrender
  • The fitness-for-purpose checklist — who, why, when, from whom, how, with what instrument, under what definitions — is how you decide whether someone else’s data can carry your conclusion
  • Used well, secondary analysis is the opening move of a diagnosis, not the whole game: it tells you where to point the primary research that follows
  • Questions and discussion are welcome

Exercise

Today’s exercise: Data scavenger hunt

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-6-exercise

Study guide

Full summary of this session, for revision: Secondary data analysis

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-6-handout