Secondary data analysis

Introduction to Social Research Methodology

Author
Affiliation

Ben Stanley

Department of Social Sciences, SWPS University

Published

November 17, 2026

Foundations of secondary analysis

Primary data are collected by the researcher, for the researcher’s own question, using an instrument the researcher designed — a survey you commission, interviews you conduct, an experiment you run. Secondary data are data that already exist, collected by someone else and usually for some other purpose: government statistics, a consultancy’s survey, or your own company’s HR records. The distinction concerns origin and purpose, not quality — excellent studies are built on each kind. Its key consequence is that with secondary data, every design decision — sampling, question wording, definitions — was made by someone else, before you arrived.

Secondary analysis is a form of research in which data collected and processed by one researcher are reanalysed, often for a different purpose, by another (Babbie). The analysis is new; the data are not. It counts as a genuine data collection method: “collection” here means locating, obtaining, and preparing existing data rather than generating fresh observations. The practice has grown enormously as governments, organisations, and large academic projects have opened their datasets to reuse. For a manager it is usually the first method to consider, because the answer may already be sitting in a filing cabinet or a public database.

Not the same as a literature review

Secondary analysis must be distinguished from the literature review covered in session 5. A review reuses other researchers’ findings and conclusions; secondary analysis reuses their raw data. The two are different operations with different outputs and different dependencies.

Literature review Secondary analysis
What you reuse Other researchers’ findings and conclusions Other researchers’ raw data
What you do Read, evaluate, synthesise Reanalyse: new questions, new comparisons, new statistics
Output A synthesis of what is known A new empirical result
You depend on The authors’ interpretations The quality of the original data collection

In short: a review tells you what others concluded; secondary analysis lets you check, extend, or contradict those conclusions with the evidence itself.

Why managers should care

Organisational questions rarely arrive with a budget for a bespoke study attached. Before commissioning anything, a competent research team asks what the organisation already has and what already exists outside it. Existing data can also establish context: is the organisation’s problem unusual, or is the whole sector experiencing it? In the Meridian case — turnover up from 14% to 27% in two years — a great deal can be learned from the company’s own records and from Poland’s official statistics before a single employee is surveyed.

Sources of secondary data

Internal organisational sources

Most organisations hold far more data than they ever analyse. The main internal sources are summarised below. All were collected for administration rather than research, which is both their weakness (they record only what the process needed) and their innocence (nobody shaped them to please a researcher).

Internal source What it typically contains What it could show at Meridian
HR records Headcount, contracts, tenure, absence, pay grades, exits by store and month Who leaves: which stores, roles, tenure bands, and managers
Sales and CRM data Transactions, complaints, satisfaction scores, loyalty-card behaviour Whether high-turnover stores also show falling sales
Exit interviews Leavers’ stated reasons for departing Patterns in reasons given — if interviews were conducted consistently
Past employee surveys Engagement or satisfaction results from previous years, often CAWI Whether warning signs preceded the turnover rise — if the same questions were asked each year

Analysing any of these is secondary analysis: the marginal cost is analyst time, not fieldwork.

External sources

External sources fall into three broad groups. First, official statistics: Statistics Poland (GUS), the national statistical office, publishes employment, wage, and labour-demand data by sector and region, much of it through the Local Data Bank; Eurostat, the EU’s statistical office, publishes harmonised labour-market data (including the EU Labour Force Survey) comparable across member states; and the OECD publishes cross-national indicators on employment, job tenure, and earnings. These are collected under statutory mandates, with published methodologies and large, carefully designed samples. Their strength is credibility and comparability; their limit is that they describe sectors and countries, not individual firms.

Second, industry-association and consultancy reports: retail federations, HR consultancies, and job portals publish studies of turnover, pay, and engagement. They are often timely and sector-specific, but the reader should always ask who paid for the study and what the publisher is selling.

Third, large academic surveys, designed for reuse from the outset. The European Social Survey (ESS) measures attitudes and behaviour across Europe every two years; the European Values Study (EVS) is a long-running study of values, including attitudes to work. Academic surveys publish their questionnaires, sampling designs, and data files openly — the gold standard of documentation.

Digital trace and big data

A newer category of secondary data is generated as a by-product of behaviour rather than by asking anyone anything: website clickstreams, till and loyalty-card transactions, staff rota and badge-swipe logs, social media posts, and online job-board activity. Such data are enormous in volume, fine-grained, and collected continuously, recording behaviour as it happens rather than as it is recalled. But they are found data, not designed data: there is no sampling frame, no consent for research, and coverage extends only to those who leave traces. Digital trace data are powerful raw material that demand more methodological caution, not less. A useful one-line summary: big data tells you what people did at scale, but is usually silent on why.

Benefits and risks

The benefits of secondary analysis are chiefly practical; the risks are chiefly methodological. Secondary analysis trades control for convenience.

Benefits Risks
Cost — the expensive phases (sampling, fieldwork, data entry) have already been paid for Fit — the data were collected for someone else’s question; the variable you need may be missing or measured differently
Speed — analysis can begin in days rather than months Unknown quality — response rates, interviewer conduct, and cleaning were outside your control
Scale — sample sizes and coverage no single organisation could fund Unit-of-analysis mismatch — data at the wrong level for your question
Longitudinal reach — series stretching back years or decades permit the study of change Missing documentation — without the instrument and definitions, numbers cannot be judged
No participant burden — nobody is surveyed or observed again; ethically lighter Changing definitions — measures may shift over time, producing artefactual trends

Three of the risks deserve elaboration.

Unit-of-analysis mismatch. The data may describe one level (the sector, the country, the region) while your question concerns another (one firm, one store, one team). Retail turnover for Poland as a whole cannot reveal whether one store’s manager is the problem, and inferring individual-level facts from aggregate data is the road to the ecological fallacy.

Missing documentation. Without the questionnaire, the sampling design, and the definitions, you cannot judge what the numbers mean. An undocumented dataset is a rumour with decimal places.

Changing definitions. Long time series are the great gift of secondary data, and definitions quietly shifting underneath them are the great trap. Statistical offices revise classifications and methods; what counts as “employed”, “vacancy”, or “retail” can change between years. Organisations do the same: if Meridian redefined “leaver” two years ago — say, from voluntary quits only to all departures including seasonal contract expiries — part of the rise from 14% to 27% may be an artefact of the definition. The rule: before interpreting any trend, ask whether the measure changed before believing the world did.

Assessing fitness for purpose

Before using any secondary dataset, interrogate it like a witness. The checklist has seven questions:

Question What to establish
Who collected it? Competence and independence: statistical office, university consortium, industry body, consultancy, vendor
Why was it collected? Purpose shapes design — data collected to administer, to persuade, and to research have different properties
When was it collected? Currency (a 2019 study describes a pre-pandemic labour market) and consistency of measurement across waves
From whom? Population, sampling method, and achieved response rate; beware self-selected panels and hidden composition in aggregates
How? Mode and procedure — PAPI, CAWI, CATI, administrative records, or digital trace — each with characteristic biases
With what instrument? The exact question wording; similar labels can hide different variables
Under what definitions? How key concepts are defined, and whether the definitions have changed — the decisive question for comparability

If the documentation cannot answer these questions, that silence is itself a finding. Purpose deserves particular attention: data collected to administer (payroll, tills) record only what the process needed; data collected to persuade (some consultancy reports) may sample and word questions conveniently; data collected to research (GUS, Eurostat, ESS) are designed for inference — but for someone else’s inference. A quick test of independence: would the collector be embarrassed by any particular result? If so, read the methodology section twice.

Definitions decide whether comparison is possible at all. Does “turnover” mean all separations, voluntary quits, or replacements hired? Does “retail” include e-commerce fulfilment? Does “employee” include agency staff? Only when definitions match can Meridian’s 27% be compared with anyone else’s figure.

The checklist applied

Suppose the team finds an industry report claiming that “retail staff turnover in Poland averages 31%” — apparently reassuring for Meridian’s board. The checklist deflates it: the publisher is a recruitment agency that profits from portraying turnover as universal (who/why); the sample is a self-selected online panel of 400 HR managers with no reported response rate (from whom); and the report’s definition of turnover includes planned seasonal contract expiries, which Meridian’s figure excludes (definitions). The comparison collapses, because the 31% and the 27% are not measuring the same thing. A number without its methodology is not evidence; fitness for purpose is established, never assumed.

Combining secondary and primary research

In organisational diagnosis the efficient sequence is almost always to exhaust the existing data before collecting new data. Secondary analysis does three preparatory jobs: it establishes context (is the problem unusual or sector-wide?), it sharpens the question (internal records show where and among whom the problem sits, so primary research can ask why), and it provides benchmarks against which the organisation’s own results will be judged. It also disciplines spending: never pay to collect what already exists.

Secondary data nonetheless have hard limits. They cannot answer questions that nobody’s data address — no public dataset contains Meridian employees’ reasons for leaving. They cannot rescue a mismatched unit of analysis — sector aggregates will never identify which stores, teams, or managers drive a problem. They cannot reveal meanings and motives — administrative records show that people left, not what leaving meant to them. And they inherit every flaw of the original collection permanently: no reanalysis can repair a biased sample or a leading question. The question is therefore never secondary or primary? but what must the primary study still do, once the secondary work is done?

A defensible combined design for Meridian proceeds in four stages, each cheaper than the next and each making the next one smarter:

  1. Internal secondary analysis — analyse HR records: turnover by store, role, tenure, and manager; check whether the definition of “leaver” changed.
  2. External secondary analysis — benchmark against GUS and Eurostat retail employment and wage data: is Meridian an outlier, or riding a sector-wide wave?
  3. Primary qualitative research — interviews with recent leavers from the worst-affected stores, targeting the why that records cannot reach.
  4. Primary quantitative research — a CAWI survey of current staff, designed around what stages 1–3 revealed.

Conclusion

Secondary analysis reuses data, not findings — that is what separates it from the literature review and what makes it a data collection method in its own right. Its sources span the organisation’s own records (HR files, sales and CRM data, exit interviews, past surveys) and a rich external ecosystem (GUS, Eurostat, OECD, industry reports, and academic surveys such as the ESS and EVS). Its benefits — cost, speed, scale, longitudinal reach, and the absence of participant burden — are bought by surrendering control over design, and all its risks flow from that surrender. The fitness-for-purpose checklist — who, why, when, from whom, how, with what instrument, under what definitions — is how a researcher decides whether someone else’s data can carry their conclusion. Used well, secondary analysis is the opening move of an organisational diagnosis, not the whole game: it tells you where to point the primary research that follows.