Introduction to Social Research Methodology

Quantitative data analysis

Ben Stanley

Department of Social Sciences, SWPS University

January 12, 2027

Today’s lecture

  • From questionnaires to a data matrix — coding, cleaning, and why level of measurement decides what you may do next
  • Describing one variable — frequency tables, percentages, central tendency, and dispersion
  • Describing relationships — crosstabs, group means, and correlation, at the level of intuition
  • Reading, presenting, and trusting numbers — honest tables, a gentle look at inference, and the classic pitfalls
  • By the end, you should be able to take a small dataset and produce findings a board could act on — and spot the standard ways such findings go wrong

From questionnaires to a data matrix

What quantitative data analysis is

  • Quantitative data analysis is the systematic application of statistical and logical techniques to describe, summarise, and compare numerical data (Babbie)
  • Its purpose is to uncover patterns, relationships, and trends that no one can see by leafing through individual questionnaires
  • It sits at stage five of the research process — after design, measurement, sampling, and collection — and inherits every decision made before it
  • Analysis cannot rescue bad measurement or a biased sample; it can only make the most of the data it is given

Meridian: the staff survey has come back — 84 completed questionnaires. The board does not want 84 questionnaires; it wants three sentences it can act on. Analysis is the bridge.

The Meridian staff survey

  • Our running dataset for today: meridian-staff-survey.csv — 84 employees across the company
  • Nine variables per respondent:
    • Who they are — division (Retail / Ecommerce / HQ), contract type, tenure in months, age group
    • What they think — engagement, pay satisfaction, and schedule satisfaction, each on a 1–5 scale
    • What they intend — do they intend to leave in the next year? (Yes / No)
  • The board’s question is simple to state: who is at risk of leaving, and what distinguishes them?
  • Every technique today will be a step towards answering it

The data matrix

  • Quantitative data live in a data matrix: one row per case (here, one employee), one column per variable
  • Each cell holds exactly one value — the case’s score on that variable
  • This rectangle is the universal format: every spreadsheet, statistics package, and survey platform assumes it
  • The matrix is only usable alongside its codebook — the document that records, for every variable, its name, its question wording, and what each value means
  • Without a codebook, a matrix is just numbers; six months later, even its author cannot read it

Meridian: row 7 reads R007, Retail, Part-time, 3, 18-24, 2, 2, 1, Yes — meaningless until the codebook tells you the columns and the scales.

Coding the data

  • Coding assigns a value to every answer so it can be stored and counted
  • Closed questions largely code themselves: the questionnaire’s response options become the values (Yes/No; 1–5 on a Likert item)
  • Decisions still have to be made, and recorded in the codebook:
    • numeric codes or text labels (1/2 vs Full-time/Part-time)
    • how to treat “don’t know” and refused answers
    • how to band continuous answers if bands are wanted (age → age groups)
  • The golden rule: code consistently and document everything — an undocumented coding decision is a future error

Meridian: we coded intention to leave as Yes/No, and age into four bands. Banding loses detail — a choice to note, not hide.

Cleaning the data

  • Real data arrive dirty; cleaning finds and fixes errors before analysis, and it is not optional
  • What to look for:
    • Impossible values (“wild codes”) — an engagement score of 7 on a 1–5 scale; a tenure of 480 months
    • Inconsistencies — an employee aged 18–24 with fifteen years of tenure
    • Duplicates — the same respondent entered twice
    • Missing values — and a policy for handling them, recorded openly
  • The frequency table is your first cleaning tool: run one for every variable and any value outside the legal range shows itself
  • Cleaning decisions are research decisions — log what you changed and why

Level of measurement constrains analysis

  • Recall session 8: how a variable is measured fixes what arithmetic makes sense
Level Example in our data What you may do
Nominal division, intends_to_leave count, mode, percentages
Ordinal age_group, engagement (1–5) all the above + order, median
Interval/ratio tenure_months all the above + mean, standard deviation
  • The scale carries downwards, not upwards: everything allowed for a nominal variable is allowed for an ordinal one, and so on
  • A mean of division is nonsense; a mean of tenure_months is fine; a mean of a 1–5 rating is common practice but strictly a convenience — the intervals are not guaranteed equal
  • First question of any analysis: what level is this variable?

Describing one variable

Univariate description: the first look

  • Univariate analysis describes one variable at a time — always the first step, never skipped
  • It answers the descriptive questions: what values occur, how often, what is typical, how spread out?
  • Three jobs at once:
    • Know your data — you cannot interpret a relationship between variables you have never looked at individually
    • Catch errors — impossible values surface here
    • Report the basics — the board’s first questions are univariate ones
  • The tools: frequency distributions, measures of central tendency, measures of dispersion

Frequency distributions and percentages

  • A frequency distribution lists each value of a variable with the number of cases taking it
  • Convert counts to percentages so distributions can be compared across groups of different sizes — but always report the base N alongside
  • “43% of Retail staff” means something very different when N = 60 than when N = 7
intends_to_leave n %
No 53 63.1
Yes 31 36.9
Total 84 100.0

Meridian: more than a third of surveyed staff intend to leave within a year — already a headline, and we have done nothing but count.

Central tendency: mode, median, mean

  • Three answers to “what is the typical value?” — and level of measurement decides which are available
Measure Definition Needs at least
Mode the most frequent value nominal
Median the middle value when cases are ordered ordinal
Mean the arithmetic average interval/ratio
  • The mode is crude but universal; the median is robust to extreme values; the mean uses all the information but is dragged around by outliers
  • Reporting all three costs nothing and often tells a story by itself

When mean and median disagree

  • If a distribution is symmetric, mean and median coincide; when they diverge, the distribution is skewed
  • Mean well above median → a long right tail: a minority of very large values pulls the average up
  • Classic right-skewed variables: income, firm size, hospital stays — and job tenure

Meridian: tenure has a mean of 33 months but a median of 22 months. Half our staff have been here under two years; a handful of long-serving veterans (up to 15 years) drag the mean up. “The average employee has almost three years’ service” would be true and misleading at once.

  • Rule of thumb: for skewed variables, lead with the median and mention the mean; for symmetric ones, the mean is fine

Dispersion: range and standard deviation

  • Two distributions can share a mean and still be utterly different — central tendency without dispersion is half a description
  • The range (maximum minus minimum) is the simplest measure, but it depends entirely on the two most extreme cases
  • The standard deviation summarises, roughly, how far cases typically sit from the mean
    • small SD → cases cluster tightly around the mean; the mean represents them well
    • large SD → cases are spread widely; the mean represents them poorly
  • No formula needed today — read the SD as a “typical distance from average”

Meridian: engagement averages 2.7 with an SD of about 1.1 — on a 1–5 scale, that is wide. The workforce is not uniformly disengaged; it is divided, which is itself a finding.

Describing relationships

Bivariate description: the question changes

  • Bivariate analysis examines two variables together; the question shifts from what is typical? to do these vary together?
  • This is where analysis meets the research question, which is almost always relational: does engagement differ by contract? does intention to leave depend on satisfaction?
  • The tool depends on the levels of measurement of the pair:
    • two categorical variables → crosstab
    • one categorical, one numeric → compare group means
    • two numeric variables → correlation
  • Association is a pattern in the data; explanation is a claim about the world — keep the two ideas apart (more shortly)

Crosstabs: two categorical variables

  • A crosstab (contingency table) counts cases for every combination of two categorical variables

Meridian: intention to leave by contract type — raw counts:

Leave: Yes Leave: No Total
Full-time 18 37 55
Part-time 13 16 29
Total 31 53 84
  • Raw counts are honest but hard to compare — 18 is more than 13, yet there are nearly twice as many full-timers
  • To compare fairly, we need percentages — and which way we percentage is the single most important decision in reading a crosstab

Percentaging a crosstab the right way

  • Rule: percentage within the categories of the independent variable, then compare across them
  • Here contract type is the (candidate) independent variable → percentage within each contract row:
Leave: Yes Leave: No Total (N)
Full-time 32.7% 67.3% 100% (55)
Part-time 44.8% 55.2% 100% (29)
  • Now the comparison is fair: 44.8% of part-timers intend to leave against 32.7% of full-timers — a 12-point gap
  • Percentaging the other way (“58% of leavers are full-time”) answers a different question — who makes up the leavers — and invites a wrong conclusion, because full-timers dominate simply by being more numerous

Comparing group means

  • When the outcome is numeric (or a rating treated as numeric) and the grouping is categorical, compare the group means
  • The logic mirrors the crosstab: split the cases by the independent variable, summarise the dependent one within each group, compare across

Meridian: mean engagement (1–5) by intention to leave:

Group Mean engagement N
Intends to stay 3.17 53
Intends to leave 1.90 31
  • A gap of 1.3 points on a 5-point scale is large; leavers are not slightly less engaged, they are a different population
  • Always report group sizes with group means — a mean over 5 cases and a mean over 53 do not deserve equal confidence

Correlation, at the level of intuition

  • For two numeric variables, correlation summarises how they move together in one number, r, between −1 and +1
    • positive — high values go with high values (tenure and age)
    • negative — high values go with low values (schedule satisfaction and intention to leave)
    • near zero — no linear pattern
  • The size of r measures strength: ±0.1 is weak, ±0.3 moderate, ±0.5 strong (rough social-science conventions)
  • Correlation captures only straight-line association, and a single number can hide very different patterns — never report an r for data you have not looked at

Meridian: engagement and pay satisfaction correlate positively — dissatisfaction clusters in the same people. Useful to know; nothing yet about which causes which.

Reading, presenting, and trusting numbers

Reading and presenting tables

  • A table is an argument; a good one can be read without the surrounding text. Every honest table has:
    • a title that states what, who, and when (“Intention to leave by contract type, Meridian staff survey, N = 84”)
    • labelled rows and columns, with units and scale ranges spelt out
    • the base N for every percentage
    • a source note, and a note of who was excluded
  • Round sensibly — one decimal place is almost always enough; false precision (“32.727%”) signals naivety, not rigour
  • State which way a crosstab is percentaged; make the 100% totals visible
  • Reading others’ tables, run the same list as a checklist: no N, no labels, no source — no trust

How sure can we be? Sampling variability

  • Our 84 respondents are a sample of Meridian’s ~450 staff; a different 84 would have produced slightly different numbers
  • This is sampling variability: sample results dance around the population value from sample to sample
  • Small samples dance more; large samples dance less — which is why base N matters so much
  • Statistical inference asks: is the pattern in our sample strong enough that it is unlikely to be a fluke of who happened to respond?
  • We need the idea, not the machinery — no formulas today

Meridian: the 12-point gap between part-timers and full-timers could partly reflect who ended up in our 84. How seriously to take it is exactly the inferential question.

Significance and confidence, in plain words

  • A confidence interval turns one number into an honest range: “36.9% intend to leave, give or take about 10 points” — the plausible values, given a sample this size
  • A statistically significant difference is one too large to be plausibly explained by sampling variability alone — “probably not a fluke”
  • Two warnings that save careers:
    • significance is not importance — with a huge sample, trivial differences become “significant”; with a tiny one, real differences fail to
    • the calculations assume a decent (probability) sample — no arithmetic repairs a biased one (recall session 8 on sampling)
  • For this course: read significance claims as “how sure can we be?”, and always ask for the interval, not just the verdict

Pitfall 1: correlation is not causation

  • An association between X and Y is compatible with at least four stories:
    • X causes Y
    • Y causes X (reverse causation)
    • Z causes both (a lurking third variable)
    • chance
  • Description tells you that variables travel together; it cannot tell you why
  • Establishing cause needs design — experiments, time order, controls (sessions 5 and 8) — not just a strong correlation

Meridian: low engagement predicts intending to leave. Does disengagement drive exit — or have staff who already decided to leave stopped caring? The crosstab cannot say. Act on the association, but do not oversell the mechanism.

Pitfall 2: percentaging the wrong way

  • The most common table error in real reports: percentaging within the dependent variable and reading the result as risk
  • “58% of employees who intend to leave are full-time” → true, and almost inevitable, because 65% of the workforce is full-time
  • The wrong-way percentage reflects group size, not group risk; the right-way percentage (32.7% vs 44.8%) reverses the apparent conclusion
  • Defence: before reading any crosstab, find the 100% totals and ask what exactly is this a percentage of?
  • If a claim compares groups, the percentages must be computed within those groups

Pitfall 3: tiny subgroups

  • Percentages hide their base: “33% of HQ respondents under 25 intend to leave” may be one person out of three
  • Small subgroups make wild percentages — one respondent changing their mind swings the figure by 33 points
  • Rules of thumb:
    • always print the N beside any subgroup percentage
    • be sceptical of any percentage on fewer than ~30 cases; below ~10, report the raw count instead (“1 of 3”)
    • resist the urge to keep splitting until something dramatic appears — with enough slicing, noise always obliges

Meridian: only 9 respondents work at HQ. “11% of HQ staff intend to leave” is one person — say “1 of 9”, not “11%”.

Tools of the trade

  • For this course, a spreadsheet is enough: Excel or Google Sheets
    • COUNTIF / AVERAGE / MEDIAN for univariate work
    • pivot tables for frequency tables, crosstabs, and group means — the workhorse of today’s exercise
  • One step further sit the statistics packages: jamovi (free, friendly), SPSS (the corporate and academic standard), R and Python (free, powerful, code-based)
  • They automate the same logic you learn in a spreadsheet — significance tests, intervals, and models on top of the descriptions
  • The tool is not the skill: knowing which table to build and how to read it transfers across every package

Conclusion

Conclusion

  • Analysis starts before statistics: a coded, cleaned, documented data matrix, with the level of measurement of every variable known
  • Univariate first — frequencies, percentages with their N, mode/median/mean chosen to fit the level, dispersion alongside
  • Bivariate next — crosstabs percentaged within the independent variable, group means with group sizes, correlation read for direction and strength
  • Inference, gently: sample numbers wobble; significance and confidence intervals measure how sure we can be, not how much it matters
  • The three classic sins — causal claims from correlations, wrong-way percentages, tiny-base drama — are all avoidable by asking what is this a percentage of, and how many cases is it based on?
  • In today’s exercise you will do all of this yourselves, on the Meridian survey, with a pivot table
  • Questions and discussion are welcome

Exercise

Today’s exercise: First look at the data

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-13-exercise

Study guide

Full summary of this session, for revision: Quantitative data analysis

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-13-handout