Quantitative data analysis
Introduction to Social Research Methodology
From questionnaires to a data matrix
Quantitative data analysis is the systematic application of statistical and logical techniques to describe, summarise, and compare numerical data (Babbie). Its purpose is to uncover patterns, relationships, and trends that no one could see by leafing through individual questionnaires. Analysis sits at stage five of the research process — after design, measurement, sampling, and data collection — and it inherits every decision made before it. This has an important corollary: analysis cannot rescue bad measurement or a biased sample. It can only make the most of the data it is given.
Our running dataset throughout this session is the Meridian staff survey: 84 completed questionnaires from employees across the company, recording each respondent’s division (Retail, Ecommerce, or HQ), contract type, tenure in months, age group, three attitude ratings on 1–5 scales (engagement, pay satisfaction, schedule satisfaction), and whether they intend to leave within a year. The board’s question is simple to state — who is at risk of leaving, and what distinguishes them? — and every technique below is a step towards answering it.
The data matrix and the codebook
Quantitative data live in a data matrix: one row per case (here, one employee) and one column per variable, with each cell holding exactly one value — that case’s score on that variable. This rectangle is the universal format assumed by every spreadsheet, statistics package, and survey platform. The matrix is only usable alongside its codebook: the document that records, for every variable, its name, the question wording it came from, and what each value means. Without a codebook a matrix is just numbers; six months later, even its author cannot read it.
Coding
Coding assigns a value to every answer so that it can be stored and counted. Closed questions largely code themselves — the questionnaire’s response options become the values (Yes/No; the points 1–5 of a Likert item) — but decisions still have to be made and recorded: whether to use numeric codes or text labels, how to treat “don’t know” and refused answers, and how to band continuous answers when bands are wanted (age into age groups, for instance — a choice that loses detail and should be noted, not hidden). The golden rule is to code consistently and document everything; an undocumented coding decision is a future error.
Cleaning
Real data arrive dirty, and cleaning — finding and fixing errors before analysis — is not optional. The things to look for are:
| Problem | Example |
|---|---|
| Impossible values (“wild codes”) | an engagement score of 7 on a 1–5 scale; a tenure of 480 months |
| Inconsistencies | an employee aged 18–24 with fifteen years of tenure |
| Duplicates | the same respondent entered twice |
| Missing values | gaps that need an explicit, openly recorded handling policy |
The frequency table is the first cleaning tool: run one for every variable, and any value outside the legal range shows itself immediately. Cleaning decisions are research decisions — log what you changed and why.
Level of measurement constrains analysis
Recall session 8: how a variable is measured fixes what arithmetic makes sense.
| Level | Example in the Meridian data | What you may legitimately do |
|---|---|---|
| Nominal | division, intends_to_leave | count, mode, percentages |
| Ordinal | age_group, engagement (1–5) | all of the above, plus ordering and the median |
| Interval/ratio | tenure_months | all of the above, plus the mean and standard deviation |
The permissions carry downwards, not upwards: everything allowed for a nominal variable is allowed for an ordinal one, and so on. A mean of division is nonsense; a mean of tenure_months is fine; a mean of a 1–5 rating is common practice but strictly a convenience, since the intervals between scale points are not guaranteed to be equal. The first question of any analysis is therefore: what level is this variable?
Describing one variable
Univariate analysis describes one variable at a time. It is always the first step and is never skipped, because it does three jobs at once: it lets you know your data (you cannot interpret a relationship between variables you have never looked at individually), it catches errors, and it answers the basic descriptive questions a board asks first.
Frequency distributions and percentages
A frequency distribution lists each value of a variable together with the number of cases taking that value. Counts are converted to percentages so that distributions can be compared across groups of different sizes — but the base N must always be reported alongside, because “43% of Retail staff” means something very different when N = 60 than when N = 7. In the Meridian survey, the frequency table for intention to leave already yields a headline before any further analysis: 31 of 84 respondents (36.9%) intend to leave within a year.
Central tendency
There are three answers to the question “what is the typical value?”, and the level of measurement decides which are available.
| Measure | Definition | Requires at least | Character |
|---|---|---|---|
| Mode | the most frequent value | nominal | crude but universal |
| Median | the middle value when cases are ordered | ordinal | robust to extreme values |
| Mean | the arithmetic average | interval/ratio | uses all the information, but dragged around by outliers |
Reporting all three costs nothing and often tells a story by itself. When mean and median coincide, the distribution is roughly symmetric; when they diverge, it is skewed. A mean well above the median indicates a long right tail: a minority of very large values pulls the average up. Income, firm size, and job tenure are classic right-skewed variables. In the Meridian data, tenure has a mean of 33 months but a median of 22: half the staff have been with the company under two years, while a handful of long-serving veterans (up to fifteen years) drag the mean upwards. “The average employee has almost three years’ service” would be simultaneously true and misleading. The rule of thumb: for skewed variables, lead with the median and mention the mean; for symmetric ones, the mean is fine.
Dispersion
Two distributions can share a mean and still be utterly different, so central tendency without dispersion is only half a description. The range (maximum minus minimum) is the simplest measure but depends entirely on the two most extreme cases. The standard deviation summarises, roughly, how far cases typically sit from the mean: a small SD means cases cluster tightly around the mean, which therefore represents them well; a large SD means they are spread widely, and the mean represents them poorly. No formula is needed at this stage — read the SD as a “typical distance from average”. In the Meridian data, engagement averages 2.7 with a standard deviation of about 1.1 — wide, for a 1–5 scale. The workforce is not uniformly disengaged; it is divided, which is itself a finding.
Describing relationships
Bivariate analysis examines two variables together, shifting the question from what is typical? to do these vary together? This is where analysis meets the research question, which is almost always relational. The appropriate tool depends on the levels of measurement of the pair:
| Pair of variables | Tool |
|---|---|
| two categorical | crosstab (contingency table) |
| one categorical, one numeric | comparison of group means |
| two numeric | correlation |
Crosstabs and the percentaging rule
A crosstab counts cases for every combination of two categorical variables. In the Meridian survey, intention to leave by contract type gives raw counts of 18 intending leavers among 55 full-timers and 13 among 29 part-timers. Raw counts are honest but hard to compare — 18 is more than 13, yet there are nearly twice as many full-timers — so we percentage. The rule: percentage within the categories of the independent variable, then compare across them. Contract type is the candidate independent variable here, so we percentage within each contract group:
| Leave: Yes | Leave: No | Total (N) | |
|---|---|---|---|
| Full-time | 32.7% | 67.3% | 100% (55) |
| Part-time | 44.8% | 55.2% | 100% (29) |
Now the comparison is fair: 44.8% of part-timers intend to leave against 32.7% of full-timers — a twelve-point gap. Percentaging the other way (“58% of leavers are full-time”) answers a different question — who makes up the leavers — and invites a wrong conclusion, because full-timers dominate that figure simply by being more numerous.
Comparing group means
When the outcome is numeric (or a rating treated as numeric) and the grouping variable is categorical, we compare group means: split the cases by the independent variable, summarise the dependent variable within each group, and compare across groups. In the Meridian data, mean engagement is 3.17 among the 53 respondents intending to stay and 1.90 among the 31 intending to leave — a gap of 1.3 points on a five-point scale, which is large. Leavers are not slightly less engaged; they look like a different population. Group sizes must always be reported with group means: a mean over 5 cases and a mean over 53 do not deserve equal confidence.
Correlation
For two numeric variables, correlation summarises how they move together in a single number, r, ranging from −1 to +1. A positive correlation means high values of one variable go with high values of the other (tenure and age); a negative correlation means high values go with low values (schedule satisfaction and intention to leave); a value near zero means no linear pattern. The size of r measures strength, with rough social-science conventions of ±0.1 for weak, ±0.3 for moderate, and ±0.5 for strong associations. Two cautions apply. Correlation captures only straight-line association, and a single number can conceal very different patterns — never report an r for data you have not looked at. And correlation says nothing about direction of influence: in the Meridian data, engagement and pay satisfaction correlate positively, which tells us dissatisfaction clusters in the same people, but nothing about which causes which.
Reading, presenting, and trusting numbers
Honest tables
A table is an argument, and a good one can be read without the surrounding text. Every honest table has: a title stating what, who, and when (“Intention to leave by contract type, Meridian staff survey, N = 84”); labelled rows and columns with units and scale ranges spelt out; the base N for every percentage; and a source note, including who was excluded. Round sensibly — one decimal place is almost always enough, and false precision (“32.727%”) signals naivety rather than rigour. State which way a crosstab is percentaged and make the 100% totals visible. When reading other people’s tables, run the same list as a checklist: no N, no labels, no source — no trust.
Statistical inference, gently
The 84 respondents are a sample of Meridian’s roughly 450 staff, and a different 84 would have produced slightly different numbers. This is sampling variability: sample results dance around the population value from sample to sample, and small samples dance more than large ones — which is one more reason the base N matters. Statistical inference asks whether a pattern in the sample is strong enough that it is unlikely to be a fluke of who happened to respond.
Two ideas carry most of the weight, and no formulas are needed for either. A confidence interval turns a single number into an honest range — “36.9% intend to leave, give or take about ten points” — the set of population values that are plausible given a sample of this size. A statistically significant difference is one too large to be plausibly explained by sampling variability alone: in plain words, “probably not a fluke”. Two warnings accompany these ideas. First, significance is not importance: with a very large sample, trivially small differences become “significant”, while with a very small one, real differences fail to. Second, the calculations assume a decent probability sample — no arithmetic repairs a biased one (recall session 8). For this course, read significance claims as answers to “how sure can we be?”, and always ask for the interval rather than just the verdict.
Three classic pitfalls
| Pitfall | What goes wrong | Defence |
|---|---|---|
| Correlation read as causation | an association between X and Y is compatible with X causing Y, Y causing X (reverse causation), a third variable Z causing both, or chance | causal claims need design — experiments, time order, controls — not just a strong association |
| Percentaging the wrong way | percentages computed within the dependent variable are read as group risk; “58% of leavers are full-time” mostly reflects the fact that 65% of the workforce is full-time | find the 100% totals and ask what exactly is this a percentage of?; risk comparisons require percentages computed within the groups being compared |
| Tiny subgroups | percentages hide their base — “33% of HQ staff under 25” may be one person in three, and one changed mind swings the figure by 33 points | print the N beside every subgroup percentage; be sceptical below ~30 cases; below ~10, report raw counts (“1 of 3”) instead |
The causation pitfall applies directly to Meridian: low engagement predicts intending to leave, but the crosstab cannot say whether disengagement drives exit or whether staff who have already decided to leave have stopped caring. The association is worth acting on; the mechanism should not be oversold. The tiny-subgroup pitfall also applies: only nine respondents work at HQ, so “11% of HQ staff intend to leave” is one person and should be reported as “1 of 9”. And beware the temptation to keep splitting the data until something dramatic appears — with enough slicing, noise always obliges.
Tools of the trade
For this course a spreadsheet is enough: Excel or Google Sheets, using COUNTIF, AVERAGE, and MEDIAN for univariate work and pivot tables for frequency tables, crosstabs, and group means. One step further sit the statistics packages: jamovi (free and friendly), SPSS (the long-standing corporate and academic standard), and R and Python (free, powerful, code-based). They automate the same logic learned in a spreadsheet, adding significance tests, confidence intervals, and statistical models on top of the descriptions. The tool is not the skill: knowing which table to build and how to read it transfers across every package.
Conclusion
Analysis starts before statistics, with a coded, cleaned, documented data matrix and a known level of measurement for every variable. Univariate description comes first — frequencies, percentages with their base N, measures of central tendency chosen to fit the level, and dispersion alongside. Bivariate description follows — crosstabs percentaged within the independent variable, group means reported with group sizes, and correlation read for direction and strength. Inference, taken gently, reminds us that sample numbers wobble: significance and confidence intervals measure how sure we can be, not how much a difference matters. The three classic sins — causal claims from correlations, wrong-way percentages, and tiny-base drama — are all avoidable by asking two questions of every number: what is this a percentage of, and how many cases is it based on?