Introduction to Social Research Methodology

Quantitative approaches and measurement

Ben Stanley

Department of Social Sciences, SWPS University

November 24, 2026

Today’s lecture

  • The logic of quantitative research — why turning the social world into numbers buys standardisation, comparison, and generalisation
  • From concept to measure — conceptualisation, dimensions and indicators, and operationalisation: the ladder from an abstract idea to a concrete measure
  • Levels of measurement — nominal, ordinal, interval, and ratio, and why the level you choose constrains everything you can do later
  • Measurement quality — validity and reliability of measures, and the difference between precision and accuracy
  • Survey research and survey modes — the workhorse of quantitative research, and the trade-offs between PAPI, CAWI, and CATI

The logic of quantitative research

Why quantify?

  • Quantitative analysis is the numerical representation and manipulation of observations in order to describe and explain the phenomena those observations reflect (Babbie)
  • Recall session 1: quantitative research is the deductive wing of social research — it specifies expectations in advance and confronts them with evidence
  • Turning observations into numbers buys three things:
    • Standardisation — every case is measured in the same way, with the same instrument
    • Comparison — standardised measures can be compared across people, units, and time
    • Generalisation — with appropriate sampling, patterns in the sample support claims about the population
  • The price is abstraction: a number strips away context — which is why we must be very careful about how the number was made

Running example: Meridian’s board wants to compare engagement and satisfaction across 32 stores and over time. Only standardised measurement makes that comparison meaningful.

The measurement problem

  • In the physical world, measurement is often direct: the distance between two objects can be read off a ruler
  • The things that matter most to organisations — engagement, satisfaction, loyalty, morale, trust — cannot be measured that way
  • You cannot hand an employee a ruler and read off their engagement; you can only observe things that indicate it
  • The problem is not that these things are unreal, but that they are abstract: they exist as ideas summarising many observable behaviours and feelings
  • Quantitative research therefore needs a disciplined procedure for getting from an abstract idea to a defensible number — that procedure is today’s subject

What can we actually observe?

  • It helps to distinguish three kinds of “thing” a researcher can measure (Babbie):
Kind Definition Meridian example
Direct observables Things we can see or count directly Whether an employee attended work; items scanned at a till
Indirect observables Characteristics indicated by traces such as questionnaire answers or records An employee’s answer to “How satisfied are you with your pay?”; tenure recorded in the HR system
Constructs Theoretical creations built by combining observables An “engagement score” combining several questionnaire items
  • Most of what organisational research measures sits in the second and third rows — which is why measurement quality matters so much

From concept to measure

Conceptualisation

  • When I say the word engagement, it evokes a mental image in your mind, just as it evokes one in mine — but they are probably not the same image
  • Each of us holds a private conception; research requires a shared, explicit concept
  • Conceptualisation is the process of coming to an agreement about what a term means for the purposes of a study — its product is a specific, agreed-on meaning
  • If Meridian wants to measure “employee engagement” or “customer satisfaction”, it must first stipulate what, exactly, counts as engagement or satisfaction for this study
  • Without that stipulation, every manager reading the results will silently substitute their own private conception — and the findings will mean different things to different readers

Defining “employee engagement”

  • Try to define employee engagement and you will quickly discover it is not one thing
  • Is an engaged employee one who works hard? One who feels enthusiastic? One who stays? One who recommends the company to friends?
  • A widely used academic definition (Schaufeli and colleagues): a positive, fulfilling, work-related state of mind characterised by vigour, dedication, and absorption
  • A consultant’s definition might instead stress discretionary effort; an HR department’s might stress intention to stay
  • None of these is the “true” definition — concepts have no true definitions, only more and less useful ones
  • The clarification of concepts is continual: examining specific forms of engagement refines your understanding of the general concept, and vice versa

In practice: the working definition you choose determines what your study can and cannot see. Choose it deliberately, and write it down.

Dimensions and indicators

  • Conceptualisation usually reveals that a concept has several dimensions — specifiable aspects or facets of the concept
  • An indicator is an observable sign of the presence or absence (or degree) of the concept — something we can actually record
  • Each dimension typically needs its own indicators; a concept measured through only one dimension is usually measured incompletely

Example — employee engagement:

Dimension Possible indicators
Vigour (energy) Self-rated energy at work; low absence frequency
Dedication (commitment) Agreement with “I am proud of the work I do”; intention to stay
Absorption (immersion) Agreement with “Time flies when I am working”

Concept formation: customer satisfaction

  • Abstract understanding: the degree to which a customer’s experience of a product or service meets or exceeds their expectations
  • Dimensions:
    • Transactional — satisfaction with a specific purchase or visit
    • Cumulative — overall evaluation of the relationship with the company over time
    • Facets often distinguished: product quality, service quality, value for money
  • Variations: satisfaction (experience matched expectations) is not the same as delight (expectations exceeded) or loyalty (continued custom) — related concepts that are frequently confused
  • Deciding which dimensions and variations matter for your research question is a conceptual decision, made before any questionnaire is written

Operationalisation

  • An operational definition specifies precisely how a concept will be measured — the exact operations that produce the number
  • Wishing to study “employee engagement”, Meridian might specify: engagement is measured by responses to six questionnaire items covering vigour, dedication, and absorption, each rated 1–5, averaged into a single score
  • In making this decision, we rule out other possible aspects: overtime worked, participation in company events, tenure, and so forth
  • Operationalisation is therefore an act of deliberate narrowing — it trades richness for measurability, and the trade should be made consciously
  • How faithfully the operational definition captures the concept is precisely the question of validity, to which we return shortly

The measurement ladder

  • The whole journey from idea to data can be summarised in four steps (Babbie):
Measurement step Example: “employee engagement”
Conceptualisation What are the different meanings and dimensions of “engagement”?
Nominal definition For our study, engagement means a positive work-related state of mind with three dimensions: vigour, dedication, absorption
Operational definition We measure it as the mean of six items, two per dimension, each rated 1–5
Measurement in the real world The questionnaire asks: “Rate your agreement: At work, I feel full of energy” (1–5)
  • Each step should follow visibly from the one above — a reader should be able to climb back up the ladder from any survey item to the concept it serves

Getting the categories right

  • Whatever the concept, the attributes composing a variable must satisfy two requirements:
  • Exhaustiveness — every observation must fit one of the categories
    • If contract type is recorded only as “permanent” or “fixed-term”, Meridian’s B2B contractors cannot be classified at all
  • Mutual exclusiveness — every observation must fit only one category
    • “Sales floor” and “e-commerce” must be defined so an employee who does both is classifiable in exactly one way (or the variable must allow multiple roles explicitly)
  • Researchers must also decide the range of variation to capture: if almost nobody at Meridian earns above 20,000 PLN monthly, a top salary category of “20,000 or more” is fine — tracking finer distinctions above it adds nothing

Levels of measurement

Why levels matter

  • Every operationalisation produces attributes at one of four levels of measurement: nominal, ordinal, interval, or ratio
  • The level is not decoration — it determines which mathematical operations are meaningful, and therefore which analyses are legitimate later (session 13)
  • Computing the “average store ID” is arithmetic nonsense; computing average tenure is not — the difference lies entirely in the level of measurement
  • The levels form a hierarchy: each higher level has all the properties of the levels below it, plus one more
  • Getting the level right at the design stage is much cheaper than discovering at the analysis stage that your data cannot answer your question

Nominal measures

  • Variables whose attributes are simply different from one another — categories with no additional structure
  • All we can say about two cases is that they are the same or different
  • Organisational examples: department, contract type, store location, reason for leaving, gender
  • Beware: nominal categories are often coded with numbers (store 1, store 2, …, store 32), but the numbers are labels, not quantities — store 16 is not “twice” store 8
  • Legitimate operations: counting cases per category, identifying the most common category (the mode)

Ordinal measures

  • Variables whose attributes can be rank-ordered: the attributes represent relatively more or less of the variable
  • We can say one case is “more” than another — more satisfied, more senior, more likely to recommend — but the distances between categories have no defined meaning
  • Organisational examples: seniority bands (junior/mid/senior), age brackets, satisfaction ratings from “very dissatisfied” to “very satisfied”, NPS categories (detractor/passive/promoter)
  • The gap between “satisfied” and “very satisfied” need not equal the gap between “neutral” and “satisfied” — the numerals 1–5 conceal this
  • Legitimate operations: everything nominal allows, plus ranking, medians, and percentiles

Interval measures

  • Variables where the distance between attributes is meaningful and constant, but there is no true zero point
  • The classic example: temperature in Celsius — the difference between 0° and 10° equals the difference between 10° and 20°, but 0° does not mean “no temperature”, so 20° is not “twice as hot” as 10°
  • Organisational examples are rare in their pure form: standardised test scores, calendar years (the year a store opened), IQ-style assessment scores
  • Legitimate operations: everything ordinal allows, plus meaningful addition and subtraction — and therefore means and standard deviations

Ratio measures

  • Variables with all the properties of interval measures plus a true zero point — zero means the complete absence of the thing measured
  • Because zero is real, ratios are meaningful: an employee with 24 months’ tenure has twice the tenure of one with 12
  • Organisational examples: tenure in months, salary in PLN, monthly sales, number of absences, number of customer complaints, queue time in seconds
  • Legitimate operations: all of the above, plus multiplication and division — the full toolkit of arithmetic
  • Most “count” and “amount” variables in organisational data are ratio-level — which is one reason administrative records are analytically precious

What each level lets you say

  • Suppose Janina earns 8,000 PLN a month at Meridian and Andrzej earns 4,000 PLN:
Level Operations What we can say
Nominal = ≠ Janina and Andrzej earn different amounts
Ordinal > < Janina earns more than Andrzej
Interval + − Janina earns 4,000 PLN more than Andrzej
Ratio × ÷ Janina earns twice as much as Andrzej
  • Read downwards: each level adds a claim the previous one could not support
  • Measured only at the ordinal level (“above/below median pay”), the last two statements become unavailable — the information is gone and cannot be recovered

Levels constrain analysis

  • You can always convert downwards — collapse exact salary into salary bands — but never upwards: banded data cannot recover exact values
  • Practical rule: measure at the highest level the phenomenon and your resources allow, then simplify later if useful
  • The level determines the statistics available in session 13: means and regression assume interval-level information; nominal data confine you to counts and cross-tabulations
  • The classic ambiguous case: Likert-type ratings (1–5 agreement scales) are strictly ordinal, but researchers routinely treat them as interval — a defensible convention, especially for multi-item averages, but a convention you should apply knowingly, not by accident
  • When a decision hangs on the result, check whether it survives treating the variable strictly as ordinal

Measurement quality

Precision and accuracy

  • Precision concerns the fineness of the distinctions a measure makes: “34 years old” is more precise than “in her thirties”
  • Precise measurements are, as a rule, superior to imprecise ones — but exact precision is not always necessary: if knowing an employee is “in her second year” satisfies the research need, chasing the exact start date is wasted effort
  • Accuracy is a different property: whether the measurement is correct
  • A store’s footfall counter may report visitor numbers to the exact person (precise) while systematically double-counting people who re-enter (inaccurate)
  • Precision is no defence against being wrong — which brings us to the two central qualities of any measure

Reliability

  • Reliability: would the same technique, applied repeatedly to the same object, yield the same result each time? (consistency)
  • Asking employees “How many days were you absent last month?” will be more reliable than “How many days have you been absent in your career?” — recent events are recalled far more consistently
  • Two accessible ways to assess it:
    • Test–retest — administer the same measure to the same people twice; a reliable measure produces similar results (assuming the thing itself has not changed)
    • Internal consistency — where several items measure one concept, check that answers to them hang together; respondents highly engaged on one item should tend to score high on the others
  • One good way to ensure reliability: use measures that have proved their reliability in previous research rather than improvising your own

Validity

  • Validity: does the measure actually measure the concept it claims to measure? (accuracy of the measure)
  • A measure of engagement should measure engagement — not hours worked, not fear of the manager, not enthusiasm for free fruit in the break room
  • Several kinds of evidence bear on it:
Type Question it asks Meridian example
Face validity Does it look like a reasonable measure, on its face? “I feel energetic at work” plausibly indicates engagement
Criterion (predictive) validity Does it relate to an external criterion it should predict? Do low engagement scores predict who actually resigns?
Construct validity Does it relate to other variables as theory expects? Engagement should correlate negatively with burnout scores
Content validity Does it cover the full range of the concept’s meaning? A measure covering vigour only, omitting dedication and absorption, is incomplete

Reliability is not validity

  • Recall the archer’s target from session 1 — it applies to every measure Meridian will ever field:
    • Reliable but not valid — arrows tightly clustered, off the bullseye: hours worked as an “engagement” measure may be beautifully consistent and consistently wrong
    • Valid but not reliable — scattered around the bullseye: a vague, ambiguously worded satisfaction item may capture the right concept erratically
    • Both — tightly clustered on the bullseye: the goal
  • A measure can be perfectly reliable and still measure the wrong thing — consistency is necessary for trust, but never sufficient
  • Both qualities must be established separately, and both before the data are used for decisions — not after

Survey research and survey modes

The workhorse: survey research

  • The survey — standardised questions administered to a sample — is the single most widely used quantitative method, in social science and in organisations alike (Babbie; Creswell & Creswell)
  • Probably the best method available for describing a population too large to observe directly: careful sampling plus a standardised questionnaire yields data in the same form from every respondent
  • Strengths: large samples are feasible; many questions can be asked at once; standardisation permits consistent measurement and reliable comparison across units and over time
  • Weaknesses: coverage of complex topics can be superficial; surveys collect self-reports of action and attitude rather than action itself; the format has a degree of artificiality — people rarely think in terms of “somewhat agree”
  • A survey can establish that Meridian’s satisfaction is falling, and where; why it is falling usually needs the qualitative complement (session 9)

Survey modes: PAPI, CAWI, CATI

  • How the questionnaire reaches the respondent — the survey mode — shapes cost, coverage, and error:
Mode What it is Cost Coverage and error trade-offs
PAPI Paper-and-pencil interviewing: printed questionnaires, self-completed or interviewer-administered Printing, distribution, and manual data entry make it slow and costly at scale Reaches respondents without devices or email; data-entry errors; no automatic routing or checks
CAWI Computer-assisted web interviewing: online questionnaire, increasingly on smartphones Cheapest per respondent; no interviewers, no printing, no data entry Misses those without ready internet access; respondents skew younger and more digital; low response rates; no interviewer to clarify — but also no interviewer to trigger social desirability
CATI Computer-assisted telephone interviewing: interviewer reads the questionnaire by phone, entering answers directly Cheaper than face-to-face; call-centre and interviewer costs remain Interviewer can clarify questions; easy for respondents to hang up on long surveys; declining willingness to answer unknown numbers; more social desirability than CAWI

Choosing a mode at Meridian

  • The mode decision is a coverage decision before it is a cost decision: who can this mode actually reach?
  • CAWI is the obvious choice for head-office and e-commerce staff — but many shop-floor employees may lack company e-mail, so a CAWI-only employee survey would systematically under-represent exactly the stores where turnover is worst
  • PAPI in the stores (completed in the break room, returned in sealed envelopes) closes that coverage gap at the price of printing, collection, and data entry
  • CATI suits the customer side — a short telephone follow-up to recent purchasers — though hang-ups impose brutal discipline on questionnaire length
  • Mixing modes is common and often right, but answers can differ by mode (the same question reads differently on paper and aloud) — a comparability cost to weigh, not ignore

Conclusion

Conclusion

  • Quantitative research earns its power — standardisation, comparison, generalisation — only if the numbers are made well: measurement is where that battle is won or lost
  • The ladder runs from conceptualisation (agree what the term means) through dimensions and indicators to an operational definition precise enough for another researcher to repeat
  • Every measure sits at a level of measurement — nominal, ordinal, interval, ratio — and the level fixes, irrevocably, what analysis can later be done
  • Trust in a measure has two separate components: reliability (consistency) and validity (measuring the right thing) — and neither implies the other
  • The survey is the workhorse that carries these measures into the field, and the mode — PAPI, CAWI, CATI — determines who is reached, at what cost, with what errors
  • Next sessions build directly on this foundation: qualitative approaches (9), sampling (10), and question wording, indexes, and scales (11)
  • Questions and discussion are welcome

Exercise

Today’s exercise: The operationalisation ladder

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-8-exercise

Study guide

Full summary of this session, for revision: Quantitative approaches and measurement

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-8-handout