Quantitative approaches and measurement
Introduction to Social Research Methodology
The logic of quantitative research
Quantitative analysis is the numerical representation and manipulation of observations for the purpose of describing and explaining the phenomena those observations reflect (Babbie). It is the deductive wing of social research introduced in session 1: expectations are specified in advance and then confronted with systematically gathered evidence. Turning observations into numbers buys three things. Standardisation means every case is measured in the same way, with the same instrument. Comparison follows from standardisation: measures taken identically can be compared across people, organisational units, and points in time. Generalisation follows when the sample has been drawn appropriately: patterns found in the sample support careful claims about the wider population. The price of these advantages is abstraction — a number strips away context — which is why the way the number was made deserves close scrutiny. For Meridian, whose board wants to compare engagement and satisfaction across 32 stores and over time, only standardised measurement makes such comparison meaningful.
The central difficulty is that the things organisations care most about cannot be measured the way physical quantities can. The distance between two objects can be read off a ruler; employee engagement, customer satisfaction, loyalty, morale, and trust cannot. These concepts are not unreal, but they are abstract: they exist as ideas summarising many observable behaviours and feelings. Quantitative research therefore needs a disciplined procedure for getting from an abstract idea to a defensible number.
It helps to distinguish three kinds of “thing” a researcher can measure:
| Kind | Definition | Meridian example |
|---|---|---|
| Direct observables | Things we can see or count directly | Whether an employee attended work; items scanned at a till |
| Indirect observables | Characteristics indicated by traces such as questionnaire answers or records | An employee’s answer to “How satisfied are you with your pay?”; tenure recorded in the HR system |
| Constructs | Theoretical creations built by combining observables | An “engagement score” combining several questionnaire items |
Most of what organisational research measures sits in the second and third rows, which is why measurement quality matters so much.
From concept to measure
Conceptualisation
A word such as engagement evokes a mental image — a private conception — in everyone who hears it, but those images differ from person to person. Research requires a shared, explicit concept. Conceptualisation is the process of coming to an agreement about what a term means for the purposes of a study; its product is a specific, agreed-on meaning. If Meridian wants to measure “employee engagement” or “customer satisfaction”, it must first stipulate what exactly counts as engagement or satisfaction for this study. Without that stipulation, every manager reading the results silently substitutes their own private conception, and the findings mean different things to different readers.
Defining employee engagement reveals immediately that it is not one thing. Is an engaged employee one who works hard, one who feels enthusiastic, one who stays, or one who recommends the company to friends? A widely used academic definition, from Schaufeli and colleagues, treats engagement as a positive, fulfilling, work-related state of mind characterised by vigour, dedication, and absorption. A consultant’s definition might instead stress discretionary effort; an HR department’s might stress intention to stay. None of these is the “true” definition, because concepts have no true definitions — only more and less useful ones. The clarification of concepts is continual: examining specific forms of engagement refines the understanding of the general concept, and the refined general understanding changes how each specific form is seen. The working definition chosen determines what the study can and cannot see, so it should be chosen deliberately and written down.
Dimensions and indicators
Conceptualisation usually reveals that a concept has several dimensions — specifiable aspects or facets. An indicator is an observable sign of the presence, absence, or degree of the concept: something that can actually be recorded. Each dimension typically needs its own indicators, and a concept measured through only one dimension is usually measured incompletely.
| Dimension of engagement | Possible indicators |
|---|---|
| Vigour (energy) | Self-rated energy at work; low absence frequency |
| Dedication (commitment) | Agreement with “I am proud of the work I do”; intention to stay |
| Absorption (immersion) | Agreement with “Time flies when I am working” |
Customer satisfaction illustrates the same logic. An abstract understanding might be: the degree to which a customer’s experience of a product or service meets or exceeds their expectations. Its dimensions include the transactional (satisfaction with a specific purchase or visit) and the cumulative (the overall evaluation of the relationship with the company over time), and facets such as product quality, service quality, and value for money are often distinguished. Satisfaction is also easily confused with related but distinct concepts: delight (expectations exceeded) and loyalty (continued custom). Deciding which dimensions and variations matter for the research question is a conceptual decision, made before any questionnaire is written.
Operationalisation
An operational definition specifies precisely how a concept will be measured — the exact operations that produce the number. Wishing to study employee engagement, Meridian might specify that engagement is measured by responses to six questionnaire items covering vigour, dedication, and absorption, each rated 1–5 and averaged into a single score. In making this decision, other possible aspects — overtime worked, participation in company events, tenure — are deliberately ruled out. Operationalisation is an act of conscious narrowing: it trades richness for measurability, and the trade should be made knowingly. How faithfully the operational definition captures the concept is precisely the question of validity, treated below.
The whole journey from idea to data can be summarised in four steps (Babbie):
| Measurement step | Example: “employee engagement” |
|---|---|
| Conceptualisation | What are the different meanings and dimensions of “engagement”? |
| Nominal definition | For our study, engagement means a positive work-related state of mind with three dimensions: vigour, dedication, absorption |
| Operational definition | We measure it as the mean of six items, two per dimension, each rated 1–5 |
| Measurement in the real world | The questionnaire asks: “Rate your agreement: At work, I feel full of energy” (1–5) |
Each step should follow visibly from the one above: a reader should be able to climb back up the ladder from any survey item to the concept it serves.
Getting the categories right
Whatever the concept, the attributes composing a variable must satisfy two requirements. They must be exhaustive: every observation must fit one of the categories. If contract type is recorded only as “permanent” or “fixed-term”, Meridian’s B2B contractors cannot be classified at all. They must also be mutually exclusive: every observation must fit only one category, so “sales floor” and “e-commerce” must be defined in a way that classifies an employee who does both in exactly one way (or the variable must explicitly allow multiple roles). Researchers must also decide the range of variation worth capturing: if almost nobody at Meridian earns above 20,000 PLN monthly, a top salary category of “20,000 or more” is adequate, and tracking finer distinctions above it adds nothing.
Levels of measurement
Every operationalisation produces attributes at one of four levels of measurement, and the level determines which mathematical operations are meaningful — and therefore which analyses are legitimate later. Computing the “average store ID” is arithmetic nonsense; computing average tenure is not, and the difference lies entirely in the level of measurement. The levels form a hierarchy: each higher level has all the properties of those below it, plus one more.
| Level | Defining property | Organisational examples | Legitimate operations |
|---|---|---|---|
| Nominal | Attributes are simply different; same/different is all we can say | Department, contract type, store location, reason for leaving | Counts per category; mode |
| Ordinal | Attributes can be rank-ordered, but distances between them have no defined meaning | Seniority bands, age brackets, satisfaction ratings, NPS categories (detractor/passive/promoter) | All of the above, plus ranking, medians, percentiles |
| Interval | Distances between attributes are meaningful and constant, but there is no true zero | Temperature in Celsius; calendar years (the year a store opened); standardised test scores | All of the above, plus addition and subtraction — hence means and standard deviations |
| Ratio | All interval properties plus a true zero, so ratios are meaningful | Tenure in months, salary in PLN, monthly sales, number of absences, queue time in seconds | All of the above, plus multiplication and division |
Two warnings are worth stating explicitly. First, nominal categories are often coded with numbers — store 1, store 2, and so on to store 32 — but the numbers are labels, not quantities: store 16 is not “twice” store 8. Second, the gap between “satisfied” and “very satisfied” on a rating scale need not equal the gap between “neutral” and “satisfied”; the numerals 1–5 conceal this.
The four levels can be compared through a single example. Suppose Janina earns 8,000 PLN a month at Meridian and Andrzej earns 4,000 PLN:
| Level | Operations | What we can say |
|---|---|---|
| Nominal | = ≠ | Janina and Andrzej earn different amounts |
| Ordinal | > < | Janina earns more than Andrzej |
| Interval | + − | Janina earns 4,000 PLN more than Andrzej |
| Ratio | × ÷ | Janina earns twice as much as Andrzej |
Each level adds a claim the previous one could not support. Had pay been measured only at the ordinal level (“above/below median”), the last two statements would be unavailable — and the information could never be recovered.
That asymmetry yields a practical rule. Data can always be converted downwards (exact salary collapsed into bands) but never upwards (banded data cannot recover exact values), so a variable should be measured at the highest level the phenomenon and the study’s resources allow, and simplified later if useful. The level also determines the statistics available in session 13: means and regression assume interval-level information, while nominal data confine the analyst to counts and cross-tabulations. The classic ambiguous case is the Likert-type rating (a 1–5 agreement scale), which is strictly ordinal but routinely treated as interval — a defensible convention, especially for averages of multiple items, but one that should be applied knowingly. When a decision hangs on a result, it is worth checking whether the result survives treating the variable strictly as ordinal.
Measurement quality
Precision and accuracy
Precision concerns the fineness of the distinctions a measure makes: “34 years old” is more precise than “in her thirties”. Precise measurements are, as a general rule, superior to imprecise ones — there are no conditions under which imprecision is intrinsically better — but exact precision is not always necessary: if knowing that an employee is “in her second year” satisfies the research requirement, effort spent establishing the exact start date is wasted. Accuracy is a different property: whether the measurement is correct. A store’s footfall counter may report visitor numbers to the exact person (precise) while systematically double-counting people who re-enter (inaccurate). Precision is no defence against being wrong.
Reliability
Reliability asks whether a particular technique, applied repeatedly to the same object, yields the same result each time — it is a matter of consistency. Asking employees “How many days were you absent last month?” will be more reliable than “How many days have you been absent in your career?”, because recent events are recalled far more consistently than a lifetime of specific actions. Two accessible ways of assessing reliability are worth knowing. Test–retest reliability is checked by administering the same measure to the same people twice: a reliable measure produces similar results, assuming the underlying thing has not changed in the meantime. Internal consistency applies where several items measure one concept: answers should hang together, so that respondents scoring high on one engagement item tend to score high on the others. One good way to ensure reliable measurement is to use measures that have proved their reliability in previous research rather than improvising new ones.
Validity
Validity asks whether the measure actually measures the concept it claims to measure. A measure of engagement should measure engagement — not hours worked, not fear of the manager. Several kinds of evidence bear on validity:
| Type | Question it asks | Meridian example |
|---|---|---|
| Face validity | Does it look like a reasonable measure of the concept, on its face? | “I feel energetic at work” plausibly indicates engagement |
| Criterion (predictive) validity | Does the measure relate to an external criterion it should predict? | Do low engagement scores predict who actually resigns? |
| Construct validity | Does the measure relate to other variables as theory expects? | Engagement should correlate negatively with burnout scores |
| Content validity | Does the measure cover the full range of meanings within the concept? | A measure covering vigour only, omitting dedication and absorption, is incomplete |
The archer’s target from session 1 applies to every measure an organisation fields. Arrows tightly clustered but off the bullseye are reliable but not valid: hours worked as an “engagement” measure may be beautifully consistent and consistently wrong. Arrows scattered around the bullseye are valid but not reliable: a vague, ambiguously worded satisfaction item may capture the right concept erratically. The goal is arrows tightly clustered on the bullseye. A measure can be perfectly reliable and still measure the wrong thing: consistency is necessary for trust, but never sufficient, and both qualities must be established separately — before the data are used for decisions, not after.
Survey research and survey modes
The survey — standardised questions administered to a sample — is the single most widely used quantitative method, in social science and in organisations alike (Babbie; Creswell & Creswell). It is probably the best method available for describing a population too large to observe directly: careful sampling combined with a standardised questionnaire yields data in the same form from every respondent. Its strengths are that large samples are feasible, that many questions can be asked at once, and that standardisation permits consistent measurement and reliable comparison across units and over time. Its weaknesses are equally characteristic: coverage of complex topics can be superficial; surveys collect self-reports of action and attitude rather than action itself; and the format has a degree of artificiality, since people rarely think in terms of “somewhat agree”. A survey can establish that Meridian’s satisfaction is falling and where; why it is falling usually needs the qualitative complement discussed in session 9.
How the questionnaire reaches the respondent — the survey mode — shapes cost, coverage, and error:
| Mode | What it is | Cost | Coverage and error trade-offs |
|---|---|---|---|
| PAPI | Paper-and-pencil interviewing: printed questionnaires, self-completed or interviewer-administered | Printing, distribution, and manual data entry make it slow and costly at scale | Reaches respondents without devices or email; data-entry errors; no automatic routing or checks |
| CAWI | Computer-assisted web interviewing: an online questionnaire, increasingly completed on smartphones | Cheapest per respondent; no interviewers, printing, or data entry | Misses those without ready internet access; respondents skew younger and more digital; response rates are low; no interviewer to clarify questions — but also no interviewer to trigger social desirability |
| CATI | Computer-assisted telephone interviewing: an interviewer reads the questionnaire by phone and enters answers directly | Cheaper than face-to-face, but call-centre and interviewer costs remain | Interviewer can clarify questions; easy for respondents to hang up on long surveys; declining willingness to answer unknown numbers; more social desirability than CAWI |
The mode decision is a coverage decision before it is a cost decision: the first question is who each mode can actually reach. CAWI is the obvious choice for Meridian’s head-office and e-commerce staff, but many shop-floor employees may lack company e-mail, so a CAWI-only employee survey would systematically under-represent exactly the stores where turnover is worst. PAPI in the stores — completed in the break room and returned in sealed envelopes — closes that coverage gap at the price of printing, collection, and data entry. CATI suits the customer side, as a short telephone follow-up to recent purchasers, though the ease of hanging up imposes brutal discipline on questionnaire length. Mixing modes is common and often the right choice, but answers can differ by mode — the same question reads differently on paper and aloud — and that comparability cost must be weighed rather than ignored.
Conclusion
Quantitative research earns its power — standardisation, comparison, generalisation — only if the numbers are made well, and measurement is where that battle is won or lost. The ladder runs from conceptualisation (agreeing what the term means), through dimensions and indicators, to an operational definition precise enough for another researcher to repeat. Every measure sits at a level of measurement — nominal, ordinal, interval, or ratio — and the level fixes, irrevocably, what analysis can later be done. Trust in a measure has two separate components, reliability (consistency) and validity (measuring the right thing), and neither implies the other. The survey is the workhorse that carries these measures into the field, and the choice of mode — PAPI, CAWI, or CATI — determines who is reached, at what cost, and with what errors. The sessions that follow build directly on this foundation: qualitative approaches (session 9), sampling (session 10), and question wording, indexes, and scales (session 11).