Foundations of social research
What is social research?
What makes social research different from everyday sense-making
At the top, the definition: social research is the systematic investigation of human behaviours, attitudes, relationships and social systems, to produce reliable knowledge and insight. Below, a comparison in four rows. Procedure: everyday sense-making is casual, with its steps left unstated; social research is explicit, transparent and repeatable. Questions framed by: personal experience, against theory, meaning existing concepts and explanations. Claims settled by: intuition, anecdote and authority, against systematically gathered evidence. Aim: a personal impression, against describing, then explaining, then informing change. At the bottom: organisations run on it too; employee surveys, customer research and market analysis are social research applied to managerial decisions, and the same standards apply.
Social research is the systematic investigation of human behaviours, attitudes,
relationships and social systems, to produce reliable knowledge and insight
Everyday sense-making
Procedure
Questions framed by
Claims settled by
Aim
casual; steps left unstated
personal experience
intuition · anecdote · authority
a personal impression
Social research
explicit · transparent · repeatable
theory: concepts and explanations
systematically gathered evidence
describe → explain → inform change
Organisations run on it too — employee surveys, customer research and market analysis
are social research applied to managerial decisions, and the same standards apply
Objectives of social research
The objectives of social research as a ladder of increasing ambition
Four steps rising from left to right. Step 1, describe: what is happening, to whom, and under what conditions. Step 2, explain: causes and consequences, moving from what to why. Step 3, predict: use established patterns to anticipate what comes next. Step 4, inform policy: the evidence on which governments, organisations and communities act. An arrow along the bottom reads: increasing ambition; each step builds on the one below.
increasing ambition: each step builds on the one below
1
Describe
what is happening, to whom,
under what conditions
2
Explain
causes and consequences:
from what to why
3
Predict
use established patterns
to anticipate what
comes next
4
Inform policy
the evidence on which
governments, organisations
and communities act
Types of social research: two axes
Types of social research: two axes, four combinations
A two-by-two grid. The columns are the first axis, the kind of data and reasoning: qualitative (deep, interpretative, concerned with meaning and experience) and quantitative (numerical, statistical, concerned with measurement and pattern). The rows are the second axis, whether the researcher intervenes: experimental (controls and varies conditions) and descriptive (observes only, without intervening). Every cell is a possible study. Qualitative and experimental: a new civic-education workshop is tried in some schools, and focus groups explore how pupils experienced it. Quantitative and experimental: half of voters, chosen at random, get a text-message reminder, and turnout in the two groups is compared. Qualitative and descriptive: in-depth interviews explore how first-time voters experienced an election. Quantitative and descriptive: a large survey maps how turnout varies with age, without intervening.
Axis 1: what kind of data and reasoning?
Qualitative
deep, interpretative: meaning and experience
Quantitative
numerical, statistical: measurement and pattern
Axis 2: does the researcher intervene?
Experimental
intervenes:
controls and
varies conditions
Descriptive
observes only:
no intervention
A new civic-education workshop
is tried in some schools; focus groups
explore how pupils experienced it
Half of voters, chosen at random,
get a text-message reminder;
turnout in the two groups is compared
In-depth interviews explore how
first-time voters experienced
an election
A large survey maps how
turnout varies with age,
without intervening
Inductive and deductive logic
Deductive and inductive logic as two halves of one cycle
Theory, the general, sits at the top; observations, the specific, at the bottom. On the left, the deductive path runs down: from theory to a hypothesis, then to data gathered to test it; it moves from the general to the specific and is typical of quantitative work. On the right, the inductive path runs up: from observations to patterns, then to theory built up from them; it moves from the specific to the general and is typical of qualitative work. In the middle: most research cycles between the two, and today's inductive insight becomes tomorrow's deductive test.
Theory
the general
Observations
the specific
Hypothesis
DEDUCTIVE
general → specific
start from theory,
derive a hypothesis,
gather data to test it
typical of
quantitative work
Patterns
INDUCTIVE
specific → general
start from observations,
look for patterns,
build theory up from them
typical of
qualitative work
most research cycles
between the two:
today's inductive insight
becomes tomorrow's
deductive test
The stages of a study
The stages of a study as a chain of seven links
Seven interlocking chain links, numbered in order: 1, choose a topic (what is worth studying, and can it be studied?); 2, design the study (turn the topic into answerable questions); 3, sample (decide who or what to study); 4, collect data (gather the evidence); 5, analyse and interpret (make sense of the evidence); 6, present (communicate what was found); 7, reflect (acknowledge the limits). Each decision constrains the ones that follow, so the chain is only as strong as its weakest link. An orange thread runs through every link: the running example followed through all the stages, why do young people vote less than older people?
1
Choose a topic
what is worth studying,
and can it be studied?
2
Design the study
turn the topic into
answerable questions
3
Sample
decide who or what
to study
4
Collect data
gather the
evidence
5
Analyse and interpret
make sense
of the evidence
6
Present
communicate
what was found
7
Reflect
acknowledge
the limits
RUNNING EXAMPLE, FOLLOWED THROUGH EVERY STAGE
Why do young people vote less than older people?
Choosing a research topic
A good research topic meets three criteria at once
Three overlapping circles; a good topic sits where all three overlap. Relevant: it speaks to a current issue or a gap in knowledge. Feasible: it is achievable with the time, resources, expertise and access available. Ethical: its benefits outweigh any risk of harm. The running example, why do young people vote less, passes all three: turnout among the young is falling; survey data on turnout and age already exist; and asking people about voting is low-risk.
Relevant
speaks to a current issue
or a gap in knowledge
Feasible
achievable with the time,
resources, expertise
and access available
Ethical
its benefits outweigh any risk of harm
a good
topic
RUNNING EXAMPLE
Why do young people vote less?
✓ turnout among the young
is falling
✓ survey data on turnout
and age already exist
✓ asking people about voting is low-risk
Research design and planning
Research design as the blueprint between a topic and the evidence
On the left, the topic: what to study. On the right, the evidence needed to answer it. Between them, the design, drawn as a blueprint with three tasks: 1, research questions or hypotheses, clear and realistically answerable; 2, scope, meaning what is studied, who is included, when and where; 3, methods and tools, qualitative, quantitative or a mixture, chosen to fit the questions. Running example: the topic "why do young people vote less?" is sharpened into the testable question "does political interest explain the age gap in turnout?"
Topic
what to study
Design: the blueprint
Evidence
needed to answer it
1
Research questions or hypotheses
clear, and realistically answerable
2
Scope
what is studied, who is included, when, where
3
Methods and tools
qualitative, quantitative or a mixture
RUNNING EXAMPLE
Why do young people vote less?
a topic
Does political interest explain the age gap in turnout?
a testable question: the first task of the design
Variables and operationalisation
Variables, their roles, and operationalisation
Top: a variable is anything that varies across cases, such as age, income, turnout or political interest. An arrow runs from the independent variable, the presumed cause, to the dependent variable, the presumed effect. In the running example, the independent variable is political interest, measured as a 0 to 10 self-rating, and the dependent variable is turnout in the last national election. Bottom: operationalisation turns an abstract concept into something measurable. The concept political engagement, which cannot be observed directly, is linked to three observable indicators: did you vote; how often do you discuss politics; are you a party member. A bracket around the indicators asks the question of validity: do the indicators really capture the concept?
Variables — anything that varies across cases: age, income, turnout, political interest
Independent variable
the presumed cause
shapes?
Dependent variable
the presumed effect
Operationalisation — turning an abstract concept into something measurable
Political
engagement
abstract: cannot be
observed directly
observable indicators
Did you vote?
How often do you discuss politics?
Are you a party member?
Validity:
do the indicators
really capture
the concept?
RUNNING EXAMPLE
political interest: a 0–10 self-rating
RUNNING EXAMPLE
turnout in the last national election
Sampling techniques
Probability and non-probability sampling from the same population
On the left, a population of 64 people, 16 of them younger voters (dark dots) and the rest older voters (light dots); the younger voters are concentrated in one corner, the part of the population that is easiest to reach. Probability sampling: every member has a known, non-zero chance of selection, for example simple random sampling; its 10 selections are scattered across the population, and the sample, with 3 younger voters, mirrors the population; it supports generalisation and statistical inference, but needs a complete list of the population. Non-probability sampling: participants are chosen by specific criteria or convenience, for example convenience, quota or snowball sampling; its 9 selections all come from the easy-to-reach corner, and the sample, with 7 younger voters, is skewed; it is quicker and cheaper, but carries a higher risk of bias. At the bottom: the central concern is representativeness, how well the sample reflects the population it is meant to describe. Running example: the population is all eligible voters, and a random sample of 1,500 lets us compare turnout across age groups.
Population
the whole group the study is meant to describe
younger voters
older voters
Probability sampling
every member has a known, non-zero chance of selection
e.g. simple random sampling
sample
✓ mirrors the population: supports generalisation
and statistical inference, but needs a complete list of it
Non-probability sampling
participants chosen by specific criteria or convenience
e.g. convenience, quota or snowball sampling
sample
✗ skewed towards whoever is easiest to reach
quicker and cheaper, but a higher risk of bias
Representativeness: how well does the sample reflect the population it is meant to describe?
ALL ELIGIBLE VOTERS
1,500 AT RANDOM
Data collection methods
Four methods of data collection, each with a strength and a limitation
Four cards. Surveys: standardised questions for many respondents, via a structured questionnaire; strength, breadth and comparability; limitation, weak on depth and context. Observation: behaviours or events recorded as they occur, with or without participating; strength, what people actually do, not what they say; limitation, labour-intensive and hard to generalise from. Interviews: direct questioning, from tightly structured to open and conversational; strength, depth and nuance; limitation, time-consuming and shaped by the relationship between interviewer and interviewee. Archival sources: pre-existing records, documents and datasets; strength, unobtrusive, so suited to the past and to sensitive topics; limitation, limited to the records that happen to exist. No method is best in the abstract. Running example: a structured survey, outlined in orange, is the natural fit for measuring turnout across a large, representative sample.
Surveys
standardised questions for many
respondents, via a structured questionnaire
+ breadth and comparability
− weak on depth and context
Observation
behaviours or events recorded as they
occur, with or without participating
+ what people actually do, not what they say
− labour-intensive; hard to generalise from
Interviews
direct questioning, from tightly
structured to open and conversational
+ depth and nuance
− time-consuming; shaped by the relationship
Archival sources
pre-existing records, documents
and datasets
+ unobtrusive: the past, sensitive topics
− limited to the records that happen to exist
RUNNING EXAMPLE: A STRUCTURED SURVEY
Ethics in social research
Three ethical principles, and ethical review before data collection
Three cards. Informed consent: participants understand what the research involves and agree freely, without coercion. Privacy and confidentiality: personal information is safeguarded and identities are protected. Avoiding harm: anticipate and minimise physical, psychological or social discomfort. Below, the safeguard as a sequence: design the study, then ethical review by an independent committee, then collect data. No data are collected until the study has passed review.
Informed consent
participants understand what
the research involves, and
agree freely, without coercion
Privacy and confidentiality
personal information
safeguarded; identities
protected
Avoiding harm
anticipate and minimise
physical, psychological
or social discomfort
The safeguard: no data are collected until the study has passed review
Design the study
Ethical review
by an independent committee
Collect data
Data analysis and interpretation
Two routes from raw data to an answer to the research question
Two routes converge on one box, the answer to the research question. Qualitative analysis starts from texts such as transcripts and identifies themes and patterns of meaning, through thematic or content analysis. Quantitative analysis starts from numbers and summarises and models them with descriptive statistics, regression, t-tests and related techniques. A warning under the answer: it is only as trustworthy as the measures behind it. Running example, in two steps: first compare turnout across age bands, which describes the pattern; then use regression to ask whether political interest accounts for the gap, which explains it.
“
QUAL
Qualitative analysis
themes and patterns of meaning:
thematic or content analysis
7 3 9 5 8 2
QUAN
Quantitative analysis
summarise and model the data:
descriptive statistics, regression, t-tests
An answer to the
research question
!
only as trustworthy as
the measures behind it
RUNNING EXAMPLE
1
Compare turnout across age bands — describes the pattern
2
Regression: does political interest account for the gap? — explains it
Validity and reliability
Validity and reliability pictured as an archer's target
Two definitions at the top. Validity means accuracy: are we measuring what we intend to measure? Reliability means consistency: would the same procedure give the same result again? Below, three archery targets. First, reliable but not valid: six arrows tightly clustered, but in the lower left, off the bullseye; consistently wrong. Second, valid but not reliable: six arrows scattered around the bullseye, centred on it on average, but none of them on it. Third, reliable and valid: six arrows tightly clustered on the bullseye.
Validity = accuracy
are we measuring what we intend to measure?
Reliability = consistency
would the same procedure give the same result again?
Reliable, not valid
tightly clustered, off the bullseye:
consistently wrong
Valid, not reliable
scattered around the bullseye,
none of them on it
Reliable and valid
tightly clustered
on the bullseye
Presenting research findings
The same finding presented two ways
Two bar charts of the same illustrative data: turnout by age group, rising from 51 per cent among 18 to 24 year olds through 59, 66, 71 and 75 per cent to 78 per cent among those aged 65 and over. Left, hard to grasp: a jargon title (Fig. 3: self-reported electoral participation, per cent, by age cohort, weighted, n equals 1,500), heavy gridlines, bars in six unrelated colours, no labels on the bars and a separate legend to decode them, and an axis titled with a questionnaire code. Right, easy to grasp, the running example: a plain-language headline (under-25s are the least likely to vote), one colour with the youngest group highlighted, values written on the bars and age groups under them, and one sentence relating the chart back to the question: the age gap our question asks about is real; next, we ask whether political interest explains it. Three numbered callouts mark the three habits: 1, a plain-language headline; 2, a clean, directly labelled chart; 3, one sentence relating it back to the question.
Hard to grasp
Fig. 3: Self-reported electoral participation
(%) by age cohort, weighted (n = 1,500)
0
10
20
30
40
50
60
70
80
90
100
% Q17 = 1
18–24
25–34
35–44
45–54
55–64
65+
Easy to grasp
RUNNING EXAMPLE · ILLUSTRATIVE DATA
Under-25s are the least likely to vote
turnout at the last national election, by age (%)
51
18–24
59
25–34
66
35–44
71
45–54
75
55–64
78
65+
The age gap our question asks about is real;
next, we ask whether political interest explains it.
1
2
3
1
plain-language headline
2
clean, directly labelled chart
3
one sentence relating it back to the question
Limitations and critiques
Acknowledging limitations, and how they move knowledge forward
Three practices at the top: 1, acknowledge the weaknesses of the design, data and analysis, openly; 2, invite critique from peers, reviewers and the wider community; 3, point forward, since the limits show what future research should address. Below, how knowledge accumulates: a study leads to the next study, and that to the next, because each study's limits set the next study's questions. Running example: our survey's limitation is that self-reported turnout is over-stated, because voting is socially approved; the next study validates self-reports against official voting records.
1
Acknowledge
the weaknesses of the design,
data and analysis, openly
2
Invite critique
from peers, reviewers and
the wider community
3
Point forward
the limits show what future
research should address
How knowledge accumulates: each study's limits become the next study's questions
A study
its limits set
the questions
The next study
its limits set
the questions
… and the next
RUNNING EXAMPLE
Our survey's limitation: self-reported turnout is over-stated, because voting is socially approved
The next study: validate self-reports against official voting records
Conclusion
Methodology matters: the quality of our conclusions depends on the quality of our methods
Social research is an evolving practice — new data sources, tools, and techniques continually reshape what is possible
The choices made at each stage — topic, design, sampling, collection, analysis, presentation — are connected and consequential
We followed one question — why young people vote less — from a vague idea to a defensible, if imperfect, answer; every study makes that same journey
Questions and discussion are welcome
Model answers: research or not?
Model answers to Task 1: research or not?
Eight cards, one per scenario, each with a verdict and the criterion that decides it. Scenario 1, store visits: not research. Missing: an explicit procedure. Stores and staff picked haphazardly; "basically fine" is an impression, not evidence. Scenario 2, engagement survey: research. Same questionnaire, same month, all 450 staff, documented: explicit, transparent, repeatable. Scenario 3, ceo's memo: not research. Missing: evidence. Two articles and the CEO's authority settle the claim; no data were gathered. Scenario 4, warehouse log: research. A predefined protocol and a structured daily log: systematic observation in context. Scenario 5, checkout a/b test: research. Random assignment and recorded outcomes: an experiment anyone could repeat. Scenario 6, four-day-week poll: research, but bad. A procedure and real data, but a self-selected sample: 78% of readers who chose to click, not of employees. Scenario 7, turnover figures: not research. Missing: theory. Orderly, repeatable record-keeping, but no question it was designed to answer. Scenario 8, team interviews: research. Same open questions, recorded with consent, coded for themes: a small qualitative study.
1
Store visits
NOT RESEARCH
Missing: an explicit
procedure. Stores and staff
picked haphazardly;
“basically fine” is an
impression, not evidence.
2
Engagement survey
RESEARCH
Same questionnaire, same
month, all 450 staff,
documented: explicit,
transparent, repeatable.
3
CEO's memo
NOT RESEARCH
Missing: evidence. Two
articles and the CEO's
authority settle the
claim; no data were
gathered.
4
Warehouse log
RESEARCH
A predefined protocol and
a structured daily log:
systematic observation
in context.
5
Checkout A/B test
RESEARCH
Random assignment and
recorded outcomes: an
experiment anyone could
repeat.
6
Four-day-week poll
RESEARCH, BUT BAD
A procedure and real data,
but a self-selected sample:
78% of readers who chose
to click, not of employees.
7
Turnover figures
NOT RESEARCH
Missing: theory. Orderly,
repeatable record-keeping,
but no question it was
designed to answer.
8
Team interviews
RESEARCH
Same open questions,
recorded with consent,
coded for themes: a small
qualitative study.
Model answers: what kind of research?
Model answers to Task 2: the research scenarios on the two axes
The two-by-two grid from the lecture: columns qualitative and quantitative, rows experimental and descriptive. Quantitative and descriptive: scenario 2, the engagement survey (standardised items, compared year on year), and scenario 6, the four-day-week poll (a bad sample, but still this type). Quantitative and experimental: scenario 5, the checkout A/B test (random halves see version A or B, so conditions are deliberately varied). Qualitative and descriptive: scenario 8, the team interviews (open questions, coded for themes). Scenario 4, the warehouse log, sits across the boundary between qualitative and quantitative in the descriptive row: it is the odd one out, combining a predefined protocol whose entries can be counted with immersion in context aimed at meaning. The qualitative and experimental cell holds no scenario.
Axis 1: what kind of data and reasoning?
Qualitative
deep, interpretative: meaning and experience
Quantitative
numerical, statistical: measurement and pattern
Axis 2: does the researcher intervene?
Experimental
controls and
varies conditions
Descriptive
observes only:
no intervention
2
Engagement survey
2: standardised items, compared yearly
5
Checkout A/B test
random halves see version A or B:
conditions deliberately varied
6
Four-day-week poll
6: a bad sample, but still this type
8
Team interviews
open questions, coded for themes
4
Warehouse log
The odd one out: a predefined protocol (countable) and immersion in context (meaning)
no scenario here
Model answers: record-keeping or research?
Model answer to Task 3: the line between record-keeping and research
An arrow runs from record-keeping on the left to research on the right. Left of a dashed vertical line, scenario 7, the turnover figures: systematic, documented and repeatable, but compiled for a reporting template, not to answer a question. Right of the line, scenario 2, the engagement survey: designed to measure a concept, engagement, and track it, asking how it is changing and where. The dashed line is labelled: the line is a research question the data are designed to answer. Underneath: the turnover figures cross the line when the same figures are used to answer a question, such as why turnover is higher in some stores than others; the engagement survey sits close to the line, so it is defensible either way: if HR only files the results it is monitoring, and if it asks why scores fell it is research.
Record-keeping
Research
7
Turnover figures
systematic, documented and repeatable —
but compiled for a reporting template,
not to answer a question
2
Engagement survey
designed to measure a concept
(engagement) and track it:
how is it changing, and where?
The line: a research question the data are designed to answer
7 crosses the line when the same figures
are used to answer a question: why is
turnover higher in some stores than others?
2 sits close to the line, so it is defensible
either way: if HR only files the results, it is
monitoring; if it asks why scores fell, research
Model answers: fixing scenario 6
Model answer to Task 3: why scenario 6 is still bad research, and how to fix it
Left, why it is still bad research. The claim, "78% of employees want a four-day week", is based on a poll of the consultancy's own newsletter subscribers who chose to click through. The main problem is who answered: self-selected readers, not a sample of employees, so the 78% describes them, not the population the claim is about. Also: who asked, since the consultancy may gain from the result, and what was asked, since the wording is not published. Right, what it would take to fix it, in four steps: 1, define the population, the group the claim is about, for example all employees in Poland; 2, draw a probability sample, at random, from a list of that population (a sampling frame); 3, ask a neutral question and publish its exact wording; 4, report the method and its limits: sample size, response rate, and who commissioned it.
Why it is still bad research
“78% of employees want a four-day week”
based on
a poll of its own newsletter subscribers
who chose to click through
✗ Who answered: self-selected readers, not a
sample of employees. The 78% describes them,
not the population the claim is about.
Also: who asked (the consultancy may gain from
the result) and what was asked (the wording
is not published)
What it would take to fix it
1
Define the population
the group the claim is about,
e.g. all employees in Poland
2
Draw a probability sample
at random, from a list of that
population (a sampling frame)
3
Ask a neutral question
and publish its exact wording
4
Report the method and its limits
sample size, response rate,
and who commissioned it