Introduction to Social Research Methodology

Sampling and participant recruitment

Ben Stanley

Department of Social Sciences, SWPS University

December 8, 2026

Today’s lecture

  • From population to sample — who we want to know about, who we can actually list, and the gap between the two
  • Probability sampling designs — simple random, systematic, stratified, and cluster sampling, and when each earns its keep
  • Non-probability sampling designs — convenience, purposive, quota, and snowball sampling, and their legitimate uses
  • Error, bias, and sample size — why a small random sample beats a huge self-selected one
  • Recruiting participants — the practical craft of getting the people you selected to actually take part

From population to sample

Why sampling matters

  • A critical part of any study is deciding what to observe and what not to observe — the process of selecting observations is called sampling
  • We almost never study everyone: surveying every customer or every citizen would be impossibly slow and expensive
  • Done well, sampling is astonishingly efficient: election polls of only 1,000 respondents routinely estimate national results with impressive accuracy
  • Done badly, it quietly wrecks a study: no analysis, however sophisticated, can repair a sample that was chosen in a biased way
  • The same logic that lets 1,000 voters stand in for a nation lets 90 employees stand in for a workforce — if they are chosen properly

Running example: Meridian’s board wants to know why turnover has doubled. We cannot interview all 450 employees — so whom do we ask?

Population, element, sample

  • An element is the unit about which information is collected — usually a person, but it can be a store, a team, or a transaction
  • The population is the entire aggregation of elements we want to draw conclusions about
  • The sample is the subset of elements we actually study, chosen to stand in for the whole
  • Defining the population sounds trivial but rarely is: “Meridian employees” could mean current staff, current and former staff, permanent contracts only, or everyone including gig couriers
  • Every claim a study makes is a claim about its population — so define it before selecting anyone

Running example: for a turnover study, the population that matters arguably includes people who have already left — the very group a staff list excludes.

The sampling frame

  • A sampling frame is the actual list of elements from which a probability sample is drawn
  • Examples: a university’s student register, a company’s HR roster, a customer database, Poland’s PESEL register of citizens
  • A properly drawn sample describes the population composing the frame — not necessarily the population you had in mind
  • Good organisational frames exist for employees, members, and registered customers; they rarely exist for informal or transient groups
  • The frame also constrains the contact mode: a frame of postal addresses suits PAPI, e-mail addresses suit CAWI, phone numbers suit CATI — and each mode reaches some people better than others

When the frame is not the population

  • The frame–population gap is one of the commonest hidden flaws in applied research
  • Meridian’s HR roster lists everyone employed today — but “everyone who worked at Meridian this year” includes dozens of leavers who are not on it
  • A turnover study sampled from the roster systematically excludes the people with the most direct knowledge of why staff leave
  • Other classic gaps: a landline frame under-covers the young; an e-mail frame under-covers warehouse staff without company accounts; a loyalty-card frame excludes occasional customers
  • Always ask: who is in the population but missing from the frame — and does their absence bias my answer?

Rule of thumb: state the population, name the frame, and list who falls through the gap — in writing, before fieldwork starts.

Representativeness

  • A sample is representative if its aggregate characteristics closely approximate the same aggregate characteristics in the population
  • If the workforce is 60% female, a representative sample contains close to 60% women; if 73% work in retail stores, so should roughly 73% of the sample
  • The basic principle: a sample tends to be representative of its population if all members of the population have an equal chance of selection
  • No sample is perfectly representative — the question is whether the deviations are small, random, and quantifiable, or large, systematic, and invisible

Why it matters to managers: an unrepresentative employee survey does not just miss the truth — it hands the board a confident, precise-looking wrong answer.

Probability sampling designs

The logic of random selection

  • In probability sampling, every element in the frame has a known, non-zero chance of selection; in random selection, chance alone decides
  • Two reasons to let chance choose:
    • It removes researcher bias — conscious or unconscious — in picking cases that suit a preferred conclusion
    • It unlocks probability theory, which lets us estimate population characteristics and say how precise those estimates are
  • Non-probability methods can be quick and useful, but only probability samples let us quantify our own uncertainty
  • The mechanics of that quantification — margins of error, confidence — come in session 13; today we need only the logic

Simple random sampling

  • The baseline design that statistical calculations assume
  • Procedure:
    1. Number every element in the sampling frame
    2. Use random numbers — once a printed table, now a random-number generator — to select elements until the sample is full
  • Every element, and every combination of elements, has an equal chance of selection
  • In practice it is used less often than you would expect: it demands a complete numbered frame and can be laborious to execute

Running example: number Meridian’s 450 employees 1–450 and let the random-number generator pick 90 of them.

Systematic sampling

  • A practical shortcut that is, in practice, virtually identical to simple random sampling
  • Compute the sampling interval k = population size ÷ sample size, then select every k-th element from the frame
  • A frame of 10,000 with a target sample of 1,000 gives k = 10: take every tenth person
  • Choose the starting point at random within the first interval — this is what keeps the method a probability design
  • One caution: if the list has a hidden periodicity (e.g. every 7th roster entry is a Sunday shift lead), the interval can align with it and bias the sample — check the ordering of the frame first

Stratified sampling

  • A refinement that increases representativeness by reducing probable sampling error
  • Divide the frame into homogeneous strata — groups that matter for your question — then sample randomly within each stratum
  • In a proportionate design, each stratum gets a share of the sample equal to its share of the population; on the stratifying variable, sampling error falls to zero
  • Organisations are naturally stratified: by division, site, seniority, contract type — and HR data usually records all of these
  • Disproportionate stratification deliberately over-samples small strata you need to say something precise about, then re-weights at the analysis stage

Running example: with 330 retail, 70 e-commerce, and 50 HQ staff, a simple random sample of 90 could catch e-commerce badly; stratifying guarantees each division its correct share.

Cluster sampling

  • Used when no single frame of elements exists — but elements come naturally packaged in listable clusters
  • Procedure: sample the clusters first, then list and sample elements within the selected clusters (multistage sampling)
  • The classic example: no list of all church members exists, but a list of churches does — sample churches, then members within them
  • Organisational version: no convenient frame of Polish retail workers, but a frame of retail companies exists — sample firms, then staff within them
  • The price: each stage adds sampling error, so cluster samples are less precise than a simple random sample of equal size — the trade is precision for feasibility

Running example: for in-store customer research, Meridian cannot list its customers — but it can list its 32 stores, sample 8 of them, and interview within those.

Comparing the four designs

Design How it selects Best when Watch out for
Simple random Random numbers over the whole frame A complete frame exists; baseline precision Laborious on large frames
Systematic Every k-th element, random start Long ordered lists Hidden periodicity in the list
Stratified Random within key subgroups Subgroups matter and are listed in the frame Choosing irrelevant strata
Cluster Sample groups, then members No element-level frame exists Extra sampling error per stage

In practice: real large-scale surveys usually combine these — e.g. stratify regions, cluster by locality, then sample systematically within clusters.

Non-probability sampling designs

When probability sampling is impossible

  • Much research is conducted in situations that simply do not permit probability sampling
  • Some populations have no frame and never will: homeless people, informal gig workers, users of a competitor’s product
  • Sometimes probability sampling is possible but not appropriate: a five-person exploratory study gains nothing from a random draw
  • Qualitative research in particular usually needs informative participants, not statistically representative ones
  • Non-probability sampling covers these cases — the designs are legitimate tools, provided you stay honest about what they can and cannot support

Running example: Meridian’s ex-employees and its self-employed delivery couriers both lack a clean frame — the roster has dropped the former and never held the latter.

Convenience sampling

  • Relying on whoever is available: people passing a street corner, visitors to a website, the employees who happen to be in the canteen
  • Sometimes called “haphazard” sampling — the journalist’s person-on-the-street interview is the classic case
  • It offers no control over representativeness: you learn about the people who were easy to reach, and only them
  • Defensible only when the convenient group is the object of interest, or as a cheap pilot before proper fieldwork
  • Even then, generalise with extreme caution — convenience is the weakest basis for any claim about a wider population

Purposive sampling

  • Also called judgmental sampling: the researcher deliberately selects cases using knowledge of the population and the study’s purpose
  • You choose participants because of who they are — the most experienced store managers, the newest hires, the teams with the best and worst retention
  • A comparative design can work well even without representativeness: sampling members of contrasting groups suffices for comparing them
  • The workhorse of qualitative research, where a handful of information-rich cases beats a random scatter of uninformative ones
  • The honest limit: purposive samples describe the cases chosen, not the population at large

Running example: to understand exceptional stores, deliberately pick Meridian’s three lowest-turnover and three highest-turnover stores and study the contrast.

Quota sampling

  • Starts from a matrix describing the target population: what proportion is male/female, young/old, in each division, on each contract type
  • Recruiters then fill each cell of the matrix with the right number of matching participants, so the sample mirrors the population’s known structure
  • The finished sample looks representative on the quota variables — this is how much commercial market research and online panel work is done
  • The catch: within each cell, selection is still non-random — whoever was easiest to recruit fills the quota
  • Looks like stratified sampling, but the family is different: stratified = random within strata; quota = convenient within cells

Snowball sampling

  • For populations whose members are hard to locate but know each other: undocumented migrants, informal networks, niche professionals
  • Procedure: find a few members of the target population, collect data, then ask each to refer you to others they know
  • The sample accumulates like a snowball rolling downhill — hence the name
  • Representativeness is questionable by construction: you reach one social network, not the population — so it is used primarily for exploratory purposes
  • In organisations: ex-employees, freelancers, and gig workers often form exactly this kind of referral-reachable population

Running example: three couriers Meridian can contact each know other couriers — a snowball is often the only way into that population.

Discuss: which design would you use?

  • For each brief, name a sampling design and defend it:
    • The board wants a trustworthy estimate of engagement across all 450 employees
    • You have two days and no budget to get any customer reaction to a new store layout
    • You need the views of the eight people who ran last year’s failed IT migration
    • You want to study informal WhatsApp groups that staff use to swap shifts
  • Notice the pattern: the question dictates the design — precision questions demand probability samples; depth, speed, and hidden-population questions justify non-probability ones
  • The cardinal sin is not using a non-probability sample — it is using one and then claiming population-level precision it cannot deliver

Error, bias, and sample size

Sampling error and sampling bias

  • Even a perfect random sample will not match the population exactly — the random, quantifiable wobble between sample and population is sampling error
  • Sampling error shrinks predictably as samples grow; probability theory tells us by how much (the mechanics arrive in session 13)
  • Sampling bias is different in kind: a systematic tendency to over- or under-represent certain sorts of people, built into how the sample was selected
  • Bias does not shrink as the sample grows — a bigger biased sample is just a more confident wrong answer
  • Error is the price of sampling; bias is a defect of design — you budget for the first and eliminate the second

Self-selection bias

  • The most common bias in organisational practice: letting participants choose themselves
  • A newspaper polling its own readers learns (a) nothing about the national mood — most citizens do not read that paper — and (b) little even about its readership, because only a certain type of reader fills in polls
  • The modern equivalent: a company polling its newsletter subscribers and reporting the result as customer opinion — subscribers are the most engaged customers, and respondents the most opinionated subscribers
  • Two filters stack: who is in the frame, then who bothers to answer — each one tilts the sample further from the population
  • Volunteers systematically differ from non-volunteers: more engaged, more opinionated, often more extreme in both directions

Running example: an open link in Meridian’s staff newsletter would over-hear the happiest and angriest employees — and miss the quietly disengaged middle who are actually drifting towards the exit.

How big does a sample need to be?

  • Intuition says a sample must be a large percentage of the population; statistics says otherwise — what matters is the sample’s absolute size
  • A well-drawn sample of 1,000 describes 38 million Poles about as precisely as it describes a city of 100,000 — the population size barely enters into it
  • Precision shows diminishing returns: going from 100 to 400 respondents helps a lot; going from 1,000 to 4,000 helps far less — quadrupling cost roughly halves error
  • Better a small sample drawn properly than a huge one drawn badly: 90 random employees beat 300 self-selected volunteers every time
  • Plan size around the smallest subgroup you must report on — 90 overall may leave only 14 e-commerce staff, which is thin for a division-level claim

Response rates and nonresponse bias

  • Selecting a sample is not the same as obtaining one: some of those invited will never respond
  • The response rate is the proportion of invited, eligible people who actually take part — decide in advance what counts as taking part (do partial completions count?)
  • Low response is dangerous not in itself but through nonresponse bias: when responders differ systematically from non-responders on what you are measuring
  • An engagement survey answered mainly by the engaged will overstate engagement — the disaffected are precisely the ones who ignore it
  • Report the response rate honestly, compare responders to the frame on known characteristics, and chase reminders — every extra reluctant respondent makes the sample less skewed

Running example: 300 CAWI invitations and 96 completes is a 32% response — and before celebrating, ask which third of the workforce answered.

Recruiting participants

Recruitment channels

  • Selecting people on paper is sampling; getting them to take part is recruitment — a craft of its own, and where surveys are won or lost
  • The main channels, each with a characteristic reach and bias:
    • Internal comms / intranet — cheap, broad, but misses non-desk staff and reads as “management asking”
    • E-mail invitations (CAWI) — targeted and trackable, but only reaches those with accounts, and competes with a full inbox
    • Access panels — pre-recruited pools of willing respondents, fast for customer research, but panellists are professional survey-takers
    • Social media — reaches beyond the organisation (ex-staff, customers), but self-selection is severe
    • In-store / on-site intercepts (PAPI) — reaches customers and shop-floor staff whom no list covers, at real fieldwork cost
  • Match the channel to the frame: a channel your target group does not use is a bias generator, not a recruitment tool

Crafting an invitation that works

  • The invitation is the single highest-leverage text in the project — most nonresponse happens at this hurdle, not mid-questionnaire
  • What a working invitation does:
    • Who and why — who is asking, and why this person was selected
    • What and how long — the topic and an honest time estimate (“10 minutes” that turns into 30 poisons the well for every future survey)
    • What happens to answers — anonymity or confidentiality, stated plainly, with who will see what
    • What it is for — the concrete decision the results will inform, and when results will be shared back
  • Keep it short, personal where possible, and send reminders — two polite reminders routinely add more respondents than the original invitation
  • Sender matters: an invitation from a neutral research team draws franker answers than the same text signed by the CEO

Incentives and their trade-offs

  • Incentives raise response rates — the question is what they do to the sample and the answers
  • Options, in rising order of cost and complication:
    • None — relies on goodwill and topic salience; fine for short, clearly useful internal surveys
    • Lottery / prize draw — cheap per respondent, mild effect, attracts the prize-motivated
    • Guaranteed small incentive (voucher, donation to charity) — the most reliable lift, and paid to everyone, not just the lucky
    • Paid participation — standard for interviews and panels and for busy or external groups; expect to pay for an hour of someone’s time
  • Trade-offs: incentives can recruit people interested in the reward rather than the topic, and paying employees for an internal survey can feel odd where work time already covers it
  • Time is the honest incentive inside an organisation: guarantee that participation happens on paid time, and much resistance dissolves

Reaching hard-to-reach groups

  • Some of the most decision-relevant voices are the hardest to recruit — and skipping them silently biases the study
  • Senior managers — gatekeepers and diaries are the obstacle: recruit top-down through a sponsor, offer scheduling flexibility, keep interviews short and sharply focused
  • Gig workers and contractors — not on the HR roster and paid per task: recruit through the app or dispatch channel they actually use, pay for their time, and consider snowballing through referrals
  • Ex-employees — no longer in any internal channel: use exit-interview consent lists, professional networks such as LinkedIn, or personal e-mail addresses retained with consent; a neutral external interviewer helps, since they owe the company nothing
  • Shift and shop-floor staff — have no desk and no quiet hour: go to them (intercepts in the back office, paper or tablet options, kiosk time during shifts)
  • The general principle: meet people in their own channel, on their own time, at a real value for their effort

Recruitment ethics

  • Recruitment is where research ethics (session 4) meets organisational hierarchy — and hierarchy can turn an invitation into an instruction
  • Consent must be genuinely voluntary: an e-mail from your line manager saying “please complete the survey” is, in practice, hard to refuse
  • Safeguards that keep participation free:
    • Invitations come from the research team, not the chain of command
    • No participation lists that managers can see; no chasing of named non-responders by supervisors
    • Explicit statement that declining carries no consequences — and organisational behaviour that makes it true
  • Anonymity promises must survive analysis: reporting results for a five-person team identifies individuals as surely as printing names
  • Incentives must not become coercion: rewards large enough that a low-paid worker cannot afford to decline undermine voluntariness

Running example: if Meridian’s store managers hand-pick which staff get interviewed about turnover, the study inherits every manager’s interest in looking good.

Conclusion

Conclusion

  • Sampling is the bridge between the people you study and the people you want to draw conclusions about — and every claim is only as strong as that bridge
  • Define the population, name the frame, and account for the gap between them before a single invitation goes out
  • Probability designs buy generalisation and quantifiable precision; non-probability designs buy feasibility, speed, and depth — the sin is claiming the first while doing the second
  • Absolute sample size beats percentage-of-population, but no sample size cures bias — self-selection and nonresponse do the quiet damage
  • Recruitment is sampling’s last mile: the right channel, an honest invitation, fair incentives, and a hierarchy kept at arm’s length
  • Next session we turn to what we actually ask the people we have recruited: questionnaire design
  • Questions and discussion are welcome

Exercise

Today’s exercise: Who do we ask?

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-10-exercise

Study guide

Full summary of this session, for revision: Sampling and participant recruitment

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-10-handout