Hypotheses and research design
Introduction to Social Research Methodology
From questions to hypotheses
A research question tells you what you want to know; it does not yet tell you what you expect to find or how you will find it. This session covers the two steps that turn a question into a workable plan: formulating a testable hypothesis, and choosing a research design capable of testing it. The running example throughout is Meridian, a Polish retail-and-services firm with 450 employees, 32 stores, and a growing e-commerce arm, where annual staff turnover has risen from 14% to 27% in two years and the board wants evidence rather than hunches.
Research assumptions
Every study rests on assumptions — statements accepted as true for the purposes of the research without being tested directly. Three kinds recur:
| Kind | Meaning | Meridian example |
|---|---|---|
| Theoretical | Propositions taken over from theory or prior research | Job dissatisfaction increases the likelihood of quitting |
| Methodological | Beliefs about how the chosen methods behave | Employees will answer an anonymous survey honestly |
| Practical | Beliefs about the data and setting | HR records of leavers are accurate and complete |
Assumptions are unavoidable, because no study can test everything at once. What distinguishes good research is that assumptions are stated explicitly, so that readers can judge whether they are reasonable. A hidden assumption is a weakness; a stated one is a boundary of the study.
What is a hypothesis?
A hypothesis is a tentative, testable statement about the relationship between two or more variables. It is not a question but a proposed answer: a specific expectation that data can support or contradict. Hypotheses are the engine of deductive research — theory suggests the hypothesis, and the study is built to test it. Compare the two forms: the question “does the quality of store management affect staff turnover at Meridian?” becomes the hypothesis “stores with poorer-quality first-line management have higher staff turnover”. The hypothesis commits the researcher: by stating what is expected, it also specifies what evidence would prove the expectation wrong.
Figure 1 draws the move from question to hypothesis: the question becomes a proposed answer, an arrow from the independent variable (quality of first-line management) to the dependent variable (staff turnover), labelled with the expected direction. The three cards underneath name what makes it a hypothesis: two named variables in a stated relationship, a commitment to what would prove it wrong, and its place as the engine of deductive research.
Properties of a good hypothesis
A good hypothesis satisfies five criteria:
| Property | Meaning |
|---|---|
| Testable and falsifiable | Observable evidence could, in principle, show it to be false |
| Clear and specific | The variables and the expected relationship are stated precisely |
| A statement of relationship | It links at least two variables rather than describing one thing |
| Grounded | It follows from theory, prior research, or informed observation |
| Value-neutral | It predicts what is, not what ought to be |
“Bad management is unacceptable” fails every test: it is a value judgement, names no measurable variables, and no evidence could falsify it. “Stores whose managers score lower on supervisor support have higher turnover” passes them all.
Types of hypotheses
Two distinctions organise the types of hypothesis. The first concerns direction. A non-directional hypothesis states that a relationship exists without specifying which way it runs (“manager quality is related to staff turnover”). A directional hypothesis specifies the expected direction (“the poorer the manager quality, the higher the staff turnover”). Directional hypotheses are riskier and therefore more informative — they can fail in more ways, so surviving the test means more. A directional form is appropriate when theory or prior evidence justifies an expectation; a non-directional form is appropriate when the direction is genuinely unknown.
Figure 2 shows why the directional form is riskier. Each panel asks which of three patterns of turnover against manager quality — falling, rising or flat — would support the hypothesis. The non-directional hypothesis is supported by either slope and refuted only by no relationship; the directional one is supported only by the falling slope, so it can fail in more ways and surviving the test means more.
The second distinction underpins statistical testing, which always works with a pair of hypotheses. The null hypothesis (H0) states that there is no relationship between the variables (“manager quality has no effect on staff turnover”). The alternative hypothesis (H1) states that a relationship does exist, and is usually the research hypothesis itself. The logic of testing is indirect: we ask how likely our data would be if the null were true, and reject the null only when the data make it implausible. We never “prove” the alternative — we accumulate evidence against the null. This cautious logic protects researchers from seeing patterns that are not there.
Figure 3 shows the pair of hypotheses and the logic of testing them: we ask how likely our data would be if H₀ were true, reject H₀ only when the data make it implausible, and never prove H₁.
From question to hypothesis at Meridian
The board’s question — why has staff turnover at Meridian doubled in two years? — is too broad to test directly. Theory and observation allow the research team to derive specific, testable expectations from it: (H1) employees who report lower supervisor support are more likely to quit within a year (directional); (H2) turnover is higher in stores than in the e-commerce division (comparative and directional); (H3) perceived pay fairness is related to intention to quit (non-directional). Each hypothesis carves off a piece of the big question that a single study can actually answer, and each demands the same three things: named variables, a stated relationship, and data that could prove it wrong.
Figure 4 draws this as carving: the broad question at the top, and three hypotheses cut from it, each with its type and its independent and dependent variables.
Variables and units of analysis
Independent, dependent, and control variables
A variable is any characteristic that varies across the cases under study. Within a hypothesis, variables play defined roles. The independent variable (IV) is the presumed cause — the thing that does the influencing. The dependent variable (DV) is the presumed effect — the outcome to be explained. In “supervisor support → intention to quit”, support is the IV and intention to quit the DV. Importantly, the roles come from the hypothesis, not from the variables themselves: turnover is a dependent variable when we explain it, but an independent variable when we ask whether turnover damages customer satisfaction.
Figure 5 shows the two roles, and that they belong to the hypothesis rather than to the variable: turnover is the dependent variable when workload explains it, and the independent variable when we ask whether it damages customer satisfaction.
A control variable is a third factor held constant, or measured and adjusted for, because it might otherwise distort the relationship of interest. Suppose stores with poor managers also pay less: is turnover driven by the manager, or by the pay? Without controlling for pay, the study risks a spurious conclusion — blaming managers for what wages are doing. Typical controls in a Meridian study would include pay level, the local unemployment rate, store size, and employee age and tenure. Identifying plausible controls is a test of how well the researcher understands the problem, since every control is a rival explanation taken seriously.
Figure 6 draws the pay example: the observed association between manager quality and turnover, and pay level pointing at both — stores with poor managers also pay less, and low pay drives people to leave — which is why, without the control, the conclusion would be spurious.
Units of analysis
The unit of analysis is the what or whom the study describes — the entities about which data are collected and conclusions drawn.
| Unit | Organisational examples |
|---|---|
| Individuals | Employees, customers, voters |
| Groups | Teams, departments, shifts |
| Organisations | Stores, firms, branches |
| Social artefacts | Exit interviews, complaint records, job adverts |
The same topic can be studied at different units: “which employees quit?” is an individual-level question, while “which stores have high turnover?” is an organisational-level one. Researchers must beware the ecological fallacy: conclusions about one unit do not automatically transfer to another. Knowing that a store has high turnover does not tell you which individual employees within it are at risk of leaving.
Figure 7 draws the units as nested boxes — individuals within groups within organisations — with social artefacts as a separate kind of unit, the same topic posed at two units, and the ecological fallacy underneath.
Putting the pieces together
For hypothesis H1 — employees who report lower supervisor support are more likely to quit within a year — the components line up as follows:
| Component | Choice |
|---|---|
| Unit of analysis | Individual employees |
| Independent variable | Perceived supervisor support |
| Dependent variable | Quitting within 12 months |
| Control variables | Pay level, tenure, contract type, store location |
Once this table is filled in, the study almost designs itself: we know whom to study, what to measure, and which rival explanations to rule out. This is the discipline the session exercise requires — for every hypothesis, name the unit, the IV, the DV, and at least one control.
Research designs
What is a research design?
A research design is the overall structure of a study: the plan that determines when, where, and on whom data are collected. It is the blueprint connecting the hypothesis to the evidence needed to test it — Creswell & Creswell describe it as the plan running from broad assumptions to detailed methods. Design is not the same as method: a survey (a method) can sit inside a cross-sectional, longitudinal, or comparative design. Five designs cover most organisational research, and the design question is always the same: what structure of evidence would allow this hypothesis to fail?
The five main designs
Cross-sectional designs collect data from many cases at a single point in time — a snapshot. They are the workhorse of organisational research: one employee survey across all 32 Meridian stores, fielded in a single week. They are fast, relatively cheap, and good for describing patterns and comparing groups at one moment, but weak on causal order: if dissatisfied employees quit, does dissatisfaction cause quitting, or do soon-to-quit employees become dissatisfied? A cross-sectional design can show that support and turnover intention are associated; it struggles to show which comes first.
In Figure 8, ten cases are drawn as timelines, all observed in one vertical band: a single moment. The boxes on the right give the Meridian version, the strengths and the weaknesses.
Longitudinal designs collect data over time, allowing change to be observed directly and causes to be placed before effects. Babbie distinguishes three variants: trend studies (the same population, different samples — annual engagement surveys of Meridian staff), cohort studies (following a defined group, such as everyone hired in 2024), and panel studies (the same individuals measured repeatedly — surveying the same 200 employees every six months). Longitudinal work is slow and expensive, and panels suffer attrition; in a turnover study, the very people we most want to follow are the ones who leave.
Figure 9 shows the three variants across three waves: a new sample each wave (trend), one defined group followed over time (cohort), and the same individuals measured every time (panel).
Case study designs investigate a single case in depth — one store, one team, one organisation — using multiple sources of evidence such as interviews, observation, and documents. They suit questions of how and why: why does the Katowice store keep its staff when comparable stores cannot? Their strength is depth, context, and the ability to surface mechanisms and generate hypotheses; their limit is that findings cannot be statistically generalised. One exceptional store proves what is possible, not what is typical.
Figure 10 shows the case study as a magnifying glass over one store in its context, with interviews, observation and documents as its sources of evidence.
Comparative designs systematically compare two or more cases, groups, or settings on the same dimensions. Meridian offers natural comparisons: stores versus e-commerce, high-turnover versus low-turnover stores, region against region. The logic is to let contrast do the explanatory work: if high- and low-turnover stores differ systematically in management style but not in pay, that is evidence worth having. Comparison approximates the logic of experiment where experiments are impossible, but compared cases differ in many ways at once, so rival explanations must be ruled out one by one.
In Figure 11, a high-turnover and a low-turnover store are compared on the same dimensions; where everything else is the same, the contrast points to what differs.
Experimental designs are the strongest for establishing cause: the researcher manipulates the independent variable and randomly assigns cases to treatment and control groups. Randomisation makes the groups equivalent on average, so a difference in outcomes can be attributed to the treatment — for instance, Meridian pilots a new onboarding programme in 16 randomly chosen stores while the other 16 continue as before. A quasi-experiment keeps the comparison but lacks random assignment (the programme goes to the stores whose managers volunteered). Quasi-experiments are common in organisations, where randomisation is often impractical, but self-selected groups may differ from the start, so causal claims require more caution.
Figure 12 contrasts random assignment, which makes the groups equivalent on average, with volunteering, where the groups may differ from the start.
Choosing among designs
| Design | Best for | Main limit |
|---|---|---|
| Cross-sectional | Describing patterns now; comparing groups | Weak on causal order |
| Longitudinal | Tracking change; ordering cause and effect | Slow, costly, attrition |
| Case study | Understanding how and why in depth | No statistical generalisation |
| Comparative | Learning from contrasts between units | Many differences at once |
| Experimental | Establishing cause | Often impractical or unethical |
No design is best in the abstract: the design must fit the question, the hypothesis, and the constraints of time, money, and access. Managerial problems usually signal the appropriate design in their own wording — now points to a cross-sectional snapshot, whether it works to an experiment or quasi-experiment, why this case to a case study, and changes over time to a longitudinal design. Managers rarely ask for a “research design”; translating their problem into one is precisely the researcher’s job.
Planning the research process
The stages of a research project
Babbie pictures a research project as a connected structure flowing from idea to application, with feedback loops throughout:
| Stage | What happens | Covered in |
|---|---|---|
| Interest, idea, theory | The problem is identified and hunches formed | Sessions 1–2 |
| Conceptualisation | The meaning of the key concepts is specified | Session 8 |
| Choice of research method | Survey, field research, existing data, and so on | Sessions 11–12 |
| Operationalisation | How the variables will actually be measured | Session 8 |
| Population and sampling | Whom the conclusions are about, and who will be observed | Session 10 |
| Observations | The data are collected | Sessions 11–12 |
| Data processing and analysis | Data are transformed, analysed, and conclusions drawn | Later sessions |
| Application | Results are reported and their implications assessed | — |
The stages loop as well as flow: the results of analysis feed back into the initial interests, ideas, and theories, and often this feedback begins another cycle of inquiry.
Schedule, resources, and practicalities
A plan is not just a list of stages — it attaches time, money, and people to each one. Building a schedule, even a rough timeline per stage, exposes whether the project fits its deadline: a board that wants answers in six weeks has just ruled out an eighteen-month panel study. The budget should cover the real costs of research: staff time, incentives for participants, software, travel, and data access. Practicalities need checking early — access to the stores, managers’ willingness to release staff for interviews, and the gatekeepers whose approval the project requires. Most research failures are planning failures: the design was fine, but the time and access it required were never there.
The research proposal
Before a study runs, its plan is usually written down for someone else’s review — a supervisor, a client, a board, or a funder. The research proposal lays out what you want to study, why it matters, and exactly how you will proceed. Writing it is not bureaucracy but the discipline of committing to decisions before data collection makes them irreversible; a proposal that cannot be written clearly is a study that has not been thought through clearly. For Meridian’s board, the proposal is the product: it turns “we should look into turnover” into a project someone can approve, fund, and hold the team to.
A short proposal typically contains the following sections:
| Section | Question it answers |
|---|---|
| Problem / objective | What exactly will you study, and why is it worth studying? |
| Literature review | What is already known, and what remains unresolved? |
| Research question & hypothesis | What do you ask, and what do you expect to find? |
| Subjects for study | Whom or what will you study, and how will they be selected? |
| Measurement | What are the key variables, and how will you measure them? |
| Data-collection method | How will the data actually be gathered? |
| Analysis | How will the data answer the question? |
| Schedule & budget | How long will it take, and what will it cost? |
In this course, the session exercise asks for a one-page skeleton — problem, question, hypothesis, design, and data needed — which later sessions will build on.
A Meridian mini-proposal
A worked example shows how compactly the whole plan can be stated. Problem: turnover has doubled to 27% in two years, and the board needs to know why before budgeting counter-measures. Question: does the quality of first-line management explain differences in turnover across Meridian’s stores? Hypothesis: stores whose employees report lower supervisor support have higher annual turnover, controlling for pay and local labour-market conditions. Design: a cross-sectional survey of all store employees (CAWI), linked to store-level HR turnover records. Data needed: survey measures of supervisor support and quit intention; HR data on leavers, pay, and tenure; regional unemployment figures. Five short entries — and the board can see the whole study, judge it, and decide whether to fund it.
Figure 13 sets out the same five entries as blocks.
Conclusion
A hypothesis converts a research question into a testable expectation — specific, falsifiable, and built from named variables. Variables have roles: the independent variable does the explaining, the dependent variable is explained, and control variables guard against spurious conclusions, all relative to a stated unit of analysis. Research designs — cross-sectional, longitudinal, case study, comparative, and experimental — are structures of evidence, and the craft lies in matching the design to the problem rather than defending a favourite technique. Finally, planning turns method into management: stages, schedule, resources, and a proposal that commits the plan to paper before the data commit you. The session followed Meridian’s turnover crisis from a board-room worry to a fundable one-page plan; the exercise, and later your own projects, make the same journey.