Introduction to Social Research Methodology

Qualitative data analysis

Ben Stanley

Department of Social Sciences, SWPS University

January 19, 2027

Today’s seminar

  • From talk to interpretation — what qualitative analysis is, and the spiral of steps that turns raw text into a defensible finding
  • Coding in practice — what a code is, how open coding works, and how a codebook keeps a team honest
  • From codes to findings — thematic analysis as the workhorse method, content analysis as its more structured cousin, and how to use quotations as evidence
  • Rigour — how qualitative analysts earn trust: intercoder comparison, negative cases, the audit trail
  • By the end, you should be able to take a pile of transcripts and produce something a sceptical board would believe

From talk to interpretation

What is qualitative data analysis?

  • The process of interpreting and making sense of non-numerical data — text, images, audio, video — to understand concepts, experiences, and social phenomena
  • The raw material is usually talk and text: interview transcripts, focus-group recordings, open-ended survey comments, field notes, documents
  • Organisations generate this material constantly: exit interviews, customer complaints, appraisal comments, meeting minutes
  • The task is to travel from raw talk to defensible interpretation — a claim about what the data mean that you could justify to a critic
  • Data never speak for themselves: someone must decide what counts as a pattern, and that someone is you

Running example: Meridian’s junior research team has twelve exit-interview transcripts from staff who quit last quarter. The board asks: what are they telling us?

Why analysis is the hard part

  • Data overload — twelve one-hour interviews produce well over a hundred pages of transcript; richness becomes overwhelming without a system
  • Subjectivity — the analyst’s own perspective shapes what they notice; two readers can honestly see different things in the same page
  • Context dependence — the same sentence can mean different things depending on who said it, when, and in response to what
  • No single algorithm — there is no formula that turns text into findings the way a regression turns numbers into coefficients
  • The answer to all four problems is the same: explicit, systematic, documented procedure — the theme of this whole course, applied to words

In qualitative analysis, validity means accuracy of interpretation, and reliability means consistency of approach across analysts (Babbie’s terms travel here too)

The analysis spiral

  • Qualitative analysis moves through recognisable stages — but as a spiral, not a straight line: you loop back constantly
    1. Familiarisation — read and re-read the material; note first impressions
    2. Coding — attach short labels to meaningful segments of the data
    3. Categorising — group related codes together
    4. Theming — identify the broader patterns the categories point to
    5. Interpreting — say what the patterns mean for the research question
  • Each pass through the data refines the one before: a theme discovered late sends you back to re-code earlier transcripts
  • This looping is not inefficiency — it is how the method achieves depth; the analysis is finished when new passes stop changing the picture

Coding in practice

What is a code?

  • A code is a word or short phrase that captures the essence of a segment of data — a label saying “this bit is about that”
  • Codes are the bridge between raw data and analysis: they let you find, compare, and count every place where the same thing comes up
  • A segment can be a phrase, a sentence, or a whole paragraph — and one segment can carry several codes at once
  • Good codes are close enough to the data to be checkable, and abstract enough to apply across more than one transcript

Running example: the sentence “my rota changed three times last month and I found out from a group chat” might carry the codes scheduling chaos and poor communication

Descriptive and interpretive codes

  • Descriptive codes stay close to the surface of what was said — they summarise topic and content
    • mentions pay, talks about rota, praises supervisor
  • Interpretive codes name what the analyst takes the passage to mean — they add a layer of inference
    • perceived pay inequity, loss of control over time, loyalty to person not firm
  • Neither type is superior: descriptive codes are safer and easier to agree on; interpretive codes carry more analytical value but need more justification
  • A common workflow: code descriptively first, then re-code interpretively once you know the data well — the spiral again

Open coding

  • Open coding (a term inherited from grounded theory) means working through the data line by line, labelling freely, without a predetermined list
  • Generate as many codes as the data demand — do not force segments into a scheme you brought with you
  • Stay open to every theoretical direction: the point of open coding is discovery, letting the data surprise you
  • Expect mess: early open coding produces dozens of overlapping labels — that is normal, and the next stage tidies it
  • Compare constantly: each newly coded segment is checked against segments already carrying the same code — do they really belong together?

Running example: open-coding Meridian transcript #1 yields pay comparison, rota chaos, good line manager, HQ distance, team friendship — before any theory of why people quit

Building a codebook

  • A codebook turns a private coding habit into a shared, checkable instrument — essential the moment more than one person codes
  • Each entry needs four things:
Element What it does
Label Short, memorable name for the code
Definition What the code means — the concept it captures
Inclusion / exclusion rules When to apply it, and when not to
Example quote An anchor showing a clear case from the data
  • The codebook is a living document: codes are merged, split, and redefined as analysis proceeds — but every change is recorded and dated
  • A codebook another analyst can apply without asking you questions is the qualitative equivalent of a reliable measurement instrument

One passage, two coders

“Honestly the money was part of it, but it’s more that nobody above store level ever asked what we thought. My manager was great — she fought for us — but she was fighting the same system we were.”

Coder A (descriptive) Coder B (interpretive)
Codes applied mentions pay, praises manager, criticises HQ voice denied, manager as buffer, systemic not personal blame
  • Both codings are defensible — they operate at different altitudes above the data
  • Coder A’s labels are easier to agree on; Coder B’s carry more explanatory promise but need evidence from elsewhere in the data
  • The disagreement is productive: discussing it forces the team to decide what the study is actually trying to capture — and to write the decision into the codebook

From codes to findings

Thematic analysis: the workhorse

  • Thematic analysis identifies, analyses, and reports patterns (themes) within qualitative data — the most widely used qualitative method in organisational research
  • It examines the perspectives of different participants, highlighting commonalities and differences and generating unanticipated insights
  • It captures collective or shared meanings: not what one person said, but what recurs across the dataset
  • Flexible: works with interviews, focus groups, open survey comments, documents — any textual material
  • Its outputs are exactly what a decision-maker needs: a small number of well-evidenced themes, each with a clear definition and supporting extracts

The six steps of thematic analysis

  • The standard sequence (associated with Braun and Clarke) runs:
    1. Familiarisation — immerse yourself; read, re-read, note initial ideas
    2. Generating initial codes — systematically code the whole dataset, collating segments under each code
    3. Searching for themes — sort codes into candidate themes; gather all data relevant to each
    4. Reviewing themes — check each theme against the coded extracts and the full dataset; draw a thematic map
    5. Defining and naming themes — write a clear definition and a telling name for each; refine the overall story
    6. Producing the report — select vivid, compelling examples and relate the themes back to the research question
  • Steps 3–5 loop: themes that fail review are split, merged, or abandoned — the spiral inside the method

What makes a good theme?

  • A theme is not just a topic: “pay” is a topic; “pay grievances are really fairness grievances” is a theme — it claims something
  • A good theme captures something important in relation to the research question — importance is not the same as frequency
    • Something said by few people can still be a theme if it is analytically central; something said by everyone can be trivial
  • Themes must be grounded: every theme needs multiple supporting extracts, ideally from multiple participants
  • Themes should be distinct but connected — minimal overlap between them, and together they tell a coherent overall story

Running example: “scheduling” (topic) becomes the theme “unpredictable hours make staff feel interchangeable” — a claim Meridian’s board can act on

Qualitative content analysis

  • Content analysis systematically analyses the content of communication — documents, media coverage, open-ended answers — using an explicit category scheme
  • The workflow is more structured than thematic analysis:
    1. Define the research question and select the material
    2. Develop categories and a coding scheme — often in advance, from theory
    3. Pre-test the scheme on a small sample; revise it
    4. Code the full sample with the fixed scheme
    5. Analyse — including counting: frequencies, co-occurrences, comparisons
  • The key differences from thematic analysis: categories are typically fixed before full coding, and results are often quantified — content analysis sits closest to the quantitative border

Content analysis in action

  • A classic study: how are women in leadership roles portrayed in the media?
  • Sample: 100 newspaper articles — 50 discussing men in the workplace, 50 discussing women
  • Coding categories fixed in advance: titles, competence descriptors, emotional descriptors, personal-life references, leadership language, appearance
  • Selected findings:
Category Women Men
Called “leaders” / “innovators” 20% of articles 50%
Family or marital status mentioned 60% 20%
Appearance mentioned 40% 5%
  • Conclusion: media descriptions of professionals reproduce gender stereotypes — competence and leadership for men; emotion, family, and appearance for women
  • Note the form of the finding: counted categories, not interpreted themes — that is content analysis’s signature

Memo-writing

  • A memo is a note the analyst writes to themselves during analysis — the running commentary of the project
  • Memos capture:
    • Hunches — “pay complaints always seem to follow a comparison with a named competitor”
    • Decisions — “merged rota chaos and shift swaps into scheduling unpredictability, because…”
    • Emerging interpretations — first drafts of what a theme might mean
  • Write memos from day one — the best analytical ideas arrive during coding and are lost if not captured immediately
  • Memos are not overhead: they become the first draft of your findings and a key part of the audit trail

Quotations as evidence

  • In the final report, quotations are your evidence — they show the reader the data behind each claim
  • A good quotation is vivid (it lands) and representative (it stands for a pattern, not just itself)
  • The great danger is cherry-picking: choosing the one dramatic quote that supports your story while ignoring the ninety pages that complicate it
  • Defences against cherry-picking:
    • State how many participants the quoted view represents (“seven of twelve leavers…”)
    • Quote the counter-voices too — the participant who disagrees
    • Never let a single quotation carry a theme alone
  • A quotation illustrates a finding; it is the pattern, honestly reported, that is the finding

Rigour

Can qualitative analysis be rigorous?

  • The standard objection: “this is just your opinion about what people said”
  • The answer is not to imitate statistics but to make the interpretive process transparent, systematic, and checkable
  • Four practical disciplines do most of the work:
    1. Intercoder comparison — do independent analysts see the same things?
    2. Negative-case analysis — have you hunted for the evidence against yourself?
    3. The audit trail — can an outsider trace how you got from data to claims?
    4. Description–interpretation discipline — is it always clear what participants said versus what you concluded?
  • Together these turn “trust me” into “check me” — which is what rigour means in any method

Intercoder comparison

  • Two (or more) analysts independently code the same material, then compare
  • Where they agree, confidence rises; where they disagree, something valuable surfaces:
    • a vague codebook definition that needs tightening
    • a genuinely ambiguous passage worth discussing
    • a difference in interpretive stance the team must resolve
  • The process: code separately → compare segment by segment → discuss every disagreement → revise the codebook → re-code
  • In quantitative content analysis this is formalised as intercoder reliability statistics; in thematic work the structured discussion matters more than the number
  • Disagreement is not failure — undiscussed disagreement is

Negative-case analysis

  • Once a theme starts to form, the temptation is to see it everywhere — confirmation bias works on coders too
  • Negative-case analysis is the deliberate hunt for data that contradicts your emerging interpretation
  • Ask of every theme: who in the dataset does this not fit — and why?
  • Three possible outcomes, all good:
    • the exception leads you to refine the theme’s boundaries (“this holds for store staff, not warehouse staff”)
    • it reveals a rival explanation you must now address
    • the theme survives scrutiny — and your claim is stronger for the test

Running example: if the theme is “people leave over pay”, the leaver who took a pay cut to go elsewhere is the most informative transcript in the pile

The audit trail

  • The audit trail is the documented record of the whole analytical journey: raw data → codes → categories → themes → claims
  • It contains:
    • all codebook versions, with dates and reasons for changes
    • the memos recording hunches, decisions, and dead ends
    • the mapping from each theme to its supporting extracts
  • Purpose: an outsider could reconstruct and challenge every analytical decision — the qualitative counterpart of sharing your data and code
  • It also protects you: six weeks later, “why did we merge those two codes?” has a written answer
  • This is our course’s oldest theme again: systematic and transparent procedure is what separates research from impression

Description versus interpretation

  • The final discipline: keep what the data say and what you conclude visibly separate
  • Description: “Nine of twelve leavers compared their pay with a named competitor, unprompted.”
  • Interpretation: “Pay dissatisfaction at Meridian is driven by external benchmarking rather than absolute pay levels.”
  • Both belong in the report — but the reader must always be able to tell which is which
  • Blurring them is how weak analyses smuggle opinions in as findings; separating them is how strong analyses invite scrutiny
  • A practical habit: for every claim you write, ask “could a participant have said this?” — if yes, it is description; if only the analyst could say it, it is interpretation and needs its evidence shown

Software helps, thinking decides

  • Dedicated packages — MAXQDA, NVivo, Atlas.ti — support qualitative analysis at scale
  • What they do well: store and organise transcripts, attach codes, retrieve every segment under a code instantly, visualise code co-occurrence, keep the audit trail tidy
  • What they do not do: decide what a code means, judge whether a theme is grounded, or interpret anything
  • The software manages the clerical burden; every analytical decision remains human
  • A useful rule: if your analysis would be weak on paper, it will be weak in NVivo — only faster

Conclusion

Conclusion

  • Qualitative analysis is the disciplined journey from raw talk to defensible interpretation — a spiral of familiarising, coding, categorising, theming, and interpreting
  • Codes are the working unit: descriptive or interpretive, generated openly, governed by a shared codebook
  • Thematic analysis is the workhorse — six steps from immersion to report; content analysis is its more structured, countable cousin
  • Quotations are evidence, not decoration — vivid, representative, and honest about prevalence
  • Rigour is earned through intercoder comparison, negative cases, the audit trail, and the description–interpretation discipline — turning “trust me” into “check me”
  • Software carries the files; you carry the thinking
  • Questions and discussion are welcome

Exercise

Today’s exercise: Coding the exit interview

QR code linking to the exercise worksheet

bdstanley.netlify.app/social-research-methodology-14-exercise

Study guide

Full summary of this session, for revision: Qualitative data analysis

QR code linking to the session handout

bdstanley.netlify.app/social-research-methodology-14-handout