Qualitative data analysis

Introduction to Social Research Methodology

Author
Affiliation

Ben Stanley

Department of Social Sciences, SWPS University

Published

January 19, 2027

From talk to interpretation

Qualitative data analysis is the process of interpreting and making sense of non-numerical data — text, images, audio, and video — in order to understand concepts, experiences, and social phenomena. Its raw material is usually talk and text: interview transcripts, focus-group recordings, open-ended survey comments, field notes, and documents. Organisations generate this material constantly, in exit interviews, customer complaints, appraisal comments, and meeting minutes. The analyst’s task is to travel from raw talk to defensible interpretation: a claim about what the data mean that could be justified to a critic. The crucial point is that data never speak for themselves. Someone must decide what counts as a pattern, and that someone is the analyst — which is precisely why the analytical procedure must be explicit and systematic.

Why analysis is the hard part

Four difficulties make qualitative analysis genuinely demanding. First, data overload: a dozen one-hour interviews produce well over a hundred pages of transcript, and richness becomes overwhelming without a system. Second, subjectivity: the analyst’s own perspective shapes what they notice, and two readers can honestly see different things in the same page. Third, context dependence: the same sentence can mean quite different things depending on who said it, when, and in response to what. Fourth, there is no single algorithm — no formula turns text into findings the way a regression turns numbers into coefficients. The answer to all four problems is the same one this course has given from its first session: explicit, systematic, documented procedure. In qualitative work the familiar vocabulary of measurement quality translates as follows: validity means the accuracy of the interpretation, and reliability means consistency of approach across different analysts and occasions.

The analysis spiral

Qualitative analysis moves through recognisable stages, but as a spiral rather than a straight line: the analyst loops back constantly.

Stage What happens
Familiarisation Reading and re-reading the material; noting first impressions
Coding Attaching short labels to meaningful segments of the data
Categorising Grouping related codes together
Theming Identifying the broader patterns the categories point to
Interpreting Saying what the patterns mean for the research question

Each pass through the data refines the one before. A theme discovered late in the process sends the analyst back to re-code earlier transcripts in its light. This looping is not inefficiency; it is how the method achieves depth. The analysis is finished when new passes through the data stop changing the picture.

Coding in practice

A code is a word or short phrase that captures the essence of a segment of data — a label saying “this bit is about that”. Codes are the bridge between raw data and analysis: they make it possible to find, compare, and count every place in the dataset where the same thing arises. A coded segment can be a phrase, a sentence, or a whole paragraph, and one segment can carry several codes at once. Good codes are close enough to the data to be checkable and abstract enough to apply across more than one transcript. For example, the sentence “my rota changed three times last month and I found out from a group chat” might carry the codes scheduling chaos and poor communication.

Descriptive and interpretive codes

Codes operate at two levels of inference. Descriptive codes stay close to the surface of what was said, summarising topic and content: mentions pay, talks about rota, praises supervisor. Interpretive codes name what the analyst takes a passage to mean, adding a layer of inference: perceived pay inequity, loss of control over time, loyalty to person not firm. Neither type is superior. Descriptive codes are safer and easier for a team to agree on; interpretive codes carry more analytical value but require more justification. A common and sensible workflow is to code descriptively first and then re-code interpretively once the data are well known — the spiral at work again.

Open coding

Open coding, a term inherited from grounded theory, means working through the data line by line and labelling freely, without a predetermined list of codes. The analyst generates as many codes as the data demand, rather than forcing segments into a scheme brought to the data in advance, and remains open to every theoretical direction: the point of open coding is discovery. Early open coding produces dozens of overlapping labels, and this mess is normal — the categorising stage tidies it. Throughout, the analyst compares constantly, checking each newly coded segment against the segments already carrying the same code to ask whether they really belong together. Open-coding a single Meridian exit interview might yield pay comparison, rota chaos, good line manager, HQ distance, and team friendship before any theory of why people quit has been formed.

Building a codebook

A codebook turns a private coding habit into a shared, checkable instrument, and it becomes essential the moment more than one person codes. Each entry needs four elements:

Element What it does
Label A short, memorable name for the code
Definition What the code means — the concept it captures
Inclusion / exclusion rules When to apply the code, and when not to
Example quote An anchor showing a clear case from the data

The codebook is a living document: codes are merged, split, and redefined as the analysis proceeds, but every change is recorded and dated. A codebook that another analyst could apply without asking its author questions is the qualitative equivalent of a reliable measurement instrument.

One passage, two coders

Consider a passage from a Meridian exit interview: “Honestly the money was part of it, but it’s more that nobody above store level ever asked what we thought. My manager was great — she fought for us — but she was fighting the same system we were.” Two coders can treat this quite differently:

Coder A (descriptive) Coder B (interpretive)
Codes applied mentions pay, praises manager, criticises HQ voice denied, manager as buffer, systemic not personal blame

Both codings are defensible; they simply operate at different altitudes above the data. Coder A’s labels are easier to agree on, while Coder B’s carry more explanatory promise but need corroboration from elsewhere in the dataset. The disagreement is productive: discussing it forces the team to decide what the study is actually trying to capture, and to write that decision into the codebook.

From codes to findings

Thematic analysis: the workhorse

Thematic analysis identifies, analyses, and reports patterns — themes — within qualitative data, and it is the most widely used qualitative method in organisational research. It examines the perspectives of different participants, highlighting commonalities and differences and generating unanticipated insights, and it captures collective or shared meanings: not what one person said, but what recurs across the dataset. It is flexible, working with interviews, focus groups, open survey comments, documents, or any other textual material. Its outputs are exactly what a decision-maker needs: a small number of well-evidenced themes, each with a clear definition and supporting extracts.

The standard sequence, associated with Braun and Clarke, runs through six steps:

Step What happens
1. Familiarisation Immersion in the data; reading, re-reading, noting initial ideas
2. Generating initial codes Systematically coding the whole dataset, collating segments under each code
3. Searching for themes Sorting codes into candidate themes; gathering all data relevant to each
4. Reviewing themes Checking each theme against the coded extracts and the full dataset; drawing a thematic map
5. Defining and naming themes Writing a clear definition and a telling name for each theme; refining the overall story
6. Producing the report Selecting vivid, compelling examples and relating themes back to the research question

Steps 3 to 5 loop: themes that fail review are split, merged, or abandoned. This is the spiral operating inside the method itself.

What makes a good theme?

A theme is not just a topic. “Pay” is a topic; “pay grievances are really fairness grievances” is a theme, because it claims something. A good theme captures something important in relation to the research question, and importance is not the same as frequency: something said by few people can still be a theme if it is analytically central, while something said by everyone can be trivial. Themes must be grounded — each needs multiple supporting extracts, ideally from multiple participants — and the set of themes should be distinct but connected, with minimal overlap and a coherent overall story. At Meridian, the topic “scheduling” becomes the theme “unpredictable hours make staff feel interchangeable” — a claim the board can act on.

Qualitative content analysis

Content analysis systematically analyses the content of communication material — documents, media coverage, open-ended answers — using an explicit category scheme. Its workflow is more structured than thematic analysis: define the research question and select the material; develop categories and a coding scheme, often in advance and from theory; pre-test the scheme on a small sample and revise it; code the full sample with the fixed scheme; and analyse the results, including by counting frequencies, co-occurrences, and comparisons. The two key differences from thematic analysis are that categories are typically fixed before full coding begins, and that results are often quantified. Content analysis therefore sits closest of the qualitative methods to the quantitative border.

A classic illustration asks how women in leadership roles are portrayed in the media. A sample of 100 newspaper articles — 50 discussing men in the workplace, 50 discussing women — was coded with categories fixed in advance: titles, competence descriptors, emotional descriptors, personal-life references, leadership language, and appearance. Selected findings:

Category Women Men
Called “leaders” or “innovators” 20% of articles 50%
Family or marital status mentioned 60% 20%
Appearance mentioned 40% 5%

The conclusion was that media descriptions of professionals reproduce gender stereotypes: competence and leadership for men; emotion, family, and appearance for women. Note the form of the finding — counted categories rather than interpreted themes. That is content analysis’s signature.

Memo-writing

A memo is a note the analyst writes to themselves during analysis — the running commentary of the project. Memos capture hunches (“pay complaints always seem to follow a comparison with a named competitor”), decisions (“merged rota chaos and shift swaps into scheduling unpredictability, because…”), and emerging interpretations, which are effectively first drafts of what a theme might mean. Memos should be written from day one, because the best analytical ideas arrive during coding and are lost if not captured immediately. They are not overhead: they become the first draft of the findings and a key part of the audit trail.

Quotations as evidence

In the final report, quotations are the evidence: they show the reader the data behind each claim. A good quotation is both vivid — it lands — and representative — it stands for a pattern, not just for itself. The great danger is cherry-picking: choosing the one dramatic quote that supports the analyst’s story while ignoring the ninety pages that complicate it. There are three practical defences. State how many participants the quoted view represents (“seven of twelve leavers…”); quote the counter-voices too, including the participant who disagrees; and never let a single quotation carry a theme alone. A quotation illustrates a finding, but it is the pattern, honestly reported, that is the finding.

Rigour

The standard objection to qualitative analysis is that it is “just your opinion about what people said”. The answer is not to imitate statistics but to make the interpretive process transparent, systematic, and checkable. Four practical disciplines do most of the work, and together they turn “trust me” into “check me” — which is what rigour means in any method.

Intercoder comparison

In intercoder comparison, two or more analysts independently code the same material and then compare their results. Where they agree, confidence rises. Where they disagree, something valuable surfaces: a vague codebook definition that needs tightening, a genuinely ambiguous passage worth discussing, or a difference in interpretive stance that the team must resolve. The process runs: code separately, compare segment by segment, discuss every disagreement, revise the codebook, and re-code. In quantitative content analysis this is formalised as intercoder reliability statistics; in thematic work the structured discussion matters more than any number. Disagreement is not failure — undiscussed disagreement is.

Negative-case analysis

Once a theme starts to form, the temptation is to see it everywhere: confirmation bias works on coders too. Negative-case analysis is the deliberate hunt for data that contradicts the emerging interpretation. Of every theme, the analyst asks: who in the dataset does this not fit — and why? Three outcomes are possible, and all are good. The exception may lead to a refinement of the theme’s boundaries (“this holds for store staff, not warehouse staff”); it may reveal a rival explanation that must now be addressed; or the theme may survive scrutiny, in which case the claim is stronger for having been tested. At Meridian, if the emerging theme is “people leave over pay”, the leaver who took a pay cut to go elsewhere is the most informative transcript in the pile.

The audit trail

The audit trail is the documented record of the whole analytical journey, from raw data through codes and categories to themes and claims. It contains all versions of the codebook, with dates and reasons for changes; the memos recording hunches, decisions, and dead ends; and the mapping from each theme to its supporting extracts. Its purpose is that an outsider could reconstruct — and challenge — every analytical decision, making it the qualitative counterpart of sharing one’s data and code. It also protects the analyst: six weeks later, the question “why did we merge those two codes?” has a written answer. This is the course’s oldest theme once more: systematic and transparent procedure is what separates research from impression.

Description versus interpretation

The final discipline is to keep what the data say and what the analyst concludes visibly separate. “Nine of twelve leavers compared their pay with a named competitor, unprompted” is description. “Pay dissatisfaction at Meridian is driven by external benchmarking rather than absolute pay levels” is interpretation. Both belong in the report, but the reader must always be able to tell which is which. Blurring the two is how weak analyses smuggle opinions in as findings; separating them is how strong analyses invite scrutiny. A practical habit: for every claim written, ask whether a participant could have said it. If yes, it is description; if only the analyst could say it, it is interpretation, and its evidence must be shown.

Software helps, thinking decides

Dedicated packages — MAXQDA, NVivo, Atlas.ti — support qualitative analysis at scale. They store and organise transcripts, attach codes, retrieve every segment under a code instantly, visualise code co-occurrence, and keep the audit trail tidy. What they do not do is decide what a code means, judge whether a theme is grounded, or interpret anything. The software manages the clerical burden; every analytical decision remains human. A useful rule of thumb: an analysis that would be weak on paper will be weak in NVivo — only faster.

Conclusion

Qualitative analysis is the disciplined journey from raw talk to defensible interpretation: a spiral of familiarising, coding, categorising, theming, and interpreting. Codes are the working unit — descriptive or interpretive, generated openly, and governed by a shared codebook. Thematic analysis is the workhorse method, running six steps from immersion to report, while content analysis is its more structured, countable cousin. Quotations are evidence rather than decoration: vivid, representative, and honest about prevalence. Rigour is earned through intercoder comparison, negative-case analysis, the audit trail, and the discipline of separating description from interpretation — the practices that turn “trust me” into “check me”. Software carries the files; the analyst carries the thinking.