Most guides to coding qualitative data explain what a code is and then leave you alone with four hundred pages of transcript. This one is about the part that actually goes wrong: not the definition, but the doing.
A code is a label, not a summary
The single most common beginner error is writing codes that summarise rather than categorise. Participant describes difficulties with the new rota system is a summary — it applies to exactly one extract and can never recur. Rota changes disrupt handover is a code: it names a pattern, and it can apply across participants.
The test: could this label apply to something a different participant said? If not, you have written a note, not a code.
Inductive or deductive — decide before you start
Inductive means codes come out of the data. You read, you notice, you name. Right when the point is to discover what is there.
Deductive means codes come from theory or a prior framework, and you apply them. Right when you are testing whether an existing model describes your setting.
Most real projects are both: a deductive skeleton from the literature, with inductive codes for what does not fit. That is fine — but say which parts came from where, because a reader cannot tell by looking.
The mechanics, in the order that works
1. Read everything first. Do not code yet.
Coding while reading for the first time produces codes shaped by the first two transcripts. Read all of it, make rough notes, then start.
2. Code openly and generously on a subset
Take a third of your material and code it without restraint. Forty or fifty codes at this stage is normal and healthy. Twenty is a warning sign — it means you have jumped ahead to themes.
Code everything interesting, not only what answers your question. The material you skip because it seems off-topic is exactly the material that would have complicated your account, and a complicated account is a convincing one.
3. Build the codebook
Now consolidate. Merge duplicates, sharpen boundaries, and write for each code:
- A name of two to five words
- A one-sentence definition a second coder could apply
- An anchor example — a verbatim extract that clearly belongs
- Where the boundary is easily confused, a counter-example
The anchor example is the part people skip and the part that makes the codebook usable. A definition without one is an opinion.
4. Apply the codebook to everything, consistently
This is the long, mechanical stretch — and where drift happens. By transcript nine you are applying codes slightly differently than at transcript one, without noticing. Two defences: recode your first few transcripts at the end, and never change a definition without going back over what you already coded under it.
5. Check your coverage
The question almost nobody asks: how much of your material received no code at all?
A little is fine — greetings, small talk. A lot means your codebook does not describe your data. Reporting coverage is a straightforward way to show a reader that nothing quietly disappeared, and uncoded material is often where the interesting thing is hiding.
Six practical rules
- Always keep the verbatim extract with the code. A code without its evidence cannot be checked, by you or anyone else. This is the single habit that separates auditable analysis from assertion.
- One extract can carry several codes. Forcing one code per segment throws away most of the meaning.
- Keep minority views. A code applied twice may be your most interesting finding. Frequency is not importance.
- Date your codebook versions. When you write the methods section you will want to describe how it evolved.
- Do not code the interviewer. Your own questions are not data about participants.
- Stop when new material stops producing new codes — and say so, with the numbers.
Where software helps, and where it does not
Applying a fixed codebook consistently across four hundred segments is mechanical work that does not benefit from your judgement — it only suffers from your fatigue. That part automates well, on one condition: every assignment must come with the verbatim extract that produced it, so you can check it rather than trust it.
Deciding what is meaningful does not automate, and a tool that claims otherwise is selling you a problem. Nor does the write-up: a results section that is eighty percent block quotes has transcribed rather than analysed.
If you do use software, say so in your methods section, and verify a sample independently. That turns "an AI helped" from a weakness into a documented procedure.
See a worked example with codebook, anchor examples and coverage — no account needed.