Back to blog
8 min readby Mathias

Intercoder Reliability When AI Does the First Pass: Three Defensible Approaches

Running the model twice and computing kappa measures determinism, not codebook clarity. Three methods that actually hold up in review, and a worked paragraph for your methods section.

Intercoder ReliabilityMethodologyCodingResearch Rigour

"How do you establish intercoder reliability if a language model did the coding?" This is the question that comes up in supervision meetings and peer review, and it deserves a better answer than the two that usually get offered — either "you don't need it any more" or a reliability coefficient computed between a model and itself.

Both are wrong, and for the same reason: they treat the model as if it were a second coder. It is not.

What intercoder reliability was ever for

Two independent coders apply the same codebook to the same material. Agreement is measured — Cohen's kappa for two coders, Krippendorff's alpha for more or for non-nominal data. High agreement is taken as evidence that the codebook is clear enough to be applied consistently rather than reflecting one person's idiosyncratic reading.

Read that again, because the important word is codebook. The measure was never really about the coders. It was always a test of whether the category definitions were unambiguous.

Worth noting: within reflexive thematic analysis, Braun and Clarke explicitly reject reliability coefficients as a quality criterion, since their framework does not treat coding as a process of measuring an objective reality. If you are working in that tradition, the honest answer to a reviewer is to say so and to demonstrate rigour differently. This article is aimed at content-analytic traditions — Mayring, Krippendorff, Schreier — where the criterion does apply.

Why running the model twice is not a reliability test

A tempting shortcut: code the material twice with the model and compute kappa between the runs. The number will look excellent, especially at a low temperature setting.

It is meaningless. You have measured the model's determinism, not the clarity of your categories. Two runs share the same systematic blind spots — if a definition is ambiguous in a way this model resolves consistently in one direction, both runs make the same mistake and agree perfectly about it. High agreement here is compatible with a thoroughly bad codebook.

Do not report it. A reviewer who understands the measure will spot it immediately, and it undermines everything else in your methods section.

Three defensible approaches

1. Human verification on a sample

The most straightforward. Draw a random sample of coded segments — 10–20% is a common range, and state how you chose it — and have a human coder independently apply the codebook to that sample, blind to the machine assignments. Compute agreement between the human and the machine coding.

This tests exactly what the original method tested: whether the codebook is clear enough that an independent application produces the same result. That the second coder happens to be software does not weaken the logic, as long as the human coded blind.

2. Human first pass, machine extension

Code a subset entirely by hand, develop the codebook from it, establish reliability between two humans in the traditional way, and only then apply the finished codebook to the remaining material automatically.

Methodologically the cleanest option, because reliability is established before any automation touches the data. It is also the most work, and it is the right choice when the stakes are high.

3. Full-coverage audit

Rather than sampling, check every assignment — feasible when each coding carries a verbatim extract, because verification becomes reading rather than re-coding. Report the proportion you corrected.

This is not intercoder reliability and should not be labelled as such. It is a systematic verification procedure, and reported honestly it is often more convincing than a coefficient, because it describes what you actually did.

The thing that makes all three possible

Every one of these depends on the same property: each coding decision must carry the exact verbatim extract that produced it.

Without that, verification means re-reading the entire transcript to find out why a segment was coded a certain way, and nobody does that at scale. With it, checking a coding takes seconds. This is also the property that makes a machine-assisted analysis auditable at all — and it is why we validate that every extract is a genuine substring of the source text and discard any that is not, rather than trusting the model to quote accurately.

Reporting it

A worked example for the methods section:

"An initial codebook was generated computationally from the full dataset and then reviewed and revised by the author, who merged overlapping categories and sharpened definitions. The revised codebook was applied to all 312 units, with each assignment accompanied by a verbatim extract. A random sample of 20% (n = 62) was independently coded by a second researcher, blind to the computational assignments, yielding Cohen's kappa = 0.79. Disagreements were resolved by discussion and led to the clarification of two category definitions, after which the full dataset was recoded."

That paragraph is defensible because it is specific: it says what was automated, what was human, how verification was done, and what changed as a result.

In short

Automation does not remove the need for reliability — it relocates it. The question shifts from "do two people agree?" to "is this codebook clear, and did anyone independently check the output?" Answer both explicitly and the method holds.

See how coded output with verbatim extracts looks, or read our walkthrough of the six phases of thematic analysis.

Try Themera on your own data

3 analyses free. No credit card. EU-hosted and GDPR-native.

Start free