Thematic analysis, automated: from transcripts to themes

What thematic analysis is, how an AI coding pass differs from manual coding, and how to audit AI themes against their source quotes in Versive.

Updated
5 min read

Thematic analysis turns qualitative responses, including interview transcripts and open survey answers, into findings someone can act on. Researchers code what each person said, group the codes into recurring themes, and support every theme with evidence. Versive automates the coding pass, analyzing responses as they arrive and grouping the results into themes on the study's Insights tab.

What thematic analysis involves

The method is two passes over the data.

The first pass is coding. You work through each response and tag individual passages with short labels, or codes, that name what is being said: "confused by pricing," "wants an undo button," "compares us to a competitor." A code is deliberately small. One interview might yield a dozen, and repeated codes across the study provide the raw signal.

The second pass is theme-building. You group related codes into patterns that recur across multiple people, rather than across several sentences from one person. A theme is more specific than a topic. "Pricing" is a topic. "Participants expect usage-based pricing and read the flat tier as hiding something" is a theme because it makes a testable claim. A finished theme has a name, a short description, and the specific quotes that support it.

There are two broad ways to run the coding pass. Inductive coding lets the codes emerge from the responses themselves, which suits discovery work where you do not yet know what matters. Deductive coding starts from a predefined set of codes based on the questions you already care about, then tags responses against it. Most applied product research uses a few predetermined buckets while leaving room for unexpected findings.

Where manual coding breaks down

The workload grows with every response. Coding ten interviews carefully can take a day of focused work; coding two hundred can take weeks. Quality can also shift across a large batch. Fatigue and earlier judgments influence the labels assigned later, so the first and last transcripts may be coded to different standards.

Manual coding is also hard to reproduce. Two researchers working from the same transcripts will often produce overlapping but different codebooks, an inter-coder reliability problem. Academic teams address it by double-coding and reconciling disagreements, which adds substantial reading time. Product teams that skip that step accept more dependence on the individual researcher's judgment.

How an AI coding pass differs

An automated pass changes consistency and timing. It applies one coding process across the dataset without fatigue-related drift. Each response is analyzed when the interview ends, so researchers can refresh the study-level analysis during fieldwork and investigate emerging patterns sooner.

In Versive, the output lands on the Insights tab as per-question themes with a summary, sentiment, and respondent count. The count shows the amount of evidence behind a theme before you read its quotes. When the built-in views do not answer your question, a custom insight can turn a prompt such as "which competitors came up unprompted, and in what context" into a saved analysis view that you can rerun as responses accumulate.

An automated pass can miss details that careful reading reveals: hedged answers, sarcasm, or contradictions across an interview. A human coder also develops hypotheses about why a pattern exists while reading. Treat AI coding as a consistent first pass that still requires judgment before it becomes a finding.

The pass can only code what the study collected. An AI question that asks adaptive follow-ups collects several exchanges per participant where a static field collects one, giving the coding pass more material to ground a theme. For a product walkthrough covering per-response summaries, sentiment, charts, exports, and sharing, see How to analyze open-ended survey responses with AI.

How to audit a theme against its quotes

Every theme Versive generates includes supporting quotes. Each quote is tied to its source interview and exact place in the conversation, with timestamps for voice and video. You can open the source directly to check the quote in context.

Five checks, in rough order of value:

A theme that passes these checks is ready to cite. If its quotes do not support the label when read in full, investigate the analysis instead of patching the wording.

Audit themes with your preferred AI agent

The Versive MCP server lets an AI agent work directly with study summaries, per-question insights, statistics, and interview transcripts. Use the built-in study-summary or study-insights prompts for a repeatable first pass, then ask focused questions with custom-insight.

An agent can help test a theme rather than simply restating it:

Ask the agent to cite the source interviews and inspect those passages before using the result in a report. MCP shortens the search and comparison work; it does not remove the need for researcher judgment. See MCP best practices for a recommended broad-to-specific analysis workflow.

When to re-code by hand

Automated coding is most useful at volume. Several situations still justify a manual pass or a manual layer on top of the automated one.

Small samples. With five or ten interviews, the speed advantage is marginal. A careful human read is more likely to catch a hedge, a pause, or an answer that contradicts an earlier one.

High-stakes calls. If a theme is about to shape a roadmap commitment or a pricing change, read the source transcripts yourself before signing off. Regenerate the analysis after late responses so the cited version reflects the full dataset.

Repeated audit failures. If one question's themes keep failing the context check above, stop regenerating and code that question's responses manually. Persistent failures often indicate ambiguous responses that require human judgment.

Vague custom prompts usually call for revision before manual work. Tighten the prompt and rerun it before falling back to hand-coding.

To calibrate the coding pass on your data, choose a theme that you would put in a readout and work through the five checks. The result will show which parts of the automated analysis hold up and where closer review is needed.

Frequently asked questions

What is AI thematic analysis?

Thematic analysis where the coding pass is automated: software tags what each response says and groups the tags into recurring themes, applying one consistent standard across the whole dataset instead of a person reading and labeling every transcript.

What is the difference between a code and a theme?

A code is a short label attached to one passage, such as "confused by pricing." A theme is a broader pattern built from related codes that recur across multiple respondents, with a name, a description, and supporting quotes.

When should I re-code responses manually instead of using AI themes?

With very small samples, before high-stakes decisions where you want to read every transcript yourself, and when a theme repeatedly fails an audit because its quotes turn out to mean something different in full context than the label claims.

Full reference

Results & insights


Keep reading

Faster research, better insights. Start now.