All posts

How to transcribe an interview: methods and examples

Megan JohnsonUpdated August 17, 20268 min
Three star-forming galaxies of the SDSSCGB 10189 group tangled together mid-collision, with fainter, more distant galaxies scattered across the dark background.
Article

Interview Transcription

the fastest transcript is one you never have to make

Hubble · SDSSCGB 10189 · 2023 · NASA

A one-hour research interview takes about four hours to transcribe by hand. Verbit's own testing puts the ratio at roughly four to one, and it climbs higher for messy audio or multiple speakers. Almost nobody running user interviews has four spare hours per session, so transcription becomes the step people rush or quietly skip.

That is a problem, because the transcript is what lets you quote a participant accurately instead of reconstructing what you think they said. It also gives you something concrete to return to when you start comparing themes across interviews.

This guide covers how to transcribe an interview four different ways, which transcription style to pick, how to format and anonymize the result, and when AI-generated transcripts can eliminate manual transcription altogether.

How long does it take to transcribe an interview?

Transcribing an interview by hand takes roughly four hours for every hour of audio. That number climbs when the recording has crosstalk, strong accents, specialized terminology, or poor audio quality. Automated tools produce a first draft in minutes, but you will still spend time fixing names, jargon, and moments where two people talk over each other.

So the real question is not whether to use software. It is how much of the work you want to do yourself, and whether your tool leaves you with a plain document or something you can actually analyze.

Type it out yourself as you listen

The manual method is exactly what it sounds like. You play the recording and type what you hear. Social scientists and field researchers have done it this way for decades.

It is free, and it teaches you something no tool can. When you transcribe your own interview, you hear your own interviewing with fresh ears. Did you ask a leading question without noticing? Did you phrase something in a way that steered the answer? Did you miss an obvious opportunity to follow up? Those habits are obvious on the second listen and invisible in the moment.

Do this once or twice early on to sharpen your technique. It is too slow to be your default, so treat it as a useful training exercise rather than an ongoing workflow.

Use dedicated transcription software

Once you have felt how long manual transcription takes, software is the obvious next step. You upload the recording, the tool returns a draft, and you correct it. Popular options include Rev, Otter, AssemblyAI, Trint, and Sonix.

Accuracy has come a long way, but it is not perfect. Expect to review anything with strong accents, industry jargon, or several people speaking at once, since those are still where automated transcription slips. A quick read-through against the audio catches most errors.

The bigger limitation is what happens next. Most standalone tools give you a transcript, but not a research workflow. You then export it and move it into a separate tool to tag, code, and analyze. For a one-off interview that is fine. Across a whole study, the export-and-reimport shuffle adds up. If you are still assembling your stack, our roundup of user research tools covers how the pieces fit together.

Use a research platform with built-in transcription

A research platform folds transcription into the study itself, so there is nothing to upload or move between systems. When a participant speaks, their words are transcribed automatically, with timestamps you can click to jump straight to a moment in the recording.

Timestamps matter more than they look. If a participant sounds hesitant answering a question, you can jump to that second and listen to the tone rather than guessing from the text. With Great Question you can highlight a passage in the transcript, add a tag like “pricing objection,” and pull that exact quote into your report without leaving the page. Your transcripts also live alongside every other session in a searchable research repository, so a quote from March is still easy to find months later.

Skip transcription entirely with AI-moderated interviews

The fastest transcript is the one you never have to make. With AI-moderated interviews, an AI interviewer runs the session, asks follow-up questions based on what the participant says, and generates the transcript as the conversation happens. The moment the interview ends, the transcript is ready. There is no separate transcription step.

Great Question's agentic moderation launched in 2026. It is designed to make AI-moderated interviews flexible and natural for participants while helping research teams gather more customer insight without adding more moderator hours. It is a different starting point from AI-moderated tools that stop at the recording, and it is worth comparing approaches directly, for example in our Great Question vs Listen Labs breakdown.

Manual and software transcription still matter for interviews you have already recorded elsewhere. But for new research, AI-moderated interviews can remove transcription as a separate task entirely.

Verbatim, clean, intelligent, or summary: which style to use

Whatever method you pick, you are choosing a transcription style, whether you realize it or not. There are four in common use, and the right one depends on how you plan to use the transcript.

  • Verbatim captures everything that is said, including filler words like “umm” and “uhh,” false starts, and pauses or laughter when relevant. Use it when hesitation and tone are part of the data, for example in usability sessions where you are watching for confusion.
  • Clean verbatim keeps what the participant said but removes filler words, stutters, and unnecessary repetition. This is the most common style for user research: readable, but still close to the participant's own words.
  • Intelligent transcription lightly edits grammar and phrasing to make the transcript easier to read. It works well for quotes headed into a public report, but not when exact wording or speech patterns are part of the analysis.
  • Summary transcription condenses the conversation into its main points rather than creating a word-for-word transcript. It is useful for a quick share-out, but not when you need to analyze or verify the original language.

Here is the same sentence in verbatim and clean verbatim, side by side:

  • Verbatim: “Um, I guess... I think the onboarding is quite simple and, umm, there are only a few steps.”
  • Clean verbatim: “I think the onboarding process is quite simple and there are only a few steps.”

The filler that clean verbatim removes is exactly what a behavioral researcher might want to keep. Someone pausing before saying “it was fine” may be telling you something the words alone do not.

How to format an interview transcript

A transcript is only useful if you can navigate it later. Use a consistent format across every interview in a study so you can find quotes, return to the recording, and compare sessions more easily.

  • Label every speaker. Use consistent tags such as “Interviewer” and “P1,” or real names if you are not anonymizing. Put the label at the start of each turn.
  • Add timestamps. Include them at regular intervals or at each speaker change so you can jump back to the original audio or video.
  • Note relevant non-verbal cues. Use brackets for context such as [laughs], [long pause], or [sighs] when it affects how you interpret the response.
  • Flag anything you cannot hear. Mark unclear sections as [inaudible] with a timestamp rather than guessing. An acknowledged gap is more reliable than an inaccurate quote.
  • Keep the format consistent. Use the same speaker labels, timestamp conventions, and notation across every transcript in the study.

Protect participant privacy in your transcripts

Transcripts contain real people saying real things, often about their employer or their own behavior. Treat them as research data that may contain personally identifiable or sensitive information.

Anonymize where you can. Replace names, company names, and other identifying details with labels like “P3” or “[employer],” and keep the key that maps labels to people separate from the transcript itself.

Make sure participants have consented to recording and transcription before the session, and store transcripts somewhere with appropriate access controls. If you are working with participants in the EU or UK, GDPR may apply to recordings and transcripts that contain personal data, making retention, access, and deletion policies important parts of the research process.

What to do after you transcribe

Transcription is the setup, not the payoff. Once you have an accurate transcript, the next step is to code, highlight, and compare responses to identify themes across interviews. That is where a plain document starts to hold you back.

In Great Question, you can tag and highlight passages as you read, then use AI theme clustering to group related findings across sessions. Each theme remains connected to the original transcript and recording, so you can return to the source evidence when reviewing or sharing a finding.

That connection between synthesis and source material matters as studies get larger. Instead of manually moving quotes and notes between transcription, analysis, and reporting tools, researchers can keep the evidence and analysis together. For more on doing this well, see our guide to AI in UX research and our guide to qualitative data analysis.

Transcribe less, learn more

Transcription is worth doing well, but it should not take more time than the research itself. If you are transcribing interviews recorded elsewhere, choose the style that fits your analysis and use software to generate the first draft.

For new research, built-in transcription can remove much of that manual work altogether. Great Question automatically generates transcripts from research sessions and keeps them connected to the recordings, highlights, themes, and analysis that follow. With AI-moderated interviews, the transcript is ready as soon as the session ends, so researchers can move directly from conversation to analysis.

Run your next study with Great Question

Frequently asked questions

How do you write a transcript of an interview?

To write an interview transcript, listen to the recording and document what each person says using consistent speaker labels and timestamps. Choose a transcription style first: verbatim captures filler words and pauses, while clean verbatim removes them for readability. Mark anything unclear as [inaudible] rather than guessing.

How long does it take to transcribe a one-hour interview?

About four hours by hand, longer for multi-speaker or noisy audio. Automated tools return a draft in minutes, though you should still review it for names, jargon, and crosstalk.

What is the best way to transcribe interviews?

For most research interviews, automated transcription is the fastest approach. If you already have a recording, transcription software can generate the first draft, while a research platform with built-in transcription keeps the transcript connected to the session and subsequent analysis. AI-moderated interviews can generate transcripts automatically as new sessions are conducted.

Can AI transcribe interviews accurately?

AI transcription can produce accurate transcripts when the recording is clear, but accuracy varies with audio quality, accents, specialized terminology, and overlapping speech. Review important transcripts against the recording before using exact quotes or making decisions based on the text.

Share