VocalFuse is a Fuse Intelligence product.

INTERVIEW TRANSCRIPTION

Interview Transcription Software & Format Guide (2026)

Every serious interview tool charges a meter — Rev at $0.25/min AI or $1.50–$1.99/min human, Sonix at $10/audio-hour, Otter Pro cut from 6,000 to 1,200 minutes at the same price. The free path that never meters: record the interview on your PC and transcribe it locally with VocalFuse — speaker-labeled, offline, flat $5/mo for unlimited hours.

What interview transcription actually costs in 2026

Interview work is uploaded-file work — recordings of candidate screens, research sessions, and source calls that already happened. The tool you pick is a meter question more than a quality question: every option below is accurate on clear audio (85–95%); the differences show up on your invoice.

Tool Pricing (verified Sep 2026) 30 interviews × 1 hr The catch
VocalFuse $0 free tier; Pro $5/mo flat $5 Windows-only, local processing — audio never leaves your PC
Rev AI $0.25/min; human $1.50–$1.99/min $450 AI / $2,700–$3,600 human Per-minute pricing punishes scale; AI drops on crosstalk, pushing you to the human meter
Otter.ai Pro $16.99/mo ($8.33 annual); free 300 min/mo $16.99 — then throttled Pro minutes cut from 6,000 to 1,200/mo without a price cut; 10 file imports/mo cap; bot joins the call
Sonix $10/audio-hour pay-as-you-go $300 Metered forever — heavy research months cost the most when you can least afford it
Descript Hobbyist $24/mo (10 hrs); Creator $35/mo (30 hrs) $24–$35, then credit burn AI credits system (late 2025) meters filler removal and Studio Sound separately from hours
Trint Starter ~$52/mo (7 files); Advanced ~$60/mo $52–$60 7-file cap on entry plan is brutal for a 30-interview study
Temi ~$0.25/min ($15/hr) $450 Rev's AI meter at lower accuracy — the worst of both

The meter column is the whole argument. A research team transcribing 30 hours of interviews pays Rev $450 in AI minutes or up to $4,500 human-reviewed; Sonix $300; Otter Pro caps at 20 hours before throttling; Descript's credits stop when the allocation does. VocalFuse is flat $5/mo — unlimited hours, no per-minute meter, no AI credits, no upload queue, because transcription runs on your own machine.

The transcript format researchers actually need

Format rules matter more than tool choice once you're coding transcripts. Rule one — and it's the #1 failure researchers hit: one speaker per paragraph, always. Never merge two speakers into one block, even for a one-word answer — it makes the transcript unusable for quotation. Second: consistent labels with a colon (INTERVIEWER: / PARTICIPANT: or I: / P1:), decided once per study. Third: bracketed tags for anything that isn't speech — [inaudible 00:14:22] with the timestamp inside the tag, [crosstalk], [unclear: Kessler?] — mark uncertainty rather than guessing a confident wrong word.

Header block (copy into every transcript)

STUDY ID: [STU-___]
PARTICIPANT ID: [P__]
DATE (YYYY-MM-DD): [____-__-__]
INTERVIEWER: [name/initials]
SESSION TYPE: [in-person / phone / video]
CONSENT: [obtained — method]

De-identified participant IDs keep files shareable; the consent line is your audit trail. The free interview transcript template (.docx) ships all of this pre-formatted.

Speaker labels that survive automation

Auto speaker labels are only as good as their per-turn accuracy — tests put AI speaker ID at 82–94%, with overlapping speakers the failure case. VocalFuse's labels are applied per turn and editable after, so the one-speaker-per-paragraph rule holds by default and a 30-second fix pass restores 100% attribution.

Line numbers for citation

Qualitative work quotes as "P4, lines 212–218" — references that are meaningless without stable numbering. Add line numbers once the transcript is final; renumbering after every correction is how quotations and line numbers drift apart. Export the corrected file, then number.

Verbatim vs clean verbatim — pick once

Verbatim (every filler, false start, stutter) when how something was said matters — legal quotes, discourse analysis. Clean verbatim (filler removed, grammar lightly fixed) for thematic analysis, journalism, podcasts. Decide once per study; mixing styles across participants breaks comparability. Cleaning filler removes ~40 tokens per ~4,500 words (TranscribeNext, measured across 934 two-speaker recordings) — so verbatim overhead is small, and clean verbatim is the default for most research.

From recording to coded data in four steps

1. Record the interview on your own machine

A local recorder captures both sides of a call or an in-person session straight to disk — no bot joins the call, no "this meeting is being recorded" announcement for a sensitive source, no cloud round-trip. For remote interviews, record system audio + microphone. The recording stays on your machine from the first second.

2. Transcribe locally with speaker labels

Drop the file into VocalFuse. A Whisper-class model runs on your PC — no upload, no queue, no per-minute charge while you wait. Speaker turns are labeled automatically and stay editable. A 60-minute interview typically finishes in minutes, and the audio never leaves your disk.

3. Clean verbatim pass (optional, 5 minutes)

If your study uses clean verbatim, strip fillers and false starts now — locally, in any editor. Because the transcript is plain text you own, there's no credit meter charging you for the cleanup pass the way Descript's AI credits do.

4. Export plain text, number lines, code

TXT/DOCX exports are the standard every QDA tool accepts (NVivo, ATLAS.ti, MAXQDA), so your transcript survives tool switches and a 5-year archive. Final pass, then line numbers, then coding. Start from the free interview transcript template — header block, tag legend, and a QA checklist pre-built.

Interview transcription questions, answered

What is the best interview transcription software?

For uploaded interview recordings, the honest 2026 answer is a pricing decision: Rev is the accuracy gold standard but meters AI at $0.25/minute and human review at $1.50–$1.99/minute; Sonix charges $10 per audio hour; Otter Pro cut its plan from 6,000 to 1,200 monthly minutes without cutting the price. VocalFuse transcribes interviews locally on your Windows PC for a flat $5/mo with unlimited hours — the only option whose bill never scales with your study size.

How much does it cost to transcribe a 1-hour interview?

Rev AI: $15/hour ($0.25/min). Rev human: $90–$120/hour. Sonix: $10/hour. Temi: ~$15/hour. Otter: included in the $16.99/mo Pro plan until the 1,200-minute monthly cap. Descript: included in $24/mo until the 10 transcription hours run out, then credit charges. A local tool like VocalFuse costs $0 on its free tier and a flat $5/mo on Pro regardless of hours.

How do I transcribe an interview for free?

Record the interview on your PC (local recorder or system-audio capture), then transcribe the file with VocalFuse on your Windows machine. The free tier handles short sessions with no account upload step; the recording and transcript never leave your disk. This works for interviews of any age — drop in an existing recording from months ago and it transcribes the same.

Is it legal to record and transcribe an interview?

In the US, federal law requires one-party consent; some states (California, Florida, Illinois and about 8 others) require all-party consent to record a private conversation. For research interviews, get written consent regardless — your consent form is both the legal cover and the audit trail cited in your transcript header. Journalism has additional source-protection considerations; see our transcription-for-journalists guide.

What format should an interview transcript use?

Header (study ID, participant ID, date, interviewer, session type, consent note), then labeled dialogue: one speaker per paragraph with consistent INTERVIEWER:/PARTICIPANT: labels, bracketed tags for non-speech like [crosstalk] and [inaudible 00:14:22], and line numbers added last (after final edits) so citations like "P4, lines 212–218" stay stable. Pick verbatim or clean verbatim once per study and don't mix.

How accurate is AI interview transcription?

Independent 2026 tests put AI tools at roughly 5–9% word error rate on clear one-on-one audio (85–95% accuracy), degrading on crosstalk and heavy accents; AI speaker-ID lands at 82–94%. Human transcription (Rev) is the only 99%+ path at $90–$120/hour. For most research and recruiting, an AI transcript plus a short human fix pass gets you publication-grade at a fraction of the cost.

Can interview transcripts be kept private and off the cloud?

Yes — that's local transcription. VocalFuse runs a Whisper-class model entirely on your Windows PC: no upload, no cloud queue, no vendor training on your recordings. For participant data governed by IRB protocols, NDA'd source calls, or HR interviews with personal data, the audio and transcript never leave the machine — the differentiator cloud meters can't offer.

Is there a free interview transcript template I can download?

Yes — the .docx on this page is free with no email gate. It includes the header block (study ID, participant ID, date, consent note), the standardized bracket-tag legend, a verbatim and clean-verbatim worked example, citation and anonymization rules, and a pre-submission QA checklist. It opens in Word 2016+, Word on the web, and Google Docs, and works as a coding source for NVivo, ATLAS.ti, and MAXQDA. The page above also gives the copy-paste plain-text version.

Related guides

The privacy rules matter most for source-protecting journalists — see transcription for journalists. For the meeting-side of the same engine, see AI note taker.

Transcribe interviews without the meter

VocalFuse runs a Whisper-class model on your Windows PC: speaker-labeled interview transcripts, fully offline, never uploaded. Free tier to start, Pro is flat $5/mo for unlimited hours — vs Rev's $450 for the same volume.