VocalFuse is a Fuse Intelligence product.

PODCAST WORKFLOW

How to Transcribe a Podcast Episode

A transcript turns one hour of audio into 8,000-12,000 indexable words, an accessibility asset, and raw material for show notes. This is the repeatable 2026 workflow — get the audio, transcribe it locally, format it, publish it — using VocalFuse (from $5/mo, Pro AI notes $10/mo) or any Whisper-based tool.

The 4-step podcast transcription workflow (2026)

Step 1 — Get the episode audio as a file

Best source, in order: the master recording from your remote studio (Riverside, StreamYard, Zencastr) — per-participant tracks if available, because separate tracks are the difference between roughly 90% and 98% speaker-label accuracy on multi-guest episodes; the mixed master MP3/WAV from your host's dashboard; or a download from the episode's public page. Apple Podcasts and Spotify both generate listener-side transcripts natively (Spotify since August 2023, Apple since January 2024) — useful for accessibility inside the app, but you cannot export them as SRT, they carry no speaker labels, and they are not publishable. Treat them as a baseline, not a workflow.

Step 2 — Transcribe it locally (or pick your meter)

Drop the file into a transcription tool. The local route: VocalFuse runs Whisper on your Windows PC — continuous mode handles hour-long episodes, nothing uploads, and unreleased audio never touches a third-party server. Local Whisper large-v3-turbo runs roughly 5x faster than real time on consumer hardware, so an hour-long episode processes in minutes. The cloud route meters you: Otter's free tier is 300 minutes/month with a 30-minute per-recording cap and only three lifetime file imports on Basic; TurboScribe's free tier is 3 files/day capped at 30 minutes each; Descript meters media hours with AI credits as a second meter. Many paid services run Whisper under the hood — you are paying for their servers and interface, not a better model.

Step 3 — Clean up: speaker labels, timestamps, jargon

Automatic transcripts get words right and structure wrong. Add or fix speaker labels (real names, consistently formatted), insert a [MM:SS] timestamp at each speaker change for scannability, correct product names and niche jargon the model mangled, and break the text into paragraphs of 1-3 sentences. Budget 10-15 minutes for a typical hour-long interview. This cleanup pass is also where you harvest quotable lines for show notes and social clips.

Step 4 — Export in the format each surface needs

One transcript, four surfaces: plain TXT/Markdown for documents and drafting; visible HTML paragraphs on the episode page for Google and AI Overviews (the indexing workflow in detail); SRT or VTT with start-and-end intervals for YouTube captions and player subtitles; JSON with word-level timestamps for programmatic use and chapter markers. Export once, publish everywhere — the same file feeds all four.

Transcript format that works everywhere

Element Convention Why it matters
Speaker labels Real name on its own line before each turn (or "Name:" prefix); one consistent format throughout Scannability, accessibility, and AI answer engines attribute quotes correctly.
Timestamps [MM:SS] at every speaker change for reading transcripts; start-end intervals only for SRT/VTT captions Readers jump to the moment; chapters and clips come free.
Paragraphs 1-3 sentences per block; break long monologues every ~2 minutes of audio Web-readable; search engines parse short blocks better than walls of text.
Cleanup level Verbatim for legal/evidence; light cleanup for publication (drop filler words, keep meaning) Published transcripts are marketing copy — readability beats stenography.
Publishing Full transcript as visible HTML on the episode page + SRT attached for captions Only crawlable text enters Google's index; accordion-only text often doesn't.

Full transcript-format walkthrough: podcast RSS transcript tag & formats.

Why transcribe podcasts locally

No upload of unreleased audio

Episodes are pre-release intellectual property. VocalFuse transcribes on your Windows PC — nothing uploads to a third-party cloud, so embargoed interviews and sponsored segments never touch someone else's servers.

No per-minute metering

Cloud transcription bills by the audio-hour or caps your free minutes; a weekly show burns through both. VocalFuse is flat $5/mo Basic, $10/mo Pro — transcribe every episode plus the entire back catalog with no meter running.

Web-ready output

Continuous mode handles hour-long episodes; Pro adds AI summaries and show-notes drafts. Paste the text straight into your episode template — it is already plain, crawlable HTML-ready text with speaker labels intact.

Deeper reading: the podcast transcription guide, best podcast transcription software (2026), how to get transcripts indexed by Google, auto-generating podcast show notes, adding podcast chapters from a transcript, and the free voice transcription landscape.

Podcast transcription FAQ

How do I transcribe a podcast episode?

Four steps: get the episode audio as a file, run it through a transcription tool, clean up the output, and export it in the format you need. The fastest free route in 2026 is local Whisper transcription — VocalFuse ($5/mo Basic) runs Whisper on your Windows PC, so you drop the MP3 in and get clean text without uploading the episode anywhere.

How long does it take to transcribe an hour-long podcast?

Modern local models transcribe far faster than real time — Whisper large-v3-turbo runs roughly 5x faster than real time on a mid-range GPU and even integrated graphics handle an hour-long episode in minutes. Add 10-15 minutes of cleanup for speaker labels and jargon, and a one-hour episode is publish-ready in well under half an hour.

How do I format a podcast transcript with timestamps and speaker labels?

Add a timestamp at every speaker change in [MM:SS] form, put the speaker name on its own line before their words, and break paragraphs every 1-3 sentences. For captions use SRT or VTT with start-and-end intervals instead; for a web transcript page, plain HTML paragraphs with inline timestamps crawl best.

Is there a free way to transcribe podcasts?

Yes — three of them. Apple Podcasts and Spotify generate transcripts natively for listeners (Spotify since August 2023, Apple since January 2024), but you cannot download them as SRT and they carry no speaker labels. Local Whisper transcription is genuinely free software with no minute cap. Cloud services meter you: Otter gives 300 minutes a month with a 30-minute per-recording cap, TurboScribe gives 3 files a day capped at 30 minutes each.

Should I publish the full transcript on my episode page?

Yes. Google indexes text, not audio — a 60-minute episode is invisible to search until it has words on a crawlable page. Full-text transcripts routinely become the highest-ranking pages on podcast sites for long-tail queries, and AI Overviews quote pages they can parse. Publish the transcript as visible HTML, not an accordion-only fragment.

What is the most accurate way to transcribe a multi-host podcast?

Start from per-participant audio tracks if your remote studio (Riverside, StreamYard, Zencastr) recorded them — separate tracks are the difference between roughly 90% and 98% speaker-label accuracy. Transcribe each track separately or feed them to a tool that supports per-track speaker assignment, then merge. Clean mixed audio with one voice per channel works nearly as well.

Transcribe your next episode in minutes

Subscribe, install the Windows app, drop in the audio, publish the transcript. Flat pricing, local processing, cancel anytime.

Explore related AI note taking guides