SPEECH TO TEXT SOFTWARE
Speech to Text Software: Every 2026 Option Compared
"Speech to text software" is one search for two different jobs: dictation (your voice becomes text right now, in whatever app you have open) and transcription (a finished recording becomes a document). Every tool below is honest about which one it does — and in 2026 accuracy is no longer the differentiator. Clear-audio word error rates cluster at 5-7% across the field; what actually separates the options is where the audio goes (cloud vs your own machine) and what the free tier really caps.
Verified 2026 pricing and caps · Windows-first · Local/offline options flagged throughout
The three jobs "speech to text software" actually means
Most roundups compare across jobs they don't separate. Pick the job first — the tool follows.
1. Dictation into any app
You talk, text lands at your cursor — email, Slack, VS Code, a browser form. The good ones auto-punctuate and adapt formatting to the app you're in. Tools: Windows Voice Typing, Wispr Flow, Aqua Voice, SuperWhisper, BlabbyAI, VocalFuse.
2. File transcription
You have a recording — interview, lecture, memo, podcast — and need it as text, usually with timestamps or speaker labels. Tools: TurboScribe, Sonix, Rev, open-source Whisper, VocalFuse's file mode.
3. Meeting capture
A bot or a local app records the call, transcribes it, and drafts summaries and action items. Tools: Otter, Fireflies, Fathom, Read AI — or bot-free local capture in VocalFuse Pro.
The 2026 field — prices, caps, and where your audio goes
Every row verified against vendor pricing pages and docs in September 2026. "Processing" is the column most roundups skip — it decides privacy, offline use, and latency.
| Tool | Job | Price | Processing | The cap you'll hit first |
|---|---|---|---|---|
| Windows Voice Typing (Win+H) | Dictation | Free built-in | Cloud (Azure) | Needs internet; no custom vocabulary; accuracy plateaus ~85-90% on fast or technical speech |
| Windows Voice Access | Dictation + PC control | Free built-in | On-device | Windows 11 22H2+ only; ~11 language locales |
| Fluid Dictation (Copilot+ PCs) | Dictation | Free built-in | On-device | English only; needs Copilot+ hardware |
| Word Dictate (M365) | Dictation in Office | Free with M365 | Cloud (Azure) | Word/Outlook only — dies the moment you switch apps |
| Google Docs Voice Typing | Dictation in Docs | Free | Cloud | Chrome + Google Docs only |
| Wispr Flow | Dictation | $15/mo ($12 annual) | Cloud | Free tier: 2,000 words/week desktop (~15 min of speech) |
| Aqua Voice | Dictation | $8/mo annual ($10 monthly) | Cloud only | Free tier: 1,000 words lifetime (~7 min), no offline mode at any tier |
| SuperWhisper | Dictation | Free tier; Pro ~$8/mo | Local on-device | Windows build new (1.0 Nov 2025); local models download first |
| BlabbyAI | Dictation | $8.49/mo, free tier | Whisper v3 Turbo, local | Windows-native; smaller model catalog than Mac-first rivals |
| Otter.ai | Meetings | Free; Pro $16.99/mo ($8.33 annual) | Cloud | Free: 300 min/mo, 30-min conversations, 3 lifetime file imports |
| TurboScribe | Files | Free; Unlimited $10/mo | Cloud | Free: 3 files/day at 30 min on the lower-accuracy Ninja engine |
| Sonix | Files | $10/audio-hour; $25-80/mo plans | Cloud | 30-minute one-time trial; per-hour metering |
| Dragon Professional v16 | Dictation (pro) | $699 one-time | Local | Windows-only; consumer Dragon Home discontinued Feb 2023 |
| VocalFuse | Dictation + files | $5/mo flat (Pro $10 adds AI notes) | Local on your PC | None on the engine — unlimited minutes and files; ~~57 MB model runs offline |
Prices re-verified Sep 2026 from otter.ai/pricing, wisprflow.ai/pricing, superwhisper.com, aquavoice.com, sonix.ai, dragon resellers (CDW $685.99), turboscribe.ai. Vendor prices move — re-verify before buying.
Cloud vs local: the axis that actually matters in 2026
The cloud tier
Win+H, Word Dictate, Wispr Flow, Aqua Voice, Otter, TurboScribe, Sonix all send your audio to vendor servers. Accuracy is excellent and setup is instant — but every word transits someone else's datacenter, none of it works offline, and most meter you in minutes, words, or files.
Microsoft's own docs are explicit that voice typing "uses online speech recognition, which is powered by Azure Speech services" and requires an internet connection. On a plane or a locked-down client network, the free built-in simply stops.
The local tier
Windows Voice Access (on-device, offline, 11 locales), Fluid Dictation (Copilot+), SuperWhisper, BlabbyAI, Dragon, and Whisper-class engines. Published Open ASR Leaderboard numbers put the best local engines at 6.34% WER (NVIDIA Parakeet TDT 0.6B v3) vs 7.44% for Whisper large-v3 — and independent 2026 testing found a local model beating every cloud dictation service on a blended accuracy-plus-speed benchmark (SuperWhisper S1 Voice: 84/100 vs Wispr Flow's raw 75).
Local means no upload, no meter, no network dependency — and on clean English audio, parity with the cloud. The trade: a model download on first run and language coverage bounded by the model you ship.
The accuracy read, honestly: on clean audio, every serious engine lands within a few points — Parakeet 6.34% vs Whisper large-v3 7.44% WER is roughly one word in a hundred. Vendor benchmarks are self-selected (SuperWhisper's own board ranks its model first; Aqua cites its own AISpeak set), and independent benchmarkers warn that clean-studio numbers don't predict real-world calls with accents and crosstalk. Test on your own audio before committing — which is exactly why free local tiers are the lowest-risk way to try.
The dictation-subscription wave (and its caps)
A 2024-2026 product wave re-invented dictation as a subscription. The pitches are real — context-aware formatting, filler-word removal, command mode — but read the free tiers before you build a habit on one:
Wispr Flow: 2,000 words/week
The free desktop cap is about 400 words a working day — one long email plus a couple of Slack replies, by the vendor-adjacent analyses' own math. Pro is $15/mo ($12 annual, $144/yr). Cloud-only; no offline tier exists.
Aqua Voice: 1,000 words lifetime
Roughly 7-8 minutes of speech — an evaluation window, not a tier. $8/mo billed annually, $10 monthly, no annual discount on the annual-vs-monthly axis you'd expect. Cloud-only at every tier; Privacy Mode is opt-in, not default.
SuperWhisper: free tier + $8/mo
The local-first pick — on-device models, SOC 2 Type II, HIPAA-compliant, 100+ languages, and a Windows 1.0 that shipped late November 2025 (still maturing vs its Mac app). Its free tier doesn't expire.
The one-time legacy: Dragon
Dragon Professional v16 is $699 one-time (resellers currently list $685.99), Windows-only, with up to 99% recognition accuracy after vocabulary training and a 25-year track record in legal and medical dictation. What's gone: the consumer tier. Dragon Home was discontinued in February 2023 with no replacement, v15 is out of support with documented Windows 11 24H2 compatibility problems, and Dragon Anywhere (mobile) was discontinued July 1, 2026. If you need HIPAA-grade front-end dictation with custom medical/legal vocabularies and voice-trained accuracy, Dragon remains the answer. If you don't, you're paying a 25-year-old pricing model for capabilities local Whisper-class apps now match at a fraction of the price.
How to choose by job
"I just want to dictate occasionally"
Use the built-ins: Win+H online, Voice Access offline (Win 11 22H2+), Fluid Dictation if you have a Copilot+ PC. Free, zero setup, and good enough for casual volume.
When Win+H isn't enough →"I dictate all day, into every app"
A dedicated dictation app: Wispr Flow or Aqua if cloud is fine and you'll pay monthly; SuperWhisper or BlabbyAI if you want local. Watch the free-tier caps — they're sized to demo, not to live in.
The Windows dictation field →"I have files to transcribe"
File transcription is a different product: TurboScribe/Sonix/Rev meter by file, day, or hour in the cloud; local engines (Whisper, or VocalFuse's file mode) batch without a meter.
Free file-transcription caps compared →"I need meeting notes"
Meeting tools (Otter, Fireflies, Fathom) bring bots, minute caps, and cloud storage. A local alternative captures and drafts notes without a bot joining the call.
Local AI note taking →The VocalFuse take: one engine, no meter, your hardware
Most "speech to text software" splits into two products — a dictation app and a transcription service — each with its own subscription and its own cap. VocalFuse collapses both jobs into one Windows app running a local Whisper-class engine (~~57 MB):
- ✓ Dictation into any app: hold the always-on-top pill, speak, release — punctuated text lands at your cursor via the clipboard, in browsers, Slack, Word, VS Code, anything with focus.
- ✓ File transcription: drop in audio or video — no upload queue, no per-minute billing, no per-file cap. The engine runs on your PC.
- ✓ Works offline after setup: the ~~57 MB model downloads once; after that, dictation and transcription run with the network cable pulled.
- ✓ $5/mo flat for unlimited local dictation and transcription (Pro $10/mo adds AI note taking: speaker labels, summaries, meeting capture). No per-seat math, no minute pools, no engine downgrade on the entry tier.
- ✓ Private by architecture: audio never leaves your machine for dictation or files — there is no cloud step to opt out of.
Related reading
Speech to text, product page
The product-level view: hold-to-speak dictation in every Windows app, offline.
Local speech to text →Every free tier's real cap
Otter 300 min, TurboScribe 3/day, Notta 3-min cap, trials that never return — compared in one table.
Free transcription compared →Offline dictation, ranked
Architecturally-offline options vs "offline recording, cloud transcription" lookalikes.
Best offline dictation →Whisper locally, WER numbers
Whisper large-v3 7.44% vs Parakeet 6.34% on real benchmarks — and what local hardware you need.
Whisper local on Windows →Explore related AI note taking guides
Speech to text software — FAQ
What is the best speech to text software in 2026?
It depends on the job. For dictation into any app, Wispr Flow ($15/mo) and Aqua Voice ($8/mo annual) lead the cloud tier while SuperWhisper and BlabbyAI lead the local tier; for file transcription, TurboScribe ($10/mo unlimited) and Sonix ($10/audio-hour) are the cloud defaults and open-source Whisper is the free local route; for meetings, Otter ($16.99/mo Pro) dominates the cloud tier. Accuracy no longer separates them — clear-audio WER clusters at 5-7% across the field — so choose on processing location (cloud vs your own machine) and what the free tier caps. VocalFuse covers dictation plus files locally on Windows from $5/mo flat with no meter.
Is there free speech to text software with no limits?
Yes, locally. OpenAI Whisper and whisper.cpp run on your own hardware with no quota, no upload, and no cost — the catch is a command-line setup and model downloads. Windows Voice Access is a free on-device middle ground (offline, but Windows 11 22H2+ and ~11 language locales). Everything cloud-based meters something: Win+H requires internet, Wispr Flow caps its free tier at 2,000 words/week, Aqua Voice at 1,000 words lifetime, Otter at 300 minutes/month with 30-minute conversations, TurboScribe at 3 files/day on a lower-accuracy engine. VocalFuse is the turnkey local option — unlimited minutes and files at $5/mo.
Which speech to text software works offline?
Windows Voice Access (on-device recognition, works offline after a one-time model download, Windows 11 22H2+), Fluid Dictation on Copilot+ PCs, SuperWhisper (on-device Whisper/Parakeet models), BlabbyAI (local Whisper v3 Turbo), Dragon Professional (local, $699), and open-source Whisper front ends. VocalFuse runs a ~57 MB local Whisper-class model that works offline after first launch — dictation into any app plus file transcription with nothing uploading. The offline trap to avoid: several "offline-sounding" tools record locally but transcribe in the cloud, so they still fail without internet.
How accurate is speech to text software?
On clean audio, the best engines cluster within a few points: NVIDIA Parakeet TDT 0.6B v3 posts 6.34% WER and Whisper large-v3 7.44% on the Open ASR Leaderboard — roughly one word in a hundred apart. Microsoft claims 3.8% average WER across 25 languages for its new MAI-Transcribe-1 model (April 2026), and proprietary models like Aqua's Avalon claim ~97% on their own AI-vocabulary benchmarks. Real-world accuracy drops on accents, crosstalk, and noise for every engine, and vendor benchmarks are self-selected — independent testers warn clean-studio numbers don't predict real call performance. Test on your own audio.
Is Windows Voice Typing (Win+H) good enough?
For occasional dictation, yes — it's free, zero-setup, and works in any text field. Its real limits: it requires internet (Microsoft's docs confirm it uses online speech recognition powered by Azure Speech services), accuracy plateaus around 85-90% on fast or technical speech, there is no custom vocabulary, and long pauses end the session. For offline work use Voice Access (Win 11 22H2+, on-device, also handles hands-free PC control); for all-day dictation with custom vocabulary and AI cleanup, a dedicated app wins. Don't confuse the two — Win+H is cloud dictation, Voice Access is on-device control.
What did Dragon cost and is it still worth it?
Dragon Professional v16 is $699 one-time (resellers list it around $685.99), Windows-only, with up to 99% accuracy after vocabulary training — still the standard for legal and medical front-end dictation. The consumer tier is gone: Dragon Home was discontinued in February 2023 with no replacement, v15 is out of support with documented Windows 11 24H2 problems, and Dragon Anywhere mobile was discontinued July 1, 2026. Worth it only if you need voice-trained custom vocabularies and compliance-grade local dictation; otherwise local Whisper-class apps match the core capability at a fraction of the price.
Dictation vs transcription software — what's the difference?
Dictation converts live speech to text as you speak, into the app you have open — single speaker, real time, latency matters. Transcription converts a finished recording into a document, usually with timestamps or speaker labels, where batch speed and accuracy on multi-voice audio matter. Wispr Flow, Aqua, SuperWhisper, and VocalFuse's hold-to-speak mode are dictation tools; TurboScribe, Sonix, Rev, and file mode are transcription tools; Otter and Fireflies are meeting-capture tools. Most roundups blur these three jobs — the tool you need depends on which one you're actually doing.
Why does local processing matter for speech to text?
Three reasons: privacy (audio never leaves your machine — no vendor retention policy to trust or opt out of), availability (no internet means no dictation with cloud tools; local keeps working on planes and locked-down networks), and economics (no per-minute or per-word meter — the engine is yours). Independent 2026 testing found local Whisper-class models already match or beat cloud dictation services on blended accuracy-plus-speed benchmarks, so local is no longer the accuracy sacrifice it was in 2023. The costs are a one-time model download (~57 MB for VocalFuse's engine) and bounded language coverage.
Pick the architecture, not the feature grid
Cloud tools meter you by minutes, words, or files and route every syllable through a datacenter. Local tools ask for a model download and then get out of your way. VocalFuse is the turnkey local option on Windows — dictation into any app plus file transcription, from $5/mo flat.
Building agents too? VibeFuse is the first free widget-based AI harness, with an open marketplace where creators earn 80% on widgets, skills, and styling packs — and its free voice-transcription dock gives developers local STT at $0.