Waver vs built-in dictation
macOS, Windows, iOS and Android all turn speech into text, for free, and they are genuinely good at it. What none of them do is edit. They write down the filler, the false start, the hedge, and both halves of a correction — faithfully, in the order you said them — and hand you a paragraph to clean up before you dare send it.
This page is about that gap. It is one gap, it has a mechanical cause, and the whole argument fits in a single sentence taken apart word by word.
One sentence, three outcomes
Nineteen words, spoken into a message to a colleague. Nothing unusual about them — no accent to blame, no background noise, no technical vocabulary. Just the way people talk when they are thinking and speaking at the same time.
so um can you tell the team the launch is gonna slip to like not friday the following monday
So, um, can you tell the team the launch is gonna slip to, like, not Friday, the following Monday.
Punctuation placement is the one thing that differs between the four platforms — some auto-punctuate, some wait for you to say "comma". The words are identical on all of them, because faithful transcription is the specification. Note that both dates survive, in the order you said them.
Can you tell the team the launch is slipping to the following Monday, not Friday?
Same recording · Turn into a reply
Heads-up — the launch is slipping. It won't make Friday; we're now looking at the following Monday.
One recording, two shapes, no second dictation. And if you want the raw version back — every "um", both dates, in the order you said them — "what I actually said" returns it untouched.
| You said | What it is | Built-in | Waver |
|---|---|---|---|
| so um | Discourse marker, then a filler | Both written out | Both dropped |
| gonna | Casual contraction | Kept as spoken | Resolved into the verb |
| like | Hedge, not a comparison | Written out, usually with commas | Dropped |
| not friday | The half you are cancelling | Kept, in spoken order | Kept, demoted to a trailing contrast |
| the following monday | The half you actually mean | Kept, in spoken order | Promoted into the main clause |
The last two rows are the interesting ones. "Not Friday, the following Monday" is a correction, and a correction has a winner. Built-in dictation preserves both candidates and leaves the arbitration to whoever reads the message — which, on a bad day, is how a team ships on the wrong date.
The gap, itemised
None of these are bugs. Built-in dictation does exactly what it says: it transcribes. Every item below is a thing it was never trying to do, which is precisely why it has not improved in the years you have been using it.
Built-in
Writes both halves, in the order you said them
Waver
Resolves which half survives
Nobody speaks in final drafts. You say a date, hear it, and fix it in the same breath — "not Friday, the following Monday", "five pm, actually six". Built-in dictation has already typed "5pm" into your message before it hears "actually six", so you get both, and the reader has to work out which one is real. It is the single most common reason a dictated message needs editing before you dare send it.
Built-in
Transcribed faithfully
Waver
Auto Cleanup, four levels
"Um", "like", "you know", "I mean", the sentence you started twice. Built-in dictation is doing its job correctly when it writes these down — faithful transcription is the specification. But nobody wants to read their own disfluencies, so the tidying lands on you. Waver's Auto Cleanup has four settings (None, Light, Medium, High) so you choose how much of your voice survives, and "what I actually said" restores the raw transcript whenever you want it back.
Built-in
One long paragraph, unless you narrate the formatting
Waver
Paragraphs and lists from a flat ramble
Every built-in dictation engine will insert a break if you say "new paragraph" out loud. That works, and it also means holding the document's shape in your head while you are trying to hold the argument in your head. Three minutes of talking through four options arrives as one continuous block. Waver reads the whole thing after you stop and gives it the shape the content already had.
Built-in
No editable dictation vocabulary
Waver
Add, pin, and auto-learn from your corrections
This is the one people underestimate, and it has its own section below. The short version: your colleague's name, your product, your internal service. Built-in dictation gives you no list to add them to. Waver does — and the list is sent with the audio, so it changes what the recogniser hears rather than find-and-replacing afterwards.
Built-in
Per-device, per-OS, not portable
Waver
On your account, everywhere you sign in
Whatever your Mac has learned about how you speak, your Android phone has not. Nor has the work laptop. There is no built-in path for teaching four separate operating systems the same forty proper nouns, so in practice people teach none of them. Waver's dictionary rides your account: add a term on the phone, and the browser, the desktop app and the tablet all have it on the next recording.
Built-in
Not a dictation feature
Waver
Transforms, including ones you write
The same two minutes of talking is a Slack message, a status update, or a note to yourself, depending on who is reading it. Waver ships Polish, Prompt engineer, Make it shorter and Turn into a reply, plus custom transforms you define once and reuse. Built-in dictation has one output: the words, as spoken. Anything else is your problem, in another app.
Built-in
None
Waver
Snippets and spoken emoji
Say a cue and a Snippet expands into the block you set up — a sign-off, an address, a standing disclaimer you retype forty times a month. Say "fire emoji" and you get the flame rather than the words. Small things, but they are the difference between dictation as a novelty and dictation as the way you actually write.
Why the gap exists
Built-in dictation is a live streaming transcriber. It commits words into your text field while you are still speaking, because that is the product: the cursor keeps up with your mouth. Once a word is in the document, it has been typed. It cannot be un-said.
Now look at what fixing a self-correction requires. To know that "Friday" was cancelled, you have to have heard "the following Monday" — information that arrives seconds after"Friday" was already committed to the screen. A streaming transcriber is being asked to revise text it has already handed over. That is not a hard engineering problem; it is the wrong shape of problem for the architecture.
The same constraint explains the rest of the list. Filler removal needs to know the sentence continued. Paragraph structure needs to know where the argument turned. Tone needs the whole thing. Every one of them requires an ending that a live transcriber does not have yet.
Built-in
Hears a word → commits it → hears the next word. The document is written forwards and never revisited.
Waver
Holds the utterance → transcribes it whole, with your dictionary → then decides what the beginning meant, knowing how it ended.
The trade is stated plainly further down, and it is a real one: you do not watch words appear as you talk. You get them when you let go of the key.
Before vs after
This is the difference people find hardest to believe until they see it, so it is worth being precise about the mechanism rather than the marketing.
Text replacement — the feature your desktop OS has had for a decade — is a string substitution. It waits for an exact run of characters and swaps it for another. That works when you can predict the mistake. You cannot: a name the recogniser has never encountered comes back as a phonetic guess, and next time it comes back as a different phonetic guess, because the audio was slightly different. There is no string to key on.
Waver sends your Personal Dictionary with the audio, before anything is recognised. The recogniser is already leaning toward your terms while it decides what it heard, so the guess never happens. The dictionary is not a spellchecker bolted on afterwards — it is an input.
Add
Type in the terms you know it will trip on — product names, colleagues, the internal service nobody outside the building has heard of.
Pin
Mark the ones that must never be dropped, however long the list gets.
Auto-learn
Fix a word once and Waver notices. Corrections feed the list without you managing it.
And it is one list. Built-in dictation is four separate products that happen to share a name — nothing your Mac has learned is available to your phone, and there is no path between them. Waver's dictionary lives on your account, so the term you add on the train is in effect on the desktop app before you sit down.
Side by side
"Built-in dictation" means whichever of these four you have in front of you. They differ in the details; on everything in the table below, they behave alike.
macOS
Dictation, in Keyboard settings
Windows
Voice typing, Win + H
iOS / iPadOS
The microphone on the keyboard
Android
Gboard voice typing
| Capability | Built-in dictation | Waver |
|---|---|---|
| Turns speech into text | Yes | Yes |
| Costs nothing to start | Yes, unlimited | Free tier: unlimited transcription, 10 AI actions a month, no card, ad-free |
| Works in any text field, system-wide | Yes | Yes on the desktop build; in-app Flow Bar on web and mobile |
| Automatic punctuation | On most platforms now | Yes |
| Removes fillers and false starts | No | Auto Cleanup — None, Light, Medium, High |
| Resolves a mid-sentence correction | No | Yes |
| Paragraphs and lists without narrating them | No — say "new paragraph" | Yes |
| Personal dictionary you can edit | No dictation vocabulary to edit | Add, pin, auto-learn from your corrections |
| Dictionary changes what is heard | No | Yes — sent with the audio, before recognition |
| Your terms sync across devices and OSes | No | Yes, on your account |
| Tone control | No | Polish · Prompt engineer · Make it shorter · Turn into a reply |
| Transforms you define yourself | No | Yes |
| Snippets — say a cue, get a block | No | Yes |
| Spoken emoji | Varies by platform | Yes |
| Keeps the raw, unedited transcript | It only ever produces raw | "What I actually said" restore |
| Language detected automatically | No — you choose it in settings | 120+ languages, auto-detected, native scripts |
| Live translation | No | Yes |
| Meeting recording with speaker labels | No | Mic + system audio, no bot joins, labels after the meeting |
| Runs on-device, with no network | Often yes on recent hardware | No — transcription needs a connection |
| Same behaviour on every platform you use | Four different products | One, everywhere you sign in |
Built-in dictation changes with OS releases. Everything here describes default behaviour at the time of writing, on the shipping versions — check your own before you take our word for any row.
The other direction
A comparison page where the other side loses every round is a comparison page nobody believes. These are the rounds we lose, and one of them may well decide it for you.
No download, no account, no tier. That is a real advantage and it is the reason most people never look further. If your dictation needs are a sentence at a time into a search box, the built-in one is the right tool and you should keep using it.
Any text field, on a machine you installed nothing on — including one you do not administer. Waver matches this on the desktop build, where the Flow Bar is system-wide. On web and mobile the Flow Bar lives in the app, and you paste out of it.
On recent Apple hardware in particular, dictation often runs on-device. Waver's transcription does not: audio leaves the machine to be transcribed. Privacy Mode and Cloud Sync are two independent toggles and with both set nothing is retained — but "not retained" is not the same claim as "never sent", and if that distinction is the one that matters for what you are about to dictate, use the built-in one.
Words appear while you are still speaking. Waver cannot do that and be right about corrections at the same time — deciding what the start of your sentence meant requires having heard the end of it. So the text arrives when you release the key, not while you talk. Some people find that pause unbearable. It is the honest cost of everything else on this page.
The honest summary: if what you dictate is short, private, offline, or into a search box, the thing already on your machine is the right answer. Waver is for the messages, notes, docs and replies that a person is going to read — where the ten minutes you spend tidying the transcript is the actual cost, not the dictation.
Past the comparison
The seven gaps above are the head-to-head. But the reason people stop reaching for the built-in one is usually something on this list, which has no built-in counterpart to compare against.
Mic and system audio together, with no bot joining the call. Speaker labels arrive after the meeting.
Speak one language, read another, without a round trip through a separate app.
Who said what, on a recording with more than one voice in it.
Turn a recording into a spaced-repetition set — the lecture, the onboarding call, the thing you have to actually retain.
Ask a question of everything you have recorded, rather than opening notes one at a time.
In the same app, from the same voice input.
All of it on iPhone, iPad, Mac and Windows — the same account, the same dictionary, the same behaviour.
Questions
We are not going to quote you a number, because we have not run a benchmark we would be willing to defend, and an invented accuracy figure is worse than no figure. Here is the honest framing instead: on one clear sentence in a quiet room, modern dictation engines — ours and the built-in ones — are all good. The gap this page is about opens after recognition, not during it. It is about fillers, self-corrections, structure, your vocabulary and tone. Judge it on the worked example, which you can reproduce yourself in about fifteen seconds.
Either. On the desktop build the hold-to-talk Flow Bar is system-wide, so it can be the one you reach for everywhere. On web and mobile the Flow Bar is in-app. Plenty of people keep both: the OS one for a quick search box, Waver for anything a person is going to read.
No. Transcription runs in the cloud, so Waver needs a connection. Built-in dictation on recent hardware often does not. On a plane, in a basement, or on a locked-down network, the built-in one wins outright and we would rather say so here than have you discover it at the gate.
That is what the cleanup levels are for. Set Auto Cleanup to None and you get a faithful transcript. Light trims obvious fillers. Medium and High do progressively more restructuring. Whichever you pick, "what I actually said" brings back the raw version — so if you are dictating a quote, a legal note or anything where the exact words matter, you never lose them.
Autocorrect, and the text-replacement features built into desktop operating systems, are string substitutions: they wait for a specific run of characters to appear and swap it. If the recogniser produces something you did not anticipate, nothing fires. Waver's dictionary goes out with the audio, so the recogniser is already leaning toward your terms while it is deciding what it heard. It also learns from your corrections, and you can pin the terms you never want it to drop.
It is a different shape. Those tools operate on text you have already got, as a separate step you invoke on a selection — so the tidying is still a thing you remember to do, in another menu, after the fact. More importantly, they never heard the audio. A name the recogniser got wrong is wrong in a way no text-only pass can recover, because the evidence is gone. Waver's dictionary acts before recognition and the rewrite acts after it. Those are two different fixes for two different failures.
120+ languages, detected automatically and written in their native script. Built-in dictation generally asks you to choose the language up front, in settings, and switch it back afterwards — which is friction if you move between languages mid-day.
Privacy Mode and Cloud Sync are two independent toggles. With both set, there is zero retention. Recordings are never sold and never used to train models, and Waver is ad-free on every plan including free. The privacy policy has the full version, and it is worth reading before you dictate anything your security team would care about.
The free tier is unlimited transcription plus 10 AI actions a month, with no card and no ads. Premium starts at $4.13/mo. It runs on iPhone, iPad, Mac and Windows.
Free forever to start. Unlimited transcription, no card, no ads.