Waver vs built-in dictation

Every OS has dictation. None of them wait for the end of your sentence.

macOS, Windows, iOS and Android all turn speech into text, for free, and they are genuinely good at it. What none of them do is edit. They write down the filler, the false start, the hedge, and both halves of a correction — faithfully, in the order you said them — and hand you a paragraph to clean up before you dare send it.

This page is about that gap. It is one gap, it has a mechanical cause, and the whole argument fits in a single sentence taken apart word by word.

See the sentence

One sentence, three outcomes

You said one thing. Watch what each one writes down.

Nineteen words, spoken into a message to a colleague. Nothing unusual about them — no accent to blame, no background noise, no technical vocabulary. Just the way people talk when they are thinking and speaking at the same time.

What you said

so um can you tell the team the launch is gonna slip to like not friday the following monday

What built-in dictation writes

So, um, can you tell the team the launch is gonna slip to, like, not Friday, the following Monday.

Punctuation placement is the one thing that differs between the four platforms — some auto-punctuate, some wait for you to say "comma". The words are identical on all of them, because faithful transcription is the specification. Note that both dates survive, in the order you said them.

What Waver writes

Can you tell the team the launch is slipping to the following Monday, not Friday?

Same recording · Turn into a reply

Heads-up — the launch is slipping. It won't make Friday; we're now looking at the following Monday.

One recording, two shapes, no second dictation. And if you want the raw version back — every "um", both dates, in the order you said them — "what I actually said" returns it untouched.

Fragment by fragment

You saidWhat it isBuilt-inWaver
so umDiscourse marker, then a fillerBoth written outBoth dropped
gonnaCasual contractionKept as spokenResolved into the verb
likeHedge, not a comparisonWritten out, usually with commasDropped
not fridayThe half you are cancellingKept, in spoken orderKept, demoted to a trailing contrast
the following mondayThe half you actually meanKept, in spoken orderPromoted into the main clause

The last two rows are the interesting ones. "Not Friday, the following Monday" is a correction, and a correction has a winner. Built-in dictation preserves both candidates and leaves the arbitration to whoever reads the message — which, on a bad day, is how a team ships on the wrong date.

The gap, itemised

Seven things built-in dictation does not do

None of these are bugs. Built-in dictation does exactly what it says: it transcribes. Every item below is a thing it was never trying to do, which is precisely why it has not improved in the years you have been using it.

01

Self-correction

Built-in

Writes both halves, in the order you said them

Waver

Resolves which half survives

Nobody speaks in final drafts. You say a date, hear it, and fix it in the same breath — "not Friday, the following Monday", "five pm, actually six". Built-in dictation has already typed "5pm" into your message before it hears "actually six", so you get both, and the reader has to work out which one is real. It is the single most common reason a dictated message needs editing before you dare send it.

02

Fillers and false starts

Built-in

Transcribed faithfully

Waver

Auto Cleanup, four levels

"Um", "like", "you know", "I mean", the sentence you started twice. Built-in dictation is doing its job correctly when it writes these down — faithful transcription is the specification. But nobody wants to read their own disfluencies, so the tidying lands on you. Waver's Auto Cleanup has four settings (None, Light, Medium, High) so you choose how much of your voice survives, and "what I actually said" restores the raw transcript whenever you want it back.

03

Structure

Built-in

One long paragraph, unless you narrate the formatting

Waver

Paragraphs and lists from a flat ramble

Every built-in dictation engine will insert a break if you say "new paragraph" out loud. That works, and it also means holding the document's shape in your head while you are trying to hold the argument in your head. Three minutes of talking through four options arrives as one continuous block. Waver reads the whole thing after you stop and gives it the shape the content already had.

04

A dictionary that changes what is heard

Built-in

No editable dictation vocabulary

Waver

Add, pin, and auto-learn from your corrections

This is the one people underestimate, and it has its own section below. The short version: your colleague's name, your product, your internal service. Built-in dictation gives you no list to add them to. Waver does — and the list is sent with the audio, so it changes what the recogniser hears rather than find-and-replacing afterwards.

05

Terms that follow you between devices

Built-in

Per-device, per-OS, not portable

Waver

On your account, everywhere you sign in

Whatever your Mac has learned about how you speak, your Android phone has not. Nor has the work laptop. There is no built-in path for teaching four separate operating systems the same forty proper nouns, so in practice people teach none of them. Waver's dictionary rides your account: add a term on the phone, and the browser, the desktop app and the tablet all have it on the next recording.

06

Tone

Built-in

Not a dictation feature

Waver

Transforms, including ones you write

The same two minutes of talking is a Slack message, a status update, or a note to yourself, depending on who is reading it. Waver ships Polish, Prompt engineer, Make it shorter and Turn into a reply, plus custom transforms you define once and reuse. Built-in dictation has one output: the words, as spoken. Anything else is your problem, in another app.

07

Shorthand

Built-in

None

Waver

Snippets and spoken emoji

Say a cue and a Snippet expands into the block you set up — a sign-off, an address, a standing disclaimer you retype forty times a month. Say "fire emoji" and you get the flame rather than the words. Small things, but they are the difference between dictation as a novelty and dictation as the way you actually write.

Why the gap exists

It is not that nobody got round to it.

Built-in dictation is a live streaming transcriber. It commits words into your text field while you are still speaking, because that is the product: the cursor keeps up with your mouth. Once a word is in the document, it has been typed. It cannot be un-said.

Now look at what fixing a self-correction requires. To know that "Friday" was cancelled, you have to have heard "the following Monday" — information that arrives seconds after"Friday" was already committed to the screen. A streaming transcriber is being asked to revise text it has already handed over. That is not a hard engineering problem; it is the wrong shape of problem for the architecture.

The same constraint explains the rest of the list. Filler removal needs to know the sentence continued. Paragraph structure needs to know where the argument turned. Tone needs the whole thing. Every one of them requires an ending that a live transcriber does not have yet.

Built-in

Hears a word → commits it → hears the next word. The document is written forwards and never revisited.

Waver

Holds the utterance → transcribes it whole, with your dictionary → then decides what the beginning meant, knowing how it ended.

The trade is stated plainly further down, and it is a real one: you do not watch words appear as you talk. You get them when you let go of the key.

Before vs after

A dictionary that changes what is heard, not what is printed

This is the difference people find hardest to believe until they see it, so it is worth being precise about the mechanism rather than the marketing.

Text replacement — the feature your desktop OS has had for a decade — is a string substitution. It waits for an exact run of characters and swaps it for another. That works when you can predict the mistake. You cannot: a name the recogniser has never encountered comes back as a phonetic guess, and next time it comes back as a different phonetic guess, because the audio was slightly different. There is no string to key on.

Waver sends your Personal Dictionary with the audio, before anything is recognised. The recogniser is already leaning toward your terms while it decides what it heard, so the guess never happens. The dictionary is not a spellchecker bolted on afterwards — it is an input.

Add

Type in the terms you know it will trip on — product names, colleagues, the internal service nobody outside the building has heard of.

Pin

Mark the ones that must never be dropped, however long the list gets.

Auto-learn

Fix a word once and Waver notices. Corrections feed the list without you managing it.

And it is one list. Built-in dictation is four separate products that happen to share a name — nothing your Mac has learned is available to your phone, and there is no path between them. Waver's dictionary lives on your account, so the term you add on the train is in effect on the desktop app before you sit down.

Side by side

The whole comparison, in one table

"Built-in dictation" means whichever of these four you have in front of you. They differ in the details; on everything in the table below, they behave alike.

macOS

Dictation, in Keyboard settings

Windows

Voice typing, Win + H

iOS / iPadOS

The microphone on the keyboard

Android

Gboard voice typing

CapabilityBuilt-in dictationWaver
Turns speech into textYesYes
Costs nothing to startYes, unlimitedFree tier: unlimited transcription, 10 AI actions a month, no card, ad-free
Works in any text field, system-wideYesYes on the desktop build; in-app Flow Bar on web and mobile
Automatic punctuationOn most platforms nowYes
Removes fillers and false startsNoAuto Cleanup — None, Light, Medium, High
Resolves a mid-sentence correctionNoYes
Paragraphs and lists without narrating themNo — say "new paragraph"Yes
Personal dictionary you can editNo dictation vocabulary to editAdd, pin, auto-learn from your corrections
Dictionary changes what is heardNoYes — sent with the audio, before recognition
Your terms sync across devices and OSesNoYes, on your account
Tone controlNoPolish · Prompt engineer · Make it shorter · Turn into a reply
Transforms you define yourselfNoYes
Snippets — say a cue, get a blockNoYes
Spoken emojiVaries by platformYes
Keeps the raw, unedited transcriptIt only ever produces raw"What I actually said" restore
Language detected automaticallyNo — you choose it in settings120+ languages, auto-detected, native scripts
Live translationNoYes
Meeting recording with speaker labelsNoMic + system audio, no bot joins, labels after the meeting
Runs on-device, with no networkOften yes on recent hardwareNo — transcription needs a connection
Same behaviour on every platform you useFour different productsOne, everywhere you sign in

Built-in dictation changes with OS releases. Everything here describes default behaviour at the time of writing, on the shipping versions — check your own before you take our word for any row.

The other direction

Four things built-in dictation does better

A comparison page where the other side loses every round is a comparison page nobody believes. These are the rounds we lose, and one of them may well decide it for you.

It is already there

No download, no account, no tier. That is a real advantage and it is the reason most people never look further. If your dictation needs are a sentence at a time into a search box, the built-in one is the right tool and you should keep using it.

It goes everywhere by default

Any text field, on a machine you installed nothing on — including one you do not administer. Waver matches this on the desktop build, where the Flow Bar is system-wide. On web and mobile the Flow Bar lives in the app, and you paste out of it.

It can work with no network

On recent Apple hardware in particular, dictation often runs on-device. Waver's transcription does not: audio leaves the machine to be transcribed. Privacy Mode and Cloud Sync are two independent toggles and with both set nothing is retained — but "not retained" is not the same claim as "never sent", and if that distinction is the one that matters for what you are about to dictate, use the built-in one.

It is live

Words appear while you are still speaking. Waver cannot do that and be right about corrections at the same time — deciding what the start of your sentence meant requires having heard the end of it. So the text arrives when you release the key, not while you talk. Some people find that pause unbearable. It is the honest cost of everything else on this page.

The honest summary: if what you dictate is short, private, offline, or into a search box, the thing already on your machine is the right answer. Waver is for the messages, notes, docs and replies that a person is going to read — where the ten minutes you spend tidying the transcript is the actual cost, not the dictation.

Past the comparison

And then there is everything that is not dictation at all

The seven gaps above are the head-to-head. But the reason people stop reaching for the built-in one is usually something on this list, which has no built-in counterpart to compare against.

Meeting recorder

Mic and system audio together, with no bot joining the call. Speaker labels arrive after the meeting.

Live translation

Speak one language, read another, without a round trip through a separate app.

Speaker separation

Who said what, on a recording with more than one voice in it.

Flashcards and review

Turn a recording into a spaced-repetition set — the lecture, the onboarding call, the thing you have to actually retain.

AI chat across your notes

Ask a question of everything you have recorded, rather than opening notes one at a time.

Image, video and music studios

In the same app, from the same voice input.

All of it on iPhone, iPad, Mac and Windows — the same account, the same dictionary, the same behaviour.

Questions

The ones worth asking

Is Waver more accurate than my Mac's dictation?+

We are not going to quote you a number, because we have not run a benchmark we would be willing to defend, and an invented accuracy figure is worse than no figure. Here is the honest framing instead: on one clear sentence in a quiet room, modern dictation engines — ours and the built-in ones — are all good. The gap this page is about opens after recognition, not during it. It is about fillers, self-corrections, structure, your vocabulary and tone. Judge it on the worked example, which you can reproduce yourself in about fifteen seconds.

Does Waver replace built-in dictation, or sit alongside it?+

Either. On the desktop build the hold-to-talk Flow Bar is system-wide, so it can be the one you reach for everywhere. On web and mobile the Flow Bar is in-app. Plenty of people keep both: the OS one for a quick search box, Waver for anything a person is going to read.

Does it work offline, like my phone's dictation does?+

No. Transcription runs in the cloud, so Waver needs a connection. Built-in dictation on recent hardware often does not. On a plane, in a basement, or on a locked-down network, the built-in one wins outright and we would rather say so here than have you discover it at the gate.

Will it change words I did not want changed?+

That is what the cleanup levels are for. Set Auto Cleanup to None and you get a faithful transcript. Light trims obvious fillers. Medium and High do progressively more restructuring. Whichever you pick, "what I actually said" brings back the raw version — so if you are dictating a quote, a legal note or anything where the exact words matter, you never lose them.

What does the personal dictionary actually do that autocorrect does not?+

Autocorrect, and the text-replacement features built into desktop operating systems, are string substitutions: they wait for a specific run of characters to appear and swap it. If the recogniser produces something you did not anticipate, nothing fires. Waver's dictionary goes out with the audio, so the recogniser is already leaning toward your terms while it is deciding what it heard. It also learns from your corrections, and you can pin the terms you never want it to drop.

My OS added an AI rewrite feature. Is that not the same thing?+

It is a different shape. Those tools operate on text you have already got, as a separate step you invoke on a selection — so the tidying is still a thing you remember to do, in another menu, after the fact. More importantly, they never heard the audio. A name the recogniser got wrong is wrong in a way no text-only pass can recover, because the evidence is gone. Waver's dictionary acts before recognition and the rewrite acts after it. Those are two different fixes for two different failures.

What about languages other than English?+

120+ languages, detected automatically and written in their native script. Built-in dictation generally asks you to choose the language up front, in settings, and switch it back afterwards — which is friction if you move between languages mid-day.

What happens to my audio?+

Privacy Mode and Cloud Sync are two independent toggles. With both set, there is zero retention. Recordings are never sold and never used to train models, and Waver is ad-free on every plan including free. The privacy policy has the full version, and it is worth reading before you dictate anything your security team would care about.

What does it cost?+

The free tier is unlimited transcription plus 10 AI actions a month, with no card and no ads. Premium starts at $4.13/mo. It runs on iPhone, iPad, Mac and Windows.

Privacy policyMeeting recorderWaver for developersNo card · Ad-free · 120+ languages

Stop typing what you could have just said.

Free forever to start. Unlimited transcription, no card, no ads.