Dictation

Why dictating is faster than typing, and when it is not

Ordinary speech runs at roughly three times a comfortable typing speed, and that gap is real. What it buys you depends on what you are writing: dictation wins on prose you already have in your head, and loses on anything short, exact, or still being worked out.

A tall stack of blank matte clay discs with one violet disc banded through it, and two loose discs lying beside it, one of them deep violet.

How much faster, really?

Dictera makes this estimate itself. Settings ▸ Statistics compares the seconds you spent recording with how long the same number of words would take to type at 40 words a minute, and shows the difference as time saved. Only the counts are kept - words, sessions, seconds - never the text. The weekly summary that reports it is off until you turn it on.

Forty is a reasonable stand-in for a fluent, non-specialist typist, and most people speak at something like three times it without trying. So the arithmetic comes out flattering, and for a long email or a first draft it is not far wrong. The interesting part is what the number quietly leaves out.

Note

Time saved counts the seconds you were recording. It does not count the pause between letting go of the key and the words landing, and it does not count anything you fixed afterwards. Treat it as the ceiling, not the result.

Where the time actually goes

A dictation is not one event. Between the key going down and the text appearing, four things happen in order:

  • You speak. This is the part the stopwatch sees, and the only part where speech is obviously ahead.
  • The audio becomes a transcript. Two of the four speech engines - Apple's Speech Analyzer and Nvidia's Nemotron - transcribe while you are still talking, so when you stop, most of the work is already done. The other two transcribe the whole recording after you stop. That choice is the single biggest lever on how long the pause feels.
  • The transcript becomes writing. Spoken commands are applied, the on-device cleanup model punctuates and removes the fillers, and your custom replacements run last so they always win. Cleanup is the slowest step, and the one you can switch off.
  • The text is inserted. Dictera reads what is immediately before your cursor, decides on the spacing and the capital, and pastes. Pasting is near-instant regardless of length; Direct Input, which types the characters one by one instead, is off by default for that reason.

None of those steps is slow on its own. Together they are why a two-word dictation can feel slower than typing two words, even though speech won the part that was measured.

When typing wins

Anything short

The steps above are a fixed cost, and a fixed cost is a tax on short text. "ok, thanks" is faster typed than dictated, and will remain so. The break-even sits somewhere around a sentence, and it moves with your setup: a streaming model with cleanup off pushes it down, a batch model that has to load first pushes it up.

Text that has to be exact

File paths, command flags, variable names, license keys, anything where a stray comma is a bug. Speech recognition is built to produce readable prose, and the cleanup model's whole job is to make what you said read better - which is the opposite of what you want here. Turning cleanup off from Quick Controls gets you the raw words, and it is genuinely useful for a code comment or a commit message, but you will still proofread. Typing thirty exact characters is faster than saying them and checking them.

Sentences you have not finished thinking

Typing lets you stall mid-clause with your hands on the keys. Speech does not. A recording that is running while you think is either a long silence in the middle of your audio, or - with stop on silence turned on - two dictations, cleaned up separately, with a paragraph break where you did not want one. This is the honest reason some people bounce off dictation in the first week. Prose you have already composed goes fast; prose you are composing as you go does not.

Editing what is already there

Dictation appends. There is no spoken backspace, and "scratch that" throws away the recording you are in the middle of rather than the sentence you regret. Fixing existing text is a different move: select it and use Correct Selection for grammar and spelling, or Edit Selection to say the change out loud - "make it shorter", "turn this into bullet points". Both are fast. Neither is dictation, and treating one as the other is how a five-minute email becomes fifteen.

The room you are in

An open-plan desk, a shared office, a train, a call you are already on. No setting fixes this one. Dictation is fastest where you can speak in a normal voice without composing a second message in your head about who can hear you, and there is no version of it that is quick in a room where you would rather whisper.

So when should I use it?

The people who get the most out of dictation are not speaking faster. They have sorted their writing into two piles and stopped thinking about it.

Reach for the keyReach for the keyboard
An email you know the shape ofA two-word reply
A first draft, notes, a long messageA password, a path, a command
Anything longer than a couple of sentencesA search box
A reply you would otherwise put offA sentence you are still deciding on
Writing in a language you speak better than you typeAnything inside a table or a form

The second habit is smaller: say the structure and let the model do the punctuation. "New paragraph" is worth saying; "comma" usually is not, because the cleanup model adds it anyway. Dictating every mark out loud costs real time and produces text no better than dictating none of them.

Making the gap bigger

Four settings account for most of the difference between a setup that feels quick and one that feels like waiting:

  • A streaming speech model, if the languages you speak are covered by one. The transcript is being written while you talk instead of after.
  • Keep Model Ready, which is on by default for the speech model and off by default for the cleanup model. Turning the second one on trades memory for a shorter pause on the first dictation after a quiet spell.
  • The right recording mode. Hold-to-talk has no end-of-recording decision to make, which suits short bursts; the double-tap suits paragraphs. Both, and stop on silence, are worth ten minutes of trying.
  • Custom replacements for the names and jargon you would otherwise correct by hand every time. A word you fix twice a day is worth a rule.

What does not help is dictating faster. Recognition accuracy falls off with speed, and every misheard word is a trip back to the keyboard, which is the expensive kind of time.

The test that matters

Words a minute is the wrong measure for most people, because most writing is not typing-bound. It is bound by the sentence you have not decided on, the reply you have been avoiding, and the wrist that hurts by four o'clock. Dictation is not much help with the first. It is a great deal of help with the other two, and neither shows up in a stopwatch.

So the useful test is not whether the words appeared faster than you could have typed them. It is whether the writing got done. Give it a week of ordinary work, look at the statistics if you like the numbers, and discount them by whatever you spent fixing things. If what is left is an hour, that is an hour. The dictation page covers what happens in that pause between the key and the words.

Try it on your Mac.

Everything unlocked for 14 days. No card, no account.

Then €5.99/month or €59.99/year · One license, 3 Macs