Apple, Parakeet, Nemotron or SenseVoice: which to choose
If you dictate in the language your Mac is set to, keep Apple Speech Analyzer, which is already active. Otherwise pick by language: Parakeet for European languages, SenseVoice for Chinese, Japanese, Korean or Cantonese, and Nemotron for anything else.
Which speech model should I use?
All four run on your Mac, and all four do the same job: they turn the recording into a raw transcript, which cleanup then punctuates and tidies. What separates them is which languages they hear, how they decide which one you are speaking, and when the text is ready.
| Model | Languages | Which language it hears | Text ready | Download |
|---|---|---|---|---|
| Apple Speech Analyzer | Set by macOS | Your Mac's system language | While you speak | Part of macOS |
| Nvidia Parakeet | 25 European languages | Works it out, unless you pin one | After you stop | About 670 MB |
| Nvidia Nemotron | 100+ | Works it out; nothing to pin | While you speak | About 670 MB |
| Alibaba SenseVoice | Chinese, Cantonese, English, Japanese, Korean | Works out which of its five | After you stop | About 1.5 GB |
Why is Apple's engine the default?
Because it needs nothing from you. It ships with macOS, macOS keeps it up to date, and it transcribes while you are still talking, so there is little left to do when you stop. It is also the only one that cannot be deleted from Dictera: its files belong to the system.
Its limit is language. Apple's engine transcribes in one language, the one your Mac is set to, and falls back to English when macOS has no speech model for it. If you dictate in English on an English Mac, you can stop reading here. If your Mac is in German and you dictate in Polish, it will hear German.
To see what it covers, click the languages label on its row in Settings ▸ Speech Model. The list comes from macOS, so it can grow with a system update.
When should I switch to Parakeet?
When your language is European and not the one your Mac is set to, or when you move between two European languages. Parakeet covers 25 of them, from English, German and French to Polish, Ukrainian and Russian.
It is also the only model that takes a language pin. Left on Auto-detect, a multilingual model can guess wrong on a short phrase and write English in Cyrillic because it sounded like Ukrainian for half a second. Pinning your language stops that. If you mix two languages in one sentence, the pin is a trade-off; the mixed-language essay explains when to leave it on Auto-detect.
The cost is timing. Parakeet transcribes the whole recording after you stop, so the pause before the text appears grows with what you said. For a sentence you will barely notice it; for three paragraphs you will.
When is SenseVoice the right one?
When you dictate in Chinese, Cantonese, Japanese or Korean, alone or mixed with English. SenseVoice is built around exactly those five languages and works out which of them it is hearing, so there is no pin to set. Outside those five it is the wrong choice.
It is the largest download of the three, at about 1.5 GB, and like Parakeet it transcribes after you stop.
What about Nemotron?
Nemotron is the one to pick when none of the others covers you: 100+ languages across Latin, Cyrillic, CJK and other scripts, detected automatically. It streams like Apple's engine, so the text is mostly written by the time you stop.
It is also a reasonable choice if you move between languages from different groups, say Ukrainian at work and Japanese at home, and would rather not switch models between them.
How do I change the speech model?
- Open Settings ▸ Speech Model. Each model has a row with its vendor, size and a label you can click for its full language list.
- Click Download on the one you want. The row shows the percentage as it goes. When the download finishes, that model becomes the active one.
- On Parakeet, set the Dictation Language. The section appears only while Parakeet is active. Pin your language unless you switch languages between dictations.
- Dictate a few sentences. Try something with names and numbers in it, the kind of text you actually write.
Once you have more than one model downloaded, you do not need Settings to move between them. The Speech Model tile in Quick Controls (⌥⌘C) switches mid-flow. Pick one there that is not downloaded yet, and Dictera opens Settings where the download is, instead of leaving your next dictation with nothing to transcribe it.
Keep Model Ready, in the same pane, is on by default: the active model is loaded at launch and kept in memory, so the first dictation starts straight away. Turn it off if memory is tight and you do not mind a short wait on the first one.
A model you stop using can be deleted from its row to get the disk space back, and downloaded again whenever you want it.
This choice is for dictation. Meetings in the AI Assistant are transcribed by Apple's engine in the one language you choose for that session, whichever model is active here.
Does any of them send my voice somewhere?
No. Whichever model is active, it runs on your Mac, and the audio is never sent to a server or written to disk as a recording. The network is used to download a model, once, and not to use it. What Dictera sends over the network lists every connection the app makes and how to check it yourself, and the security page covers the rest.


