Dictation

Two languages in one sentence: what keeps both intact

You can switch languages halfway through a sentence and get each word back in its own language, as long as two things hold: the speech engine you picked covers both languages, and nothing forces the transcript into a single alphabet. Out of the box, the first one usually doesn't hold, and the setting that sounds like it should help can break the second.

What happens to a sentence in two languages?

It passes through three stages on your Mac, and each one treats the mixing differently.

  1. The speech engine turns the audio into words. This is where a mixed sentence is won or lost. If the engine never heard the second language, no later step can bring it back.
  2. The cleanup model fixes punctuation, removes fillers and repairs small mistakes. Its instructions assume a passage may mix several languages, and tell it to keep every word in the language it was spoken in.
  3. Your custom replacements run last and fix the words you have already taught Dictera, in any alphabet.

So most of what decides the result is a choice you make once: which of the four speech engines is active.

Which speech engine hears both languages?

They differ most on exactly this question. Here is what each one is told about language when you dictate:

EngineLanguagesHow it handles language
Apple Speech AnalyzerSet by macOSTranscribes in one language: the one your Mac is set to, or English if macOS has no model for it
Nvidia Parakeet25 European languagesWorks the language out itself, unless you pin one
Nvidia Nemotron100+Works the language out itself; there is nothing to pin
Alibaba SenseVoiceChinese, Cantonese, English, Japanese, KoreanWorks out which of its five it is hearing; there is nothing to pin

Apple's engine is the default on a fresh install, because it is part of macOS. It is also the only one that is given a single language. Dictera passes it the language of your Mac, and there is no setting for adding a second. If you only ever dictate in your Mac's language, that costs you nothing. If you switch whole clauses between two languages, it is the engine to move away from.

The other three are downloads, under Settings ▸ Model, and they are the ones built for this. Pick by your pair:

  • Two European languages, say Polish and English, or Ukrainian and English: Parakeet covers both.
  • Chinese, Japanese or Korean with English: SenseVoice is built around exactly that set.
  • Anything else, or a pair that crosses those groups: Nemotron, which has the widest coverage of the four.

The engine list on the features page has the rest of the detail. Once an engine is downloaded, the Speech Model tile in Quick Controls (⌥⌘C) switches to it mid-flow.

Why pinning a language can break a mixed sentence

Parakeet is the one engine that takes a language pin, under Settings ▸ Model or from the language on the Dictate tile in Quick Controls. The pin exists for a real problem. A multilingual model that is left to guess can get a short utterance wrong in a very visible way: an English phrase comes back spelled in Cyrillic, because for half a second it sounded like Ukrainian. The note under the setting describes exactly that, and suggests you pin your language unless you switch languages between dictations.

What the pin does is narrower than it sounds. It does not tell the model which language to expect. It tells the model which alphabet to write in. Every time the model's first choice for the next piece of text is in another alphabet, it is swapped for the best choice in the pinned one.

For a single-language speaker, that is precisely the fix. For a mixed sentence it depends on the pair:

  • Same alphabet, no harm. German and English, Spanish and English, Polish and Czech all use Latin letters. Pin either language and the other still comes through, because the filter only compares alphabets.
  • Different alphabets, real harm. Pin Ukrainian and say "я вже купив новий laptop", and "laptop" can't be written in Latin letters. It arrives as whatever Cyrillic the model rated next best. Pin English and the Ukrainian half suffers the same fate the other way round.

So the rule for Parakeet is simple. If both of your languages use the same alphabet, pin the one you speak most. If they don't, leave it on Auto-detect, and accept the occasional short phrase that lands in the wrong alphabet as the price of keeping both. Nemotron, SenseVoice and Apple's engine take no pin at all, which is why the language picker only works while Parakeet is active.

What the cleanup model does with a mixed sentence

By the time text reaches cleanup, the hard part is done. The model's job is to edit, and for a mixed passage the instruction that matters most is that it never translates. Every word stays in the language it was spoken in, and punctuation and capitals follow the conventions of whichever language that part of the sentence is in.

The instructions include mixed examples rather than describing the idea in the abstract. "ich habe das neue macbook bestellt and it arrives tomorrow" comes back as "Ich habe das neue MacBook bestellt, and it arrives tomorrow." Both halves are kept, and each gets its own repairs.

Cleanup can also undo some of the alphabet damage from the previous section. When a word was spoken in one language but written phonetically in another alphabet, and the intended word is unambiguous, the model is allowed to restore it: "спотифай" becomes "Spotify". When in doubt, the model is told to leave the text alone, so do not count on it for every word.

After the model, a check compares its output with the transcript and throws away anything that could not be an edit. One of the signals is the whole passage switching alphabet, which is what a translation looks like. That check is deliberately tolerant of mixed text, because a sentence that is half Cyrillic can tip to mostly Latin on a few restored letters and still be yours. Two models per dictation goes through that check, and the rest of the cleanup model's rules, in more detail.

When the same word keeps coming out wrong

If a mixed sentence goes wrong the same way every time, it is usually one word: a colleague's name, the tool your team uses, a term that the engine insists on spelling in the other alphabet. That is what custom replacements are for.

A rule matches a whole word or phrase, ignores capitals, and works in any alphabet. So a rule from "спотифай" to "Spotify" catches that spelling wherever it appears, capitalised or not, and because replacements run after the cleanup model, the rule always wins. It is the dependable fix for the cases where cleanup was not sure enough to act.

What does not switch languages with you

Two things stay fixed however you mix.

Spoken commands are English. "New line", "new paragraph", "question mark" and the rest are matched as English text after transcription, whatever language you are dictating in. Say the punctuation covers how reliably an engine hears an English command dropped into another language. The short version is that the structural ones, said with a pause either side, hold up best.

Mixing is not translating. A mixed sentence comes back mixed. If what you actually want is the whole thing in one language, that is the separate translate shortcut. It takes a sentence in any language, or several, and inserts it in one of 29 target languages. Your custom replacements run before the translation in that mode, so a name is fixed before the translator ever sees it.

A setup that works for most bilingual speakers

  1. In Settings ▸ Model, download the engine that covers your pair and make it active: Parakeet for two European languages, SenseVoice for Chinese, Japanese or Korean with English, and Nemotron for everything else.
  2. On Parakeet, pin a language only if both of yours use the same alphabet. Otherwise leave it on Auto-detect.
  3. Keep cleanup on. It is what keeps both languages punctuated properly and repairs the brand names that were written in the wrong alphabet.
  4. After a week, look through History for the words that keep going wrong, and turn each one into a replacement rule.

None of it needs a connection. Every speech engine runs on your Mac, and so does the built-in cleanup model, so a sentence in two languages is handled the same way on a plane as at your desk. The dictation page has the rest of what happens between pressing the key and the words appearing.

Try it on your Mac.

Everything unlocked for 14 days. No card, no account.

Then €5.99/month or €59.99/year · One license, 3 Macs