Guides

Dictating into Slack, VS Code and Notion

Dictated text used to land in Slack and VS Code with a doubled space, a capital in the middle of a sentence, or a full stop where the sentence carried on. Those apps hide their text from macOS until an assistive app asks for it. Dictera now asks, and this is what to do when the join still looks wrong.

A row of plain grey blocks standing on a pale surface with one gap in it, a violet block tilting into the gap.

What Dictera is doing at the cursor

Before anything is inserted, Dictera reads a short window of text on either side of your cursor and decides three things about the join: whether a space goes in front of the new text, whether its first letter keeps the capital the speech model gave it, and whether the full stop at the end belongs in a sentence that was already running.

Those come up constantly, because a transcriber hands back a trimmed, capitalized, sentence-ended string every time - right when you are starting a sentence, wrong when you are finishing one. Landing "Store on Main" after "I went to the " should not produce a second capital; landing ", and then" after "word " should close up the space you already typed rather than add one.

Why Slack, VS Code and Notion were the hard ones

Reading around the cursor goes through the macOS Accessibility API, and most Mac apps answer it. Chromium-based apps do not, at least not to start with. Slack, VS Code, Notion, Discord and Teams are all built on Electron, and Chromium builds its accessibility tree lazily: until an assistive client asks for it, the focused text field reports nothing at all.

Nothing is not the same as empty, and Dictera treats the two differently. An empty field takes no separator, because there is nothing to separate from. A field that will not say takes no guessed separator either - a space dropped into the middle of a word is worse than a space missing between two - so the text goes in exactly as dictated, capital and full stop included. In a chat app, that is most of what you say.

Dictera now sends the request Chromium is waiting for, once per app, the first time you dictate into it after that app launches. The tree appears, the window either side of the cursor becomes readable, and the join is fitted the same way it is in Mail or Pages. It arrived in 0.8.3.

If the text still lands plain

  • Dictate a second time. The tree does not always appear the instant it is asked for. The first dictation after you launch Slack can still come back unfitted, and the next one will be right.
  • Check Accessibility. It is the permission dictation runs on, so if nothing is being inserted at all, that is where to look - the first-run walkthrough asks for it, and the security page lists what each permission is used for.
  • Check you are on the current version. This is the newest part of the insertion path.

The one thing that will still look inconsistent

When an app genuinely will not answer, Dictera falls back to remembering the tail of what it inserted last. That is enough to keep a run of dictations joining correctly in a field it cannot read. It is offered only while the same app is still frontmost, within two minutes, and only if no keystroke or click has reached that app since - the moment you type or click, the memory is dropped.

So two dictations in a row read as one paragraph, and a dictation after you have clicked somewhere lands verbatim, with its own capital and stop. That is deliberate: a guess made from where the cursor used to be is confidently wrong, and a missing space is easier to fix by hand than a stray capital and full stop mid-sentence.

When the insertion itself is the problem

Fitting the text is one question; delivering it is another. By default Dictera puts the text on the clipboard and sends ⌘V, which is fast and holds up well for long dictations. Some apps want real keystrokes instead. Turn on Direct Input in Settings under Voice, and Dictera types the characters into the focused app rather than pasting them.

Two details that follow from that choice. Any insertion that starts or ends with a space is typed rather than pasted whatever the setting says, because an app that receives a paste feels free to trim the edges off it, and that space is the whole point. And unless you have turned on Save Text to Clipboard in General, whatever you had copied is put back a moment after the paste, so dictating into Slack does not cost you the link you were about to share.

Note

The text either side of your cursor is surrounding document text, not something you dictated, so it is held to a stricter rule than the rest of the pipeline. Most joins are decided by fixed rules with no model involved at all. The handful the rules admit they are guessing at - is that full stop ending a sentence or an abbreviation? - are passed to a model only when one is already running for this dictation, and only when that model is the built-in one or a server on your own machine. It is never sent to a cloud provider.

What to expect once it is working

Nothing, which is the point. You dictate into a Slack thread mid-sentence and the words join on in lowercase without a stop; you dictate into a VS Code comment after a colon and get one space; you add a clause to a Notion paragraph and the sentence reads as one sentence. Spoken punctuation works there the same as anywhere - the commands are in say the punctuation - and so do the replacements that fix your names and jargon, which run last, after everything else has had its say.

The rest of what happens between speaking and typing is on the dictation page.

Try it on your Mac.

Everything unlocked for 14 days. No card, no account.

Then €5.99/month or €59.99/year · One license, 3 Macs