AI Assistant

Getting an answer during a live call without leaving it

Stop going to the answer and let it come to you. That takes a panel that already heard the question, knows who you are, stays out of your screen share, and answers without a click. Dictera's Assist mode works that way. This is what that takes, and where it still needs you.

A matte grey clay tube lying on a pale surface, with a deep violet ball resting just outside its open end.

Why doesn't a chat tab work mid-call?

Everyone has tried it. Someone asks about last quarter's churn, or the difference between two API versions, and you open a browser tab with a chatbot in it. By the time you have typed the question, you have stopped listening, the person has added a second part you missed, and your silence has gone on long enough to notice.

The tab fails in four ways, and they add up:

  • It did not hear the question. You retype it from memory, shortened, without the part that made it hard.
  • It does not know who is answering. A generic answer to "how would you handle this?" is not your answer, and sounds like it.
  • You look away. On camera, typing into another window looks like typing into another window.
  • It may be on the shared screen. If you are presenting, everyone sees the tab too.

Each of these is a thing the tool has to take off your hands. A faster model fixes none of them.

What has to be true for an answer to arrive in time?

First, it has to be listening already. Assist transcribes both sides of the call as it happens: your microphone for you, and the audio your Mac is playing for everyone else, captured through macOS rather than through a bot in the meeting. The transcript is there before you need it, so the question never has to be retyped. How that capture works, and what it cannot do, is its own post.

Second, it has to know the situation. A template holds the scenario, who you are and any background: a job interview with your CV summarised, a sales call with the pricing you are allowed to quote, a standup with the project you are on. The answer is written from that background and kept consistent with it. You write the template once and pick it before the call.

Third, it must not need a click. When the other side asks a question, a suggested answer appears on its own. That is Auto mode, on by default, and it can be switched off from the panel's Auto button. If you prefer to decide when to ask, ⌥⌘A answers the last question from any app, so the call window keeps focus and you are never seen reaching for a different one. That shortcut, and ⌥⌘L to start or stop listening, are opt-in, so they do not collide with anything until you assign them.

Fourth, it has to stay off the shared screen. The panel is excluded from screen capture by default, with a one-click toggle in the panel itself if you do want it shown.

How does it know a question was asked?

From the punctuation. The speech engine punctuates what it hears, and when a line from the other side comes back with a question mark in it, Auto mode answers. That works in any language the session is set to, including the full-width question mark used in Chinese and Japanese.

It also means the trigger is literal, and it helps to know what that misses:

  • "Walk me through your last project." is a request with a full stop. No automatic answer. Press Answer, or use the shortcut.
  • "Right?" at the end of a statement is a question mark. You may get an answer to something that did not need one.
  • A question you ask yourself does not trigger anything. Auto mode listens to the other side only.

If a second question lands while an answer is still being written, it is not dropped. One more pass runs as soon as the first finishes.

What if the question I care about was three minutes ago?

By default the answer goes to the most recent question the other side directed at you. Conversations do not always work that way: the hard question comes early, then there is small talk, then someone says "no rush on that". Hover any line in the transcript and answer exactly that one, however far back it was said.

That includes lines on your own side. Your microphone hears whatever is in the room, so a line under your label is either something you said out loud or, in a conversation in person, the other person's voice. Answer it and the model works out which, then either writes what you should say or answers you directly.

What if the answer is the wrong shape?

Usually the direction is right and the length is not. Reading four paragraphs while someone waits is not an option. Under every answer are four rewrites, one press each: Shorter, Simpler, More detail, Example. For anything else, type a follow-up under the answer, like "only the second option" or "as a number, not a range", and the answer is reworked with the transcript still in view.

Answers are written to be read at a glance: the direct answer in the first sentence or two, then the reasons and steps, with a short code snippet where the question is about code. A simple factual question gets a short answer, not a structured essay. If the panel is still too much on a crowded screen, compact mode shrinks it to the answer alone.

What about the question that is on the screen, not in the call?

Someone shares a stack trace, a spreadsheet or a diagram and asks what you make of it. Capture the whole display or just the window in front, and it goes with your next request. Dictera's own windows are always left out of the capture. One window is usually the better choice: the same image budget spent on a tenth of the area is the difference between a model that can read the error and one that can only see that there is text. This needs a model that accepts images, which not every local model does.

Which model should answer?

Each mode picks its own, so Assist does not have to share with the write-up at the end. Live answers reward speed over depth: an answer that arrives after the silence has already ended is not much use. The built-in Gemma model runs on your Mac and needs no setup, and a local server such as Ollama or LM Studio works too. Either way, the transcript goes nowhere you do not run yourself.

A cloud provider can write a stronger answer to a hard technical question, and the trade is real: the transcript, and any screenshot you attach, goes to that provider on every request. The security page lists every host the app can contact. If you already use the Claude Code command-line tool, it can answer from a project folder on your Mac, signed in with your own Claude account and read-only by design, since what it is responding to is whatever was said on a call.

What language does the answer come in?

The session's language. Assist runs in one language you choose before the call, transcribes in it and answers in it, and it does not detect a switch mid-call. If a call will move between two languages, a minute of preparation decides which one it follows.

Where does it still need you?

  • It can be wrong. A suggested answer is a draft of what you might say, written from a transcript that can mishear a name or a number. Read it before you repeat it.
  • It cannot tell the other voices apart. Everyone on the far side shares one label. On a panel interview, the answer goes to the latest question, whoever on the panel asked it.
  • It only knows what was said and what you told it. It has no access to your email or your documents unless you put the relevant part in the template, or point Claude Code at a folder.
  • Reading still takes attention. An answer that appears on its own saves the typing, not the reading. The rewrites exist because the shorter version is usually the one you can use.

And the part no setting covers. The panel being invisible to screen sharing does not make every use of it fair. Some interviews and exams forbid outside help of any kind, and some places require everyone's consent before a call is transcribed. Check the rules of the conversation you are in, the same way you would before recording one.

Answers now, notes later

Assist is for the moment a question lands. The write-up is a separate job: start Meeting Notes while Assist is already listening, and the same call is transcribed into both, each in its own session. If the question has nothing to do with the call, double-tap right ⇧ for a single field in the middle of the screen, ask it there, and the answer opens in Chat without touching the recording. Everything the panel does is on the AI Assistant page.

Try it on your Mac.

Everything unlocked for 14 days. No card, no account.

Then €5.99/month or €59.99/year · One license, 3 Macs