Everything Tara and Tara Flow run on. On-device by default, always — the one cloud option below is off unless you turn it on yourself, and it only ever sends the single thing it needs to.
English speech recognition, runs resident in memory for fast repeat transcriptions.
NVIDIA's model, int8 quantized. Tara Flow's own pipeline, independent of Whisper — downloads only if you turn Flow on.
Turns raw transcript into a clean reply or a tidied-up message — filler removed, punctuation fixed, self-corrections resolved. Four on-device options to choose from, plus one cloud option in Tara Flow.
Fastest replies, lowest memory use. No reasoning-mode toggle to fight — the 2507-Instruct line skips it entirely, which made it more reliable than the alternatives below at this size.
Noticeably sharper than the 4B for about the same speed. Good upgrade if you're not sure which to pick.
Different training data and phrasing style than Qwen — worth trying if Qwen's voice doesn't feel right for you.
Strongest reasoning of the four — best for tricky corrections or multi-step requests. Slower, more memory.
Swaps Tara Flow's tidy-up pass to a larger cloud model for better quality. Your transcript leaves your Mac for this one request only — everything else about Tara stays exactly as private as always. Off by default; a paid feature when it's on.
Open-weight text-to-speech, multiple voices to choose from in Settings.
Every model above except Gemini 3.1 Pro runs entirely on your Mac — nothing you say is ever sent anywhere by default. Smart Mode is the one deliberate exception, and it stays off until you turn it on. That's the whole cloud story, not a footnote to it.