{"slug": "speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux", "title": "Speakoflow – Free Voice Dictation and AI Assistant for Windows, macOS, and Linux", "summary": "SpeakoFlow launched as a free, open-source voice dictation and AI assistant for Windows, macOS, and Linux that transcribes speech locally by default and works in any app via keyboard shortcuts. The tool ships with 65 speech models in its catalog, including Parakeet for English, Nemotron for 28 languages with automatic detection, and Whisper for 99 languages, and lets users route the assistant through a built-in offline model, their own Ollama or LM Studio server, or cloud providers such as OpenAI, Groq, Mistral, Azure AI Speech, and OpenRouter with their own API key. Optional features include screen vision, web search through Serper, Brave, Tavily, Exa, SerpAPI, or TinyFish, reminders, and profiles with memory that is off by default and stored on the user's computer.", "body_md": "Talk and it types, in any app. Ask and it answers. Start a conversation and talk it through.\n\nIt also writes up your meetings. Free, open source, and local by default.\n\nYou think faster than you type. SpeakoFlow lets you work with your voice instead, from any app, with a keyboard shortcut:\n\n| **Dictate** `Left Ctrl` +`Left Win` Hold the keys, talk, let go. What you said is typed at your cursor, in any app that takes text. | **Ask** `Left Ctrl` +`Left Alt` Select some text if you like, then ask out loud. Translate it, reply to it, explain it. Copy the answer, insert it, or replace the selection. | **Converse** `Left Ctrl` +`Left Alt` +`C` A hands-free conversation with the assistant. Talk a problem through and it answers out loud. Press the keys again to end it. | \n\n<sub>Windows defaults. The macOS and Linux keys are under\n[Keyboard shortcuts](#keyboard-shortcuts), and you can change every one.</sub>\n\nSpeech is transcribed on your own computer unless you pick a cloud service. The assistant uses whichever model you choose: a built-in one that works offline, your own Ollama or LM Studio server, or a cloud provider with your own API key.\n\nI built SpeakoFlow while studying alone for exams. I was paying for a dictation app that could hear me but couldn't help me, so I made one that does both.\n\nHold the keys, talk, and let go. Your words appear wherever the cursor is, and the recording pill can show them as you speak. Transcription runs on your graphics card or processor with a local model: Parakeet by default for English, Nemotron for 28 languages with automatic detection, Whisper for 99, and 65 speech models in the catalog altogether. If you'd rather use the cloud, ElevenLabs and Deepgram stream text while you talk, and OpenAI, Groq, Mistral, Azure AI Speech, and OpenRouter work too.\n\n- **Undo.** Cancelled a recording by accident? The pill offers Undo for a few\nseconds, and History keeps the recording so you can recover it later. A\nfailed transcription offers Try again.\n- **Your words, spelled your way.** Add names and jargon to the Dictionary, or\nwrite text replacements. On Windows it can learn the words you correct.\n- **Translate to English.** Whisper, Canary, Granite Speech, and Voxtral\nmodels, and OpenAI or Groq in the cloud, can turn speech in another language\ninto English text.\n- **Hold or tap.** Hold the keys while you talk, or switch to tap so one press\nstarts and the next one stops.\n\nSelect some text, or don't. Hold `Left Ctrl` + `Left Alt` and\nsay what you want: \"translate this to Spanish\", \"write a polite reply saying I\ncan't make Thursday\", \"explain this\". The answer streams into a card over the\napp you're in, and you can copy it, insert it at the cursor, or put it in\nplace of the text you selected.\n\nIt can do more if you let it:\n\n- **Screen vision.** Ask about the error in your terminal or the chart in your\nspreadsheet. It's off until you switch it on, and even then the model\ndecides per question whether it needs to look. A screenshot it doesn't use\nnever leaves your computer.\n- **Web search** through Serper, Brave, Tavily, Exa, SerpAPI, or TinyFish.\n- **Reminders.** \"Remind me to send the invoice in twenty minutes.\" Reminders\nsurvive a restart and pop up without taking your keyboard.\n- **Profiles and memory.** Give it different personas, each with its own reply\nlength, and let it remember how you like to work. Memory is off by default,\nstays on your computer, and you can edit or erase it.\n\nPress `Left Ctrl` + `Left Alt` + `C` and talk. The\nassistant answers out loud, and you can cut in while it's speaking. Esc stops a\nreply without ending the conversation, and pressing the keys again ends it.\nNeed to type something in the middle? Dictate as usual. The conversation waits\nwhile you do and picks up again afterwards.\n\nReplies are spoken by a voice on your computer (Kokoro, Kitten, Pocket TTS, or Supertonic) or by a cloud voice from OpenAI, ElevenLabs, Deepgram, Cartesia, Google, Azure, and others.\n\nPress Start recording before a call. SpeakoFlow records your microphone and your computer's audio as two streams and transcribes them as people speak, so it always knows which words were yours. No bot joins the meeting, and it works with any meeting app. On Windows it can offer to record when it notices a call.\n\nWhen the call ends it writes the notes, with a summary, decisions, and next steps with owners, using a template you pick (General, Standup, One-on-one, Interview, or Action items). Other voices are labelled Speaker 1, Speaker 2, and so on. Afterwards you can ask questions about the meeting, or choose Discuss this meeting to talk it over in a conversation.\n\nOn macOS, recording the other side of a call needs a virtual audio device such\nas BlackHole. See [Troubleshooting](#troubleshooting).\n\nSpeakoFlow Mini is a small model we trained for one job: turning what you said into clean text. It removes filler words, fixes grammar and punctuation, and follows spoken edits, so \"scratch that\" or \"actually, eleven\" does what you meant instead of being typed out. It's a 795 MB download, runs on your computer, and handles English for now. Any other local or cloud model can do the job instead, including Apple Intelligence on Apple silicon Macs.\n\nCleanup is off until you turn it on. It then gets its own shortcut (the dictation keys plus Shift) or runs on every dictation. On top of it you can add a writing style: Professional, Friendly, Concise, Formal, Casual, or one you write yourself.\n\n- **Insights.** Words dictated, your speaking speed, time saved compared with\ntyping at 40 words a minute, and six months of activity with streaks.\n- **History.** Your dictations, questions, and conversations. Play a recording\nback, transcribe it again, recover a dismissed one, or\ncontinue a chat as a conversation. Old recordings can delete themselves after\na set number, days, or months.\n- **Models you already have.** Add a`.gguf` or Whisper`.bin` file, or link a\nfolder and every model in it shows up. Nothing is copied or moved. Downloads\nthat do happen fetch eight chunks at once and resume where they stopped.\n- **Generate with Flow.** Start a dictation with \"Hey Flow\" and describe what\nyou want written, and the finished text is pasted instead of your words. It's\noff by default, under Settings → Dictation.\n- **20 interface languages.**\n\nEach feature has its own page in the [documentation](https://www.speakoflow.com/docs).\n\n| <sub>**Home.** Your shortcuts, the models doing each job, and what you dictated last.</sub> | <sub>**Assistant.** Its model, its voice, and what it's allowed to do.</sub> | \n| <sub>**AI cleanup.** Say it messy, get it clean, in the style you pick.</sub> | <sub>**Insights.** How much you dictate, and how much typing it saved.</sub> | \n| <sub>**Models.** Each job runs on this computer or in the cloud.</sub> | <sub>**Voices.** Four local voices and a dozen cloud ones.</sub> | \n\n| Action | Windows | macOS | Linux | \n|---|---|---|---|\n| Dictate | `Left Ctrl + Left Win` | `Fn` (🌐) | `Ctrl + Space` | \n| Ask the assistant | `Left Ctrl + Left Alt` | `Fn + Ctrl` | `Ctrl + Alt + Space` | \n| Start or end a conversation | `Left Ctrl + Left Alt + C` | `Fn + Ctrl + C` | `Ctrl + Alt + C` | \n| Dictate and clean up <sup>1</sup> | `Left Ctrl + Left Win + Shift` | `Fn + Shift` | `Ctrl + Shift + Space` | \n| Cancel | `Esc` | `Esc` | Not available yet | \n\n<sup>1</sup> Only while AI cleanup is on and has its own shortcut.\n\nThe pattern is the same on every platform. Add Shift to the dictation keys to\ndictate and clean up, and add C to the ask keys to start a conversation.\nRecording shortcuts work while you hold them; switch the Home page from **Hold\nto talk** to **Tap to toggle** and one press starts, the next one stops.\n\nEsc only cancels while something is running, like a recording or a reply being read aloud, so other apps keep their Esc the rest of the time. To change a shortcut, click its keys. Cancel and the conversation shortcut can also be turned off from there.\n\nOn a Mac, set **System Settings → Keyboard → Press 🌐 key to** to **Do\nNothing**, or the globe key opens the emoji picker as well. Macs that were on\nthe older Option + Space default keep it after updating.\n\nScripts and window managers can control SpeakoFlow with\n[command-line flags](https://www.speakoflow.com/docs/settings/cli) such as\n`--toggle-transcription`.\n\nEvery job can run on your computer or with a provider you choose. Cloud providers use your own API key, stored in your system keychain.\n\n| Job | On your computer | In the cloud, with your key | \n|---|---|---|\n| Speech to text | Parakeet, Nemotron, Canary, Cohere Transcribe, Whisper, Moonshine, Voxtral, Qwen3-ASR, GigaAM, Granite Speech, SenseVoice, and more (65 in the catalog) | ElevenLabs, Deepgram, OpenAI, Groq, Mistral (Voxtral), Azure AI Speech, OpenRouter, or any OpenAI-compatible server | \n| Assistant and cleanup | Built-in engine (llama.cpp, fully offline), Ollama, LM Studio, SpeakoFlow Mini for cleanup, Apple Intelligence for cleanup on Apple silicon | OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, OpenRouter, Groq, Cerebras, xAI, DeepSeek, Mistral, Moonshot, Together AI, Fireworks AI, Perplexity, Z.AI, or any OpenAI-compatible endpoint | \n| Spoken replies | Kokoro, Kitten, Pocket TTS, Supertonic | OpenAI, ElevenLabs, OpenRouter, Deepgram, Cartesia, Google Cloud, Azure AI Speech, Groq, xAI, Mistral, Inworld, or your own server | \n| Web search (optional) |  | Serper, Brave, Tavily, Exa, SerpAPI, TinyFish | \n\nBy default your voice is transcribed on your computer and never uploaded. There is no telemetry, no analytics, and no account.\n\nData leaves your computer only for things you set up yourself:\n\n- **A cloud speech service** , if you pick one instead of a local model. It\nreceives your recordings.\n- **The assistant's provider** , if it isn't a local one. It receives your\nquestions, any text you selected, a screenshot when screen vision is on and\nthe model asks for one, and a meeting's transcript when you generate notes or\nask about it.\n- **Web search** , if you turn it on. The search provider receives the query.\n- **Feedback** , if you send it from the app. It sends exactly what the dialog\nshows you, to a private issue tracker only the developer can read.\n\nAPI keys live in your system keychain. Memory is off until you turn it on, and\nit stays on your computer where you can view, edit, or erase it. More detail is\non the [privacy page](https://www.speakoflow.com/docs/reference/privacy).\n\nDownload the latest build from the\n[Releases](https://github.com/AbhishekBarali/SpeakoFlow/releases) page. On\nfirst launch you pick a speech model, and a short tour shows the shortcuts\nwhile it downloads.\n\nAlready on 1.4 or earlier? Those versions can't update themselves to 2.0, so download 2.0 once from Releases and install it over the old one. Your settings and history are kept. From 2.0 on, updates install from inside the app.\n\nRun the `.exe` installer. Windows may show a SmartScreen notice because the\ninstaller isn't signed by a known publisher yet; choose **More info → Run\nanyway**.\n\nDownload the `.dmg` for your Mac (`aarch64` for Apple silicon, `x64` for Intel)\nand drag **SpeakoFlow** into Applications. The app isn't signed by Apple yet,\nso macOS says it \"is damaged and can't be opened\". It isn't damaged. Clear the\nblock once with this command in Terminal, then open the app normally:\n\n```\nxattr -dr com.apple.quarantine /Applications/SpeakoFlow.app\n```\n\nSpeakoFlow then asks for **Microphone** and **Accessibility** permission so it\ncan hear you and type into other apps.\n\n## More about the macOS install\n\nThe \"damaged\" message is what macOS shows for any app it can't trace to a paid\nApple Developer account. Signing costs $99 a year, which this project doesn't\nhave yet. macOS 15 and later removed the old right-click → **Open** bypass, and\nthis message is the one case where System Settings offers no **Open Anyway**\nbutton, so Terminal is the only way through. The command removes the\n\"downloaded from the internet\" tag from that copy of the app.\n\nYou run it once per download. Updates installed from inside the app aren't\ntagged, so they don't need it. If you download a new `.dmg` by hand, run it\nagain for that copy.\n\nAfter an update, macOS sometimes keeps showing SpeakoFlow as allowed under\nAccessibility, Microphone, or Screen Recording while no longer honouring it. If\na permission screen keeps waiting, use its **Reset permission** button, then\nswitch SpeakoFlow on again in System Settings.\n\nThe Intel build needs macOS 14 Sonoma or later. It runs on the processor only, so transcription is slower than on Apple silicon, but everything works. CI launches every Intel build on a real Intel Mac before it's released.\n\n- **Arch Linux.** Install`speakoflow-bin` from the AUR, for example with`yay -S speakoflow-bin` .\n- **Debian 13+, Ubuntu 24.04+, Mint 22+, Pop!_OS.** Install the`.deb` , which\nalso adds the app icon and menu entry:\n\n```\nsudo apt install ./SpeakoFlow_*_amd64.deb\n```\n\n- **Other distributions, including Fedora and openSUSE.** Use the AppImage:\nmake it executable with`chmod +x` and run it. Tools like Gear Lever or\nAppImageLauncher add it to your app menu.\n\nBoth packages are built for x86_64 and ARM64 on Ubuntu 24.04, so they need\nglibc 2.39 or newer. That rules out Ubuntu 22.04, Debian 12, Mint 21, and\nRHEL 9 and its rebuilds. There's no `.rpm` yet, because the packaging doesn't\nbundle the speech engine correctly, and a package that installs but can't\ntranscribe would be worse than none.\n\nSpeakoFlow checks for new versions in the background and installs them from\nSettings → About, after verifying each one against the project's signing key.\nThe AUR package updates through your package manager instead. To hear about\nreleases on GitHub, click **Watch → Custom → Releases** at the top of this\npage.\n\n```\ngit clone https://github.com/AbhishekBarali/SpeakoFlow.git\ncd SpeakoFlow\nbun install\nmkdir -p src-tauri/resources/models\ncurl -o src-tauri/resources/models/silero_vad_v4.onnx https://blob.handy.computer/silero_vad_v4.onnx\nbun run tauri dev\n```\n\nOn Arch-based distributions, `bun run install:arch` builds the current checkout\nand installs it under `~/.local` with a desktop entry and a `speak` command.\n[BUILD.md](https://github.com/AbhishekBarali/SpeakoFlow/blob/main/BUILD.md) has the setup for each platform.\n\nThe app is [Tauri 2](https://tauri.app) with a Rust backend and a React and\nTypeScript frontend. Speech runs on transcribe.cpp, whisper.cpp, and ONNX\nRuntime with Silero VAD; the assistant and cleanup on a bundled llama.cpp\nengine or any OpenAI-compatible API; local voices on [Kokoro](https://github.com/hexgrad/kokoro)\nin the app's window and [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)\non the processor; and meeting speaker labels on WeSpeaker voice embeddings.\n\n|  | SpeakoFlow | Wispr Flow | Superwhisper | Handy | \n|---|---|---|---|---|\n| Price | Free | Free up to 2,000 words a week on desktop, then $15/month | Free tier, Pro at $8.49/month or $249.99 once | Free | \n| Source code | Open (MIT) | Closed | Closed | Open (MIT) | \n| Linux | Yes | No | No | Yes | \n| Transcribes offline | Yes | No | Yes | Yes | \n| AI assistant | Yes | Yes | No | No | \n\n<sub>Prices and platforms from each product's own site, checked October 2026.</sub>\n\nSpeakoFlow's dictation core comes from [Handy](https://github.com/cjpais/Handy),\nwhich is a good choice if dictation is all you need. More detail:\n[SpeakoFlow vs Wispr Flow](https://www.speakoflow.com/blog/speakoflow-vs-wispr-flow)\nand [free and open-source Wispr Flow alternatives](https://www.speakoflow.com/blog/best-free-open-source-wispr-flow-alternatives).\n\nThe common problems are below. For anything else, see the\n[troubleshooting docs](https://www.speakoflow.com/docs/reference/troubleshooting)\nor [open an issue](https://github.com/AbhishekBarali/SpeakoFlow/issues).\n\n## **macOS: a meeting only records my side of the call**\n\nmacOS gives apps no direct way to record the sound your computer plays. Windows has WASAPI loopback and Linux has your PulseAudio or PipeWire monitor source, but a Mac needs a virtual audio device in between.\n\nInstall a loopback driver such as [BlackHole](https://github.com/ExistentialAudio/BlackHole),\ncreate a Multi-Output Device in Audio MIDI Setup that sends sound to both your\nspeakers and BlackHole, and make it your output. SpeakoFlow can then record the\nother side of the call. Your microphone is recorded either way.\n\n## **Linux: the recording overlay won't stay on top of other apps**\n\nA window can only float above the others on Linux through the `wlr-layer-shell`\nprotocol (wlroots compositors like Sway and Hyprland, and KDE Plasma) or X11\n\"keep above\" stacking. Native GNOME on Wayland supports neither, so when\nSpeakoFlow detects it, it runs under XWayland, where the overlay floats\nnormally. That needs no setup, and X11 and KDE or wlroots Wayland work as they\nare.\n\n- To force native Wayland anyway, launch with `SPEAKOFLOW_ALLOW_WAYLAND=1` . The\noverlay may not stay on top.\n- If the overlay misbehaves under a layer-shell compositor, launch with\n`SPEAKOFLOW_NO_GTK_LAYER_SHELL=1` .\n\n## **Linux: shortcuts do nothing and the log repeats \"Permission denied\"**\n\nA log full of `rdev grab error: ... PermissionDenied` means the app can't read\nyour input devices. This only affects the **SpeakoFlow Keys** keyboard engine,\nwhich reads `/dev/input/event*` (needs your user in the `input` group) and\nre-sends keys through `/dev/uinput` (root-only by default on many\ndistributions, Ubuntu included, so the group alone is not enough). Tauri is the default engine on Linux, so you'd only\nsee this after switching. The **Shortcuts** card on the Home page says when\nthis is the case and gives the exact command.\n\n- Grant both, then log out and back in:\n\n```\nsudo usermod -aG input \"$USER\"\necho 'KERNEL==\"uinput\", GROUP=\"input\", MODE=\"0660\"' | sudo tee /etc/udev/rules.d/70-speakoflow-uinput.rules\nsudo udevadm control --reload && sudo udevadm trigger /dev/uinput\n```\n\n- Or switch the keyboard engine back to **Tauri** in Settings → Advanced. It\nneeds no permissions but registers shortcuts through X11, so on native\nWayland it only hears them while an X11 window has focus.\n\nOn Wayland the dependable option is a shortcut owned by your desktop. On a\nWayland session the Shortcuts card shows a **Set up** button that lists the\ncommand for each action, ready to copy. Add a custom shortcut in GNOME or KDE\nsettings, or a `bind` line in Sway or Hyprland, that runs\n`speakoflow --toggle-transcription` (for an AppImage, its path followed by the\nsame flag). `--toggle-post-process`, `--toggle-assistant`, `--toggle-call`\n(start or end a conversation), and `--cancel` work the same way. These start\nwith one press and stop with the next, like **Tap to toggle**.\n\n## **Linux: the app crashes on a touchpad pinch-to-zoom**\n\nA crash with `Received invalid message: 'DrawingArea_CommitTransientZoom'` in\nthe log is a WebKitGTK bug that affects many apps built on it, tracked in\n[tauri#13115](https://github.com/tauri-apps/tauri/issues/13115) and\n[wry#544](https://github.com/tauri-apps/wry/issues/544). Until it's fixed\nupstream, avoid pinching inside the window. Updating `webkit2gtk-4.1` to the\nlatest version can help.\n\n- Code signing for Windows and macOS\n- More one-click local models\n- More community translations\n- Dictation tuned for agentic coding\n- Help writing prompts: describe what you want to build and get a solid prompt back\n- Voice commands that take actions for you\n\nContributions are welcome. [CONTRIBUTING.md](https://github.com/AbhishekBarali/SpeakoFlow/blob/main/CONTRIBUTING.md) explains how to\nget started, and [CONTRIBUTING_TRANSLATIONS.md](https://github.com/AbhishekBarali/SpeakoFlow/blob/main/CONTRIBUTING_TRANSLATIONS.md)\ncovers translating the app.\n\nFound a bug or have an idea? Use **Send feedback** in the app (the **?** next\nto Settings), or [open an issue](https://github.com/AbhishekBarali/SpeakoFlow/issues).\n\nSpeakoFlow is released under the [MIT License](https://github.com/AbhishekBarali/SpeakoFlow/blob/main/LICENSE).\n\nThe dictation core comes from [Handy](https://github.com/cjpais/Handy) by CJ\nPais, used under the MIT licence. Thanks to CJ for making it open. The\nassistant, conversations, meetings, screen vision, Generate with Flow,\ntranslation, spoken replies, and memory are SpeakoFlow's own.\n\nThanks also to [Tauri](https://tauri.app), whisper.cpp, llama.cpp, ONNX Runtime,\n[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx), Silero VAD, WeSpeaker,\n[Kokoro](https://github.com/hexgrad/kokoro), and Kyutai's Pocket TTS.\n\nCloud transcription uploads are compressed with the [LAME](https://lame.sourceforge.io)\nMP3 encoder, via [mp3lame-encoder](https://github.com/DoumanAsh/mp3lame-encoder).\nBoth are LGPL-3.0 and are statically linked; their source, and this app's, are\npublic, so a build against a modified LAME is always possible.\n\nMade by [Abhishek Barali](https://github.com/AbhishekBarali) · [speakoflow.com](https://www.speakoflow.com)", "url": "https://wpnews.pro/news/speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux", "canonical_source": "https://github.com/AbhishekBarali/SpeakoFlow", "published_at": "2026-10-09 11:51:59+00:00", "updated_at": "2026-10-09 12:24:12.775820+00:00", "lang": "en", "topics": ["ai-tools", "ai-products", "natural-language-processing", "generative-ai", "ai-agents"], "entities": ["SpeakoFlow", "Parakeet", "Nemotron", "Whisper", "Ollama", "LM Studio", "OpenAI", "ElevenLabs"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux", "markdown": "https://wpnews.pro/news/speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux.md", "text": "https://wpnews.pro/news/speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux.txt", "jsonld": "https://wpnews.pro/news/speakoflow-free-voice-dictation-and-ai-assistant-for-windows-macos-and-linux.jsonld"}}