{"slug": "show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland", "title": "Show HN: TuxWhisper – Offline push-to-talk dictation for Linux, works on Wayland", "summary": "A developer released TuxWhisper, an open-source offline push-to-talk dictation tool for Linux that runs speech recognition locally via faster-whisper and types text into any app on X11 and Wayland, including GNOME on Wayland. The tool transcribes in about half a second after key release on an NVIDIA GPU, requires about 1.5 GB of free VRAM (or CPU mode), and optionally routes transcripts through a local Ollama model for punctuation cleanup. TuxWhisper requires Linux with systemd, Python 3.12, uv, and xclip or wl-clipboard, and its first start downloads a roughly 1.5 GB Whisper model to ~/.cache/huggingface.", "body_md": "**Offline, hotkey-driven dictation for Linux.** Press a key, speak, and your words are\ntyped wherever your cursor is: a browser text box, an LLM chat prompt, your editor or\nyour terminal.\n\nSpeech recognition runs locally with [faster-whisper](https://github.com/SYSTRAN/faster-whisper);\nnothing is sent to the cloud.\n\n- **Types anywhere.** Works in any app that accepts paste, on X11 and Wayland, including\nGNOME on Wayland, where most typing tools don't work.\n- **Fast.** About half a second from releasing the key to text on screen, with an NVIDIA GPU.\n- **Tap or hold.** Tap to start and stop, or hold the key to talk and release it to transcribe.\n- **Optional LLM cleanup.** A second key passes the transcript through a local[Ollama](https://ollama.com) model to fix punctuation and remove \"um\"s, without answering\nthe questions you dictate.\n- **Builds a voice dataset.** One key saves the last recording with its transcript, ready\nfor fine-tuning.\n- **Leaves your clipboard alone.** Whatever you had copied is restored after each paste.\n- **Works on any keyboard layout.** It pastes with Shift+Insert, which doesn't depend on\nQWERTY.\n\nDeveloped and tested on GNOME (Wayland) with an NVIDIA GPU, and in CPU mode. Other\ndesktops should work but are untested; see [Other desktops](#other-desktops).\n\n| Key | Action | \n|---|---|\n| **F5** (tap) | Start recording. Tap again to stop, transcribe and paste. | \n| **F5** (hold) | Push-to-talk: speak while holding, release to transcribe and paste. | \n| **F3** | Same as F5, but the transcript is cleaned up by a local LLM before pasting. | \n| **F4** | Save the last take (audio + transcript) to `recordings/` . | \n\nThe keys are only suggestions; you choose them when you set up the shortcuts.\n\n- **Wait for the \"🎙 Listening…\" notification before speaking.** The microphone takes about\n0.3 s to open, and anything said before that is lost.\n- **Very short clips are ignored.** Anything under 0.3 s produces no text.\n\nThe keys run small commands, which you can also use from a terminal or script:\n\n```\n./asr toggle    # start / stop dictation\n./asr rewrite   # start / stop dictation with LLM cleanup\n./asr save      # save the last take\n./asr serve     # run the daemon in the foreground (normally systemd does this)\n```\n\n**Requirements:** Linux with systemd, Python 3.12, [uv](https://docs.astral.sh/uv/), an\nNVIDIA GPU with about 1.5 GB of free VRAM (or see [CPU mode](#cpu-mode)), and `xclip`\n(or `wl-clipboard` on Wayland without XWayland).\n\n**1. Clone and install**\n\n```\ngit clone https://github.com/ialmajai/tuxwhisper.git ~/tuxwhisper\ncd ~/tuxwhisper\nuv venv --python 3.12 .venv\nuv pip install --python .venv -r requirements-cuda.txt   # no NVIDIA GPU: requirements.txt\n```\n\n**2. Allow the virtual keyboard** (needs sudo once; read the [security note](#security-note))\n\n```\necho 'KERNEL==\"uinput\", TAG+=\"uaccess\", OPTIONS+=\"static_node=uinput\"' | sudo tee /etc/udev/rules.d/60-uinput.rules\nsudo udevadm control --reload && sudo udevadm trigger --name-match=uinput\n```\n\n**3. Start the daemon at login**\n\n```\ncp tuxwhisper.service ~/.config/systemd/user/\nsystemctl --user enable --now tuxwhisper\n```\n\nThe service expects the repo at `~/tuxwhisper`. If it's elsewhere, edit `ExecStart` in the\ncopied file. The first start downloads the Whisper model (about 1.5 GB) to\n`~/.cache/huggingface`, so it takes a while; after that, startup takes about 15 s.\n\n**4. Bind the keys.** On GNOME: Settings → Keyboard → Keyboard Shortcuts → Custom Shortcuts.\n\n| Shortcut | Command | \n|---|---|\n| F5 | `/home/<you>/tuxwhisper/asr toggle` | \n| F3 | `/home/<you>/tuxwhisper/asr rewrite` | \n| F4 | `/home/<you>/tuxwhisper/asr save` | \n\nUse the full path; shortcut commands don't expand `~`. For other desktops, see\n[Other desktops](#other-desktops).\n\n**5. Optional: LLM cleanup.** Install Ollama ([Linux install guide](https://docs.ollama.com/linux)) and the cleanup model:\n\n```\ncurl -fsSL https://ollama.com/install.sh | sh\nollama pull llama3.2\n```\n\nClick into any text box, press F5 and speak.\n\nTip\n\nBinding F5 overrides page refresh in browsers. Ctrl+R still refreshes.\n\nF3 sends the transcript to a local Ollama model (`llama3.2` by default). The model fixes\npunctuation and capitalization and removes filler words (\"um\", \"uh\", \"like\"), repeated words\nand false starts.\n\n```\nWhisper:  um so can you like uh explain how the the attention mechanism works\nPasted:   So can you explain how the attention mechanism works?\n```\n\n- **It cleans; it doesn't answer.** A dictated question is pasted as a question, not\nanswered.\n- **Speed:** about 0.2–1 s once the model is loaded. The model unloads after 30 minutes\nidle to free GPU memory, so the next F3 takes a few seconds longer.\n- **Always pastes something:** if Ollama is unreachable, or the model doesn't fit in free\nGPU memory, the raw transcript is pasted and a notification says why.\n- **Not perfect:** it sometimes drops real words. Use F5 when the exact wording matters.\n\nOllama can run on another machine; point `OLLAMA_URL` at it (see\n[Configuration](#configuration)).\n\nF4 saves the most recent take to `recordings/`:\n\n```\nrecordings/\n├── 20260101-120000.wav    # 16 kHz, mono, 16-bit\n└── metadata.csv           # file_name,transcription\n```\n\nThis is the Hugging Face `audiofolder` layout, ready for fine-tuning a speech model.\n\n- Only the **latest** take can be saved, and only once. Pressing F4 again shows \"Nothing to\nsave\" until you dictate again.\n- The saved transcript is Whisper's **raw** output, even for F3 takes, because that's what\nmatches the audio. Corrections you make in the text box aren't saved; edit`metadata.csv` if needed.\n\nSettings are environment variables. Add them to the `[Service]` section of\n`~/.config/systemd/user/tuxwhisper.service`, e.g. `Environment=ASR_LANGUAGE=en`.\n\n| Variable | Default | Meaning | \n|---|---|---|\n| `ASR_LANGUAGE` | auto-detect | Language code, e.g. `en` . Setting it avoids misdetection on short clips. | \n| `ASR_DEVICE` | `cuda` | `cuda` or`cpu` | \n| `ASR_MODEL` | `large-v3-turbo` (GPU),`small` (CPU) | Any faster-whisper model name or path | \n| `ASR_PROMPT` | a short punctuated sentence | Style example for Whisper; keeps capitals and punctuation. Set to empty to disable. | \n| `ASR_PASTE_KEY` | `shift+insert` | `shift+insert` ,`ctrl+v` or`ctrl+shift+v` | \n| `ASR_REWRITE_MODEL` | `llama3.2` | Ollama model used by F3 | \n| `OLLAMA_URL` | `http://localhost:11434` | Ollama server used by F3 | \n\nThen apply the changes:\n\n```\nsystemctl --user daemon-reload && systemctl --user restart tuxwhisper\n```\n\nAudio comes from your default input device, which you can change in your desktop's sound settings.\n\nWithout an NVIDIA GPU, install from `requirements.txt` and set `Environment=ASR_DEVICE=cpu`.\nTuxWhisper then uses the smaller `small` model: a 3–4 s clip takes about 1.4 s on a recent\ndesktop CPU.\n\nThe default, Shift+Insert, pastes in browsers, editors and terminals on any keyboard\nlayout. If a particular app doesn't paste, try `ctrl+v`, which works in most GUI apps but\nnot in terminals.\n\nBind the same commands in your desktop's shortcut settings:\n\n| Desktop | Binding | Hold to talk | \n|---|---|---|\n| GNOME | Settings → Keyboard → Custom Shortcuts | ✅ Yes | \n| KDE Plasma | System Settings → Shortcuts → Add New → Command | ❔ Untested | \n| Hyprland | `binde = , F5, exec, ~/tuxwhisper/asr toggle` | ✅ Should work ( `binde` repeats while held) | \n| Sway | `bindsym F5 exec ~/tuxwhisper/asr toggle` | ❌ Tap only | \n| i3 | `bindsym F5 exec --no-startup-id ~/tuxwhisper/asr toggle` | ❌ Tap only | \n\nWhere holding doesn't work, tap to start and tap again to stop.\n\nThe systemd service starts with `graphical-session.target`, which some window managers (i3,\nor Sway without extra setup) never start. On those, skip step 3 and start the daemon from\nyour WM config instead, e.g. `exec ~/tuxwhisper/asr serve`.\n\n```\n F5 / F3 / F4 ──▶ asr toggle|rewrite|save ──▶ Unix socket ──▶ asr serve (daemon)\n                                                                 │\n     mic ──▶ record ──▶ faster-whisper ──▶ (Ollama cleanup) ─────┤\n                                                                 ▼\n        your app ◀── Shift+Insert (virtual keyboard) ◀── clipboard\n```\n\n- **Daemon:**`asr serve` keeps the Whisper model loaded and listens on a Unix socket in`$XDG_RUNTIME_DIR` , which only your user can access.\n- **Typing:** Wayland doesn't let apps type into other windows, so the text goes on the\nclipboard (`xclip` , or`wl-clipboard` without XWayland) and a virtual keyboard\n(`/dev/uinput` ) presses the paste key. The previous clipboard is restored about 0.3 s\nlater (text only; a copied image is lost).\n- **Hold detection:** while a key is held, the desktop re-runs the shortcut on auto-repeat.\nA burst of presses counts as a held key, and recording stops when the repeats end.\n\nWarning\n\nStep 2 gives your user write access to `/dev/uinput`. TuxWhisper needs it to press the\npaste key, but it also lets **any program you run create a virtual keyboard or mouse and\nsend input to any window**, including terminals and password prompts. Tools like\n`ydotool` require the same access.\n\nThe `uaccess` tag limits this to the user logged in at the active local session, and only\nwhile that session is active. Other user accounts don't get access, but every process\nrunning as you does, including ones started over SSH while you're logged in.\n\nIf you don't want this, skip step 2: transcription still works, but the text isn't pasted.\nTo undo it later, run `sudo rm /etc/udev/rules.d/60-uinput.rules` and reboot.\n\n```\nsystemctl --user status tuxwhisper    # is the daemon running?\njournalctl --user -u tuxwhisper -f    # live log: each transcription, timings and errors\nsystemctl --user restart tuxwhisper   # restart (the model takes ~15 s to load)\n```\n\n## **Nothing happens when I press the key**\n\nCheck `systemctl --user status tuxwhisper`. Right after login, the model may still be loading\n(watch for `ready` in the log). If you see \"daemon is not running\", start it with\n`systemctl --user start tuxwhisper`. Also check that the shortcut command uses the full path.\n\n## **It transcribes but nothing is pasted**\n\nRun `getfacl /dev/uinput`; it should list `user:<you>:rw-`. If not, check the udev rule from\nstep 2, including the file name (it must start with a number below 73).\n\n## **It doesn't paste in one particular app**\n\nThat app may not treat Shift+Insert as paste. Try `ASR_PASTE_KEY=ctrl+v`.\n\n## **My old clipboard is pasted instead of the transcript**\n\nThe app read the clipboard after it had already been restored. Increase the `0.3` s delay in\n`paste()` in `asr.py`.\n\n## **The daemon exits with \"cuda unavailable\"**\n\nThe GPU or the CUDA libraries couldn't be used. Check that you installed from\n`requirements-cuda.txt` and that `nvidia-smi` works, or switch to [CPU mode](#cpu-mode).\n\n## **CUDA out of memory**\n\nAnother program is using the GPU. Free some memory, or run Ollama for F3 on another machine\nwith `OLLAMA_URL`.\n\n## **F3 pastes the raw transcript**\n\nThe notification says why: either Ollama isn't reachable (check `ollama list` and\n`OLLAMA_URL`), or there isn't enough free GPU memory for the cleanup model.\n\n```\nasr                    launcher wrapper (use this, not asr.py directly)\nasr.py                 daemon and client commands\ntuxwhisper.service     systemd user service template\nrequirements.txt       Python dependencies\nrequirements-cuda.txt  + NVIDIA CUDA libraries\nrecordings/            saved takes (created by F4, not committed)\n```\n\n[MIT](https://github.com/ialmajai/tuxwhisper/blob/main/LICENSE) © 2026 Ibrahim Almajai", "url": "https://wpnews.pro/news/show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland", "canonical_source": "https://github.com/ialmajai/tuxwhisper", "published_at": "2026-10-07 20:53:27+00:00", "updated_at": "2026-10-07 21:19:56.549762+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools", "ai-products"], "entities": ["TuxWhisper", "faster-whisper", "Ollama", "GNOME", "Wayland", "Linux", "NVIDIA", "llama3.2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland", "markdown": "https://wpnews.pro/news/show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland.md", "text": "https://wpnews.pro/news/show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland.txt", "jsonld": "https://wpnews.pro/news/show-hn-tuxwhisper-offline-push-to-talk-dictation-for-linux-works-on-wayland.jsonld"}}