Show HN: TuxWhisper – Offline push-to-talk dictation for Linux, works on Wayland A developer released TuxWhisper, an open-source offline push-to-talk dictation tool for Linux that runs speech recognition locally via faster-whisper and types text into any app on X11 and Wayland, including GNOME on Wayland. The tool transcribes in about half a second after key release on an NVIDIA GPU, requires about 1.5 GB of free VRAM (or CPU mode), and optionally routes transcripts through a local Ollama model for punctuation cleanup. TuxWhisper requires Linux with systemd, Python 3.12, uv, and xclip or wl-clipboard, and its first start downloads a roughly 1.5 GB Whisper model to ~/.cache/huggingface. Offline, hotkey-driven dictation for Linux. Press a key, speak, and your words are typed wherever your cursor is: a browser text box, an LLM chat prompt, your editor or your terminal. Speech recognition runs locally with faster-whisper https://github.com/SYSTRAN/faster-whisper ; nothing is sent to the cloud. - Types anywhere. Works in any app that accepts paste, on X11 and Wayland, including GNOME on Wayland, where most typing tools don't work. - Fast. About half a second from releasing the key to text on screen, with an NVIDIA GPU. - Tap or hold. Tap to start and stop, or hold the key to talk and release it to transcribe. - Optional LLM cleanup. A second key passes the transcript through a local Ollama https://ollama.com model to fix punctuation and remove "um"s, without answering the questions you dictate. - Builds a voice dataset. One key saves the last recording with its transcript, ready for fine-tuning. - Leaves your clipboard alone. Whatever you had copied is restored after each paste. - Works on any keyboard layout. It pastes with Shift+Insert, which doesn't depend on QWERTY. Developed and tested on GNOME Wayland with an NVIDIA GPU, and in CPU mode. Other desktops should work but are untested; see Other desktops other-desktops . | Key | Action | |---|---| | F5 tap | Start recording. Tap again to stop, transcribe and paste. | | F5 hold | Push-to-talk: speak while holding, release to transcribe and paste. | | F3 | Same as F5, but the transcript is cleaned up by a local LLM before pasting. | | F4 | Save the last take audio + transcript to recordings/ . | The keys are only suggestions; you choose them when you set up the shortcuts. - Wait for the "πŸŽ™ Listening…" notification before speaking. The microphone takes about 0.3 s to open, and anything said before that is lost. - Very short clips are ignored. Anything under 0.3 s produces no text. The keys run small commands, which you can also use from a terminal or script: ./asr toggle start / stop dictation ./asr rewrite start / stop dictation with LLM cleanup ./asr save save the last take ./asr serve run the daemon in the foreground normally systemd does this Requirements: Linux with systemd, Python 3.12, uv https://docs.astral.sh/uv/ , an NVIDIA GPU with about 1.5 GB of free VRAM or see CPU mode cpu-mode , and xclip or wl-clipboard on Wayland without XWayland . 1. Clone and install git clone https://github.com/ialmajai/tuxwhisper.git ~/tuxwhisper cd ~/tuxwhisper uv venv --python 3.12 .venv uv pip install --python .venv -r requirements-cuda.txt no NVIDIA GPU: requirements.txt 2. Allow the virtual keyboard needs sudo once; read the security note security-note echo 'KERNEL=="uinput", TAG+="uaccess", OPTIONS+="static node=uinput"' | sudo tee /etc/udev/rules.d/60-uinput.rules sudo udevadm control --reload && sudo udevadm trigger --name-match=uinput 3. Start the daemon at login cp tuxwhisper.service ~/.config/systemd/user/ systemctl --user enable --now tuxwhisper The service expects the repo at ~/tuxwhisper . If it's elsewhere, edit ExecStart in the copied file. The first start downloads the Whisper model about 1.5 GB to ~/.cache/huggingface , so it takes a while; after that, startup takes about 15 s. 4. Bind the keys. On GNOME: Settings β†’ Keyboard β†’ Keyboard Shortcuts β†’ Custom Shortcuts. | Shortcut | Command | |---|---| | F5 | /home/