VOICEIO_ GitHub
// guide · linux

Voice dictation on Linux

To dictate on Linux, install a speech-to-text tool that types into the focused window. voiceio does this on Wayland and X11: press a hotkey, speak, and the text streams into whatever app has focus, with a live underlined preview through IBus. Speech is decoded on your own computer with Whisper, so it works offline and nothing is uploaded. It is free and MIT licensed.

Wayland + X11runs locallyMITPython 3.11+
// how to

Set up dictation in four steps

  1. Install the system packages (a C toolchain, PortAudio, IBus). On Debian or Ubuntu:
    $ sudo apt install pipx build-essential python3-dev \
        portaudio19-dev ibus gir1.2-ibus-1.0 python3-gi
    Fedora, Arch and NixOS commands are in the README.
  2. Install voiceio with the desktop extra (microphone, hotkeys, tray):
    $ pipx install 'python-voiceio[desktop]'
  3. Run the guided setup. It downloads the speech model, sets your hotkey and installs a systemd user service so voiceio starts on login:
    $ voiceio setup
  4. Dictate. Press the hotkey you picked during setup in any app, speak, and press it again. voiceio doctor shows what works; voiceio doctor --fix repairs what it can.

Prefer to hand it off? The install section has a prompt for Claude Code, Codex or any coding agent, which follows the runbook in llms.txt.

// desktops

What works where

DesktopText injectionLive previewStatus (README)
GNOME (Wayland)IBus, ydotool fallbackYesTested daily
GNOME (X11)IBus, xdotool fallbackYesSupported
KDE Plasma (Wayland)IBus, ydotool fallbackYes, with IBusShould work
Hyprland, swayIBus, else wtype, then ydotoolYes, with IBusShould work
i3, other X11IBus, xdotoolYes, with IBusShould work

From the voiceio README. If fcitx5 is your input method, voiceio leaves it alone and types the final text with wtype or ydotool.

// accuracy

Getting names and jargon right

Most dictation errors are proper nouns: people, products, your stack. voiceio has three tools for them, from cheapest to strongest:

  • Vocabulary. voiceio vocab add Kalshi passes your terms to Whisper as hotwords, ranked by how often you use them. Whisper's hotword channel fits about 35 terms per decode.
  • Corrections. Find-and-replace rules for words that still come out wrong, or say "correct that" right after a mistake.
  • Fine-tuning. voiceio learn trains a LoRA fine-tune on your own kept recordings, on CPU, and switches to it only when it beats the current model on clips it never trained on. voiceio learn schedule on runs it when the machine is idle.
// alternatives

Other ways to dictate on Linux

  • Voxtype: a local, MIT-licensed Rust dictation tool for Linux with many speech engines and GPU builds.
  • nerd-dictation: a single-file Python script using the offline VOSK engine.
  • Talon: hands-free computer control by voice; on Linux it runs on X11 only.
  • Whisper, whisper.cpp, faster-whisper: speech models and libraries that transcribe audio but do not type into apps by themselves.

Aqua Voice, Wispr Flow and Dragon do not offer Linux apps, and Superwhisper has a prerelease Linux build for Omarchy only (all checked 2026-10-10). See the comparison hub.

// faq

Questions

Does voice dictation work on Wayland?

Yes. voiceio types through IBus, which works on Wayland and X11, including GNOME, KDE Plasma, Hyprland and sway. Where IBus is not available it falls back to wtype, ydotool or xdotool. See Whisper dictation on Wayland.

Is there a free dictation app for Linux?

voiceio is free and MIT licensed. Other free options include Voxtype (MIT) and nerd-dictation (GPL-3.0).

Does it need an internet connection?

Only to download the speech model the first time. After that, speech is decoded on your own computer and nothing is uploaded. See offline speech to text.

Which Linux distributions are supported?

voiceio is a Python package (3.11+) installed with pipx. The README has package commands for Debian/Ubuntu, Fedora and Arch, and a flake for NixOS (not yet built, per the README).

Can I dictate in languages other than English?

The default Whisper models are multilingual; set model.language to a language code or auto. Fine-tuning and the vocabulary tools have been tested mostly on English.

// get it

Try voiceio

Free and MIT licensed. Install it with pipx on Linux, or try the dictation in your browser first: the demo on the home page runs a small speech model inside the tab.