VOICEIO_ GitHub
// guide · wayland

Whisper dictation on Wayland

To dictate with Whisper on Wayland, you need two pieces: a local Whisper decoder, and a way to put text into other apps that Wayland allows. voiceio uses faster-whisper for the first and an IBus input method for the second, so words appear in the focused app as an underlined live preview while you speak and are committed when you stop. It works on GNOME, KDE Plasma, Hyprland and sway, and on X11.

IBus preeditfaster-whisperwtype · ydotool fallback
// how to

Set it up

  1. Install the system packages (a C toolchain, PortAudio, IBus). On Debian or Ubuntu:
    $ sudo apt install pipx build-essential python3-dev \
        portaudio19-dev ibus gir1.2-ibus-1.0 python3-gi
    Fedora, Arch and NixOS commands are in the README.
  2. Install voiceio with the desktop extra (microphone, hotkeys, tray):
    $ pipx install 'python-voiceio[desktop]'
  3. Run the guided setup. It downloads the speech model, sets your hotkey and installs a systemd user service so voiceio starts on login:
    $ voiceio setup
  4. Dictate. Press the hotkey you picked during setup in any app, speak, and press it again. voiceio doctor shows what works; voiceio doctor --fix repairs what it can.

Hotkeys on Wayland

The evdev backend reads the keyboard directly. voiceio doctor --fix installs a udev rule that gives the logged-in user access, with no logout needed. On GNOME, setup can add a shortcut instead. On Hyprland or sway, bind voiceio toggle:

# ~/.config/hypr/hyprland.conf
bind = CTRL ALT, V, exec, voiceio toggle
# ~/.config/sway/config
bindsym Ctrl+Alt+v exec voiceio toggle
// typing

How the text gets in

  • IBus (preferred). voiceio installs its own IBus engine and activates it only while you dictate. The preview is shown as preedit text and committed at the end.
  • fcitx5. Left alone; final text is typed with wtype or ydotool.
  • wtype, ydotool, xdotool. Keystroke fallbacks, picked by probing what works. doctor --fix installs a uaccess rule for /dev/uinput so ydotool works.
  • Clipboard. For terminals that ignore input methods.
// the model

Which Whisper

The default is Whisper small through faster-whisper, about 5x realtime on a laptop CPU. medium is better on names and slower; large-v3-turbo wants a GPU. Long dictations are finalized in pieces while you talk, so stopping a multi-minute note does not mean waiting for the whole recording to decode again. Read how voiceio relates to the model itself in voiceio and Whisper.

// alternatives

Other Whisper tools on Wayland

Voxtype also runs Whisper (via whisper.cpp) and other engines on Wayland, typing with wtype, dotool or ydotool. nerd-dictation uses VOSK rather than Whisper. Talon does not support Wayland.

// faq

Questions

Why is dictation hard on Wayland?

Wayland does not let one app type into another or read global hotkeys the way X11 did. Tools work around it with input methods (IBus, fcitx5), virtual keyboards (wtype, ydotool, dotool) or the compositor's own key bindings.

Why IBus instead of ydotool or wtype?

An input method can show an underlined preview that is corrected in place and then committed. Keystroke tools can only type and backspace, which is slower and can drop characters on rapid corrections. voiceio uses IBus first and falls back to wtype or ydotool.

Does it work with fcitx5 (Omarchy)?

Yes. If fcitx5 is your input method, voiceio leaves it alone and types the words with wtype or ydotool as you speak, without the in-place preview.

How do I bind the hotkey on Hyprland or sway?

Use the built-in evdev hotkey, or bind voiceio toggle in your compositor, e.g. bind = CTRL ALT, V, exec, voiceio toggle on Hyprland.

// get it

Try voiceio

Free and MIT licensed. Install it with pipx on Linux, or try the dictation in your browser first: the demo on the home page runs a small speech model inside the tab.