voiceio and Whisper
voiceio does not compete with Whisper; it is built on it. OpenAI's Whisper is a speech recognition model (MIT licensed), and whisper.cpp and faster-whisper are fast ways to run it. None of them types into other apps: they turn audio into text. voiceio is the dictation layer on top: a hotkey, a microphone, streaming decode with a live preview in the focused app, vocabulary and corrections, and fine-tuning Whisper on your own voice. It runs faster-whisper by default and can use a whisper.cpp server.
Which should you use?
Choose Whisper / libraries if…
- You want to transcribe audio files or build your own pipeline: use Whisper, faster-whisper or whisper.cpp directly.
- You need a specific runtime or platform: whisper.cpp runs on macOS, iOS, Android, Windows, WebAssembly and more, with Metal, CUDA and Vulkan backends.
- You want a quick microphone test: whisper.cpp's stream example transcribes the mic continuously.
Choose voiceio if…
- You want to dictate into any app on Linux with a hotkey, without writing glue code.
- You want a live preview while you speak and fast commits for long dictations.
- You want vocabulary hotwords, correction rules and spoken numbers handled for you.
- You want Whisper fine-tuned on your voice, evaluated on your own held-out clips.
Feature comparison
| voiceio 1.5 | Whisper / libraries | |
|---|---|---|
| Platforms | Linux: Wayland and X11 (GNOME, KDE, Hyprland, sway, i3). Untested best-effort path on Windows/macOS | Libraries and CLIs: Python (Whisper, faster-whisper); C/C++ for many platforms (whisper.cpp) |
| Price | Free, MIT | Free, MIT (code and weights for Whisper) |
| Where speech is decoded | On your computer. Optional text-only cloud cleanup, off by default | Wherever you run it |
| Speech models | faster-whisper small (default), medium, large-v3-turbo, distil-large-v3; experimental Parakeet and whisper.cpp server | tiny, base, small, medium, large-v3, large-v3-turbo (and .en variants) |
| Open source | Yes (MIT) | Yes (MIT) |
| Live typing into apps | Yes: underlined preview in the focused app through IBus, corrected in place, committed when you stop | No typing into apps; whisper.cpp has a mic stream example |
| Custom vocabulary | Hotwords ranked by use (about 35 per decode) plus find-and-replace corrections | initial_prompt (all); hotwords (faster-whisper) |
| Learns from you | Yes: optional LoRA fine-tune on your own recordings, on CPU, kept only if it wins on held-out clips | Not built in; fine-tuning is up to you |
| Voice commands | Dictation commands only ("new line", "scratch that", "correct that") | None |
Whisper facts are from its official site, docs or repository, read on 2026-10-10 (sources below). Prices are as listed on that date and can change. voiceio facts are from its README. Spotted something out of date? Open an issue.
Questions
Is voiceio just Whisper?
voiceio uses Whisper models for recognition. What it adds is the dictation app: audio capture, hotkeys, streaming decode, typing into the focused app through IBus, text cleanup, vocabulary ranking, corrections and a fine-tuning pipeline.
Which Whisper model does voiceio use?
Whisper small through faster-whisper by default. You can switch to medium, large-v3-turbo or distil-large-v3 with voiceio models use, or point it at a whisper.cpp server.
Why does voiceio still use Whisper rather than a faster model?
Most dictation errors are names and jargon, so the ability to bias the model matters more than raw speed. voiceio's contributors measured Parakeet as faster, but its biasing mode lost speech on real clips; the reasoning is in CONTRIBUTING.md.
Can I use Whisper for dictation on Wayland?
Yes. See Whisper dictation on Wayland.
Try voiceio
Free and MIT licensed. Install it with pipx on Linux, or try the dictation in your browser first: the demo on the home page runs a small speech model inside the tab.