Offline speech to text
Offline speech to text means the speech model runs on your own computer, so audio never goes to a server and dictation works without a connection. On Linux, voiceio does this with Whisper: after a one-time model download it decodes your speech locally, types the text into the focused app, and has no account and no telemetry.
Set up offline dictation
- Install the system packages (a C toolchain, PortAudio, IBus). On Debian or Ubuntu:
Fedora, Arch and NixOS commands are in the README.$ sudo apt install pipx build-essential python3-dev \ portaudio19-dev ibus gir1.2-ibus-1.0 python3-gi - Install voiceio with the desktop extra (microphone, hotkeys, tray):
$ pipx install 'python-voiceio[desktop]' - Run the guided setup. It downloads the speech model, sets your hotkey and installs a systemd user service so voiceio starts on login:
$ voiceio setup - Dictate. Press the hotkey you picked during setup in any app, speak, and press it again.
voiceio doctorshows what works;voiceio doctor --fixrepairs what it can.
Once the model is downloaded, unplug the network and dictate: nothing changes. To pick a different local model:
$ voiceio models list # known models and which is active
$ voiceio models use medium # better on names, about 2.5x slower
Local models voiceio can use
| Model | Runtime | Notes (README) |
|---|---|---|
| small | faster-whisper | Default. About 5x realtime on a laptop CPU |
| medium | faster-whisper | Better on names, about 2.5x slower |
| large-v3-turbo | faster-whisper | Most accurate Whisper option for dictation; wants a GPU |
| distil-large-v3 | faster-whisper | English only |
| Parakeet TDT 0.6b v3 | sherpa-onnx | Experimental. Fast, no vocabulary biasing yet |
| whisper.cpp server | whisper.cpp | Experimental. E.g. a Vulkan build on an AMD iGPU |
voiceio ships no model weights; each model downloads from its own source under its own license.
What stays on your disk
Everything voiceio keeps is plain files in ~/.config/voiceio, ~/.local/state/voiceio and ~/.local/share/voiceio, created readable only by you. Recordings are off unless you turn them on; you need them only for fine-tuning. Logs record sizes and timings, never your words.
For a phone or another device, python -m voiceio.server serves the same engine over HTTP and WebSocket on a machine you control. It binds to localhost by default.
Other offline options
- Voxtype: a local, MIT-licensed Rust dictation tool for Linux with many speech engines and GPU builds.
- nerd-dictation: a single-file Python script using the offline VOSK engine.
- Talon: hands-free computer control by voice; on Linux it runs on X11 only.
- Whisper, whisper.cpp, faster-whisper: speech models and libraries that transcribe audio but do not type into apps by themselves.
Aqua Voice, Wispr Flow and Dragon do not offer Linux apps, and Superwhisper has a prerelease Linux build for Omarchy only (all checked 2026-10-10). See the comparison hub.
Questions
Can speech to text work without internet?
Yes. Models such as Whisper run on an ordinary CPU. voiceio downloads the model once, from its own source, and then decodes everything on your computer.
Is offline dictation as accurate as cloud dictation?
It depends on the model and your speech. Larger local models (medium, large-v3-turbo) are more accurate but slower without a GPU. voiceio's answer for names and jargon is to fine-tune the local model on your own voice; voiceio learn eval measures the result on your clips.
What does voiceio send over the network?
Nothing by default: no account, no telemetry, and audio never leaves the machine. Two optional, off-by-default features send text (never audio) to an LLM endpoint you configure.
Do I need a GPU?
No. The default small model runs at about 5x realtime on a laptop CPU, per the README, and fine-tuning runs on CPU too. A GPU makes large-v3-turbo practical.
Try voiceio
Free and MIT licensed. Install it with pipx on Linux, or try the dictation in your browser first: the demo on the home page runs a small speech model inside the tab.