VOICEIO_ GitHub
// guide · offline

Offline speech to text

Offline speech to text means the speech model runs on your own computer, so audio never goes to a server and dictation works without a connection. On Linux, voiceio does this with Whisper: after a one-time model download it decodes your speech locally, types the text into the focused app, and has no account and no telemetry.

no uploadno accountno telemetryCPU is enough
// how to

Set up offline dictation

  1. Install the system packages (a C toolchain, PortAudio, IBus). On Debian or Ubuntu:
    $ sudo apt install pipx build-essential python3-dev \
        portaudio19-dev ibus gir1.2-ibus-1.0 python3-gi
    Fedora, Arch and NixOS commands are in the README.
  2. Install voiceio with the desktop extra (microphone, hotkeys, tray):
    $ pipx install 'python-voiceio[desktop]'
  3. Run the guided setup. It downloads the speech model, sets your hotkey and installs a systemd user service so voiceio starts on login:
    $ voiceio setup
  4. Dictate. Press the hotkey you picked during setup in any app, speak, and press it again. voiceio doctor shows what works; voiceio doctor --fix repairs what it can.

Once the model is downloaded, unplug the network and dictate: nothing changes. To pick a different local model:

$ voiceio models list          # known models and which is active
$ voiceio models use medium    # better on names, about 2.5x slower
// models

Local models voiceio can use

ModelRuntimeNotes (README)
smallfaster-whisperDefault. About 5x realtime on a laptop CPU
mediumfaster-whisperBetter on names, about 2.5x slower
large-v3-turbofaster-whisperMost accurate Whisper option for dictation; wants a GPU
distil-large-v3faster-whisperEnglish only
Parakeet TDT 0.6b v3sherpa-onnxExperimental. Fast, no vocabulary biasing yet
whisper.cpp serverwhisper.cppExperimental. E.g. a Vulkan build on an AMD iGPU

voiceio ships no model weights; each model downloads from its own source under its own license.

// your data

What stays on your disk

Everything voiceio keeps is plain files in ~/.config/voiceio, ~/.local/state/voiceio and ~/.local/share/voiceio, created readable only by you. Recordings are off unless you turn them on; you need them only for fine-tuning. Logs record sizes and timings, never your words.

For a phone or another device, python -m voiceio.server serves the same engine over HTTP and WebSocket on a machine you control. It binds to localhost by default.

// alternatives

Other offline options

  • Voxtype: a local, MIT-licensed Rust dictation tool for Linux with many speech engines and GPU builds.
  • nerd-dictation: a single-file Python script using the offline VOSK engine.
  • Talon: hands-free computer control by voice; on Linux it runs on X11 only.
  • Whisper, whisper.cpp, faster-whisper: speech models and libraries that transcribe audio but do not type into apps by themselves.

Aqua Voice, Wispr Flow and Dragon do not offer Linux apps, and Superwhisper has a prerelease Linux build for Omarchy only (all checked 2026-10-10). See the comparison hub.

// faq

Questions

Can speech to text work without internet?

Yes. Models such as Whisper run on an ordinary CPU. voiceio downloads the model once, from its own source, and then decodes everything on your computer.

Is offline dictation as accurate as cloud dictation?

It depends on the model and your speech. Larger local models (medium, large-v3-turbo) are more accurate but slower without a GPU. voiceio's answer for names and jargon is to fine-tune the local model on your own voice; voiceio learn eval measures the result on your clips.

What does voiceio send over the network?

Nothing by default: no account, no telemetry, and audio never leaves the machine. Two optional, off-by-default features send text (never audio) to an LLM endpoint you configure.

Do I need a GPU?

No. The default small model runs at about 5x realtime on a laptop CPU, per the README, and fine-tuning runs on CPU too. A GPU makes large-v3-turbo practical.

// get it

Try voiceio

Free and MIT licensed. Install it with pipx on Linux, or try the dictation in your browser first: the demo on the home page runs a small speech model inside the tab.