2.1 KiB
2.1 KiB
voice_linux (C99)
Local hotkey dictation for Linux/X11 using:
- PulseAudio capture
- whisper.cpp transcription
- X11/XTest text injection
Build deps (Debian/Qubes AppVM)
sudo apt update
sudo apt install -y build-essential cmake git wget libpulse-dev libx11-dev libxtst-dev libgtk-3-dev pkg-config
Build app
Use build.sh:
# CPU whisper build
bash ./build.sh
# CUDA whisper build
WHISPER_CUDA=1 bash ./build.sh
# build without whisper integration
WITH_WHISPER=0 bash ./build.sh
./voice_linux is the single app binary and always launches the GTK window.
Download whisper models
Download all supported test models:
bash ./scripts/download_models.sh
This fetches:
ggml-tiny.en.binggml-base.en.binggml-small.en.binggml-medium.en.binggml-large-v3-turbo.binggml-large-v3.bin
Run
./voice_linux
On startup, the app opens a dialog where you choose:
- CPU or GPU mode
- Whisper model from
./models/
Default hotkey: Ctrl+Alt+V
- Press once to start recording
- Press again to stop recording, transcribe, and type into the focused window
If you see [stt] [BLANK_AUDIO], recording worked but audio energy was near silent.
Check your input source with:
pactl list short sources
Then set audio_device= in config.ini.
Versioning workflow
This project uses semantic git tags (vMAJOR.MINOR.PATCH) plus header macros in src/version.h.
Use increment_and_push.sh to increment version, update VOICE_LINUX_VERSION, commit, tag, and push:
# patch bump (default)
./increment_and_push.sh "Fix VAD trigger behavior"
# minor bump
./increment_and_push.sh -m "Add new GUI debug indicators"
# major bump
./increment_and_push.sh -M "Breaking config change"
# release mode (also runs bash ./build.sh and creates tarball)
./increment_and_push.sh -r -m "Release minor update"
Notes
- This implementation targets X11 (not native Wayland).
- In template-based AppVMs, system packages may not persist unless installed in the template.