Push-to-talk dictation for Linux apps inside ChromeOS Crostini, inspired by digimata/parrot (macOS).
Hold Right Alt, speak, release. The transcript types itself into the focused X11 app (VS Code, Linux terminals, …) via xdotool and also lands on the clipboard. Transcription is fully local via faster-whisper on the CPU — nothing leaves the machine after the one-time model download.
git clone https://github.com/exbald/parrot-linux.git && cd parrot-linux
sudo apt update && sudo apt install -y xdotool xclip wl-clipboard libportaudio2
pip3 install --break-system-packages -r requirements.txt(Or use a venv instead of --break-system-packages.)
ChromeOS mic access (required): ChromeOS Settings → search "Linux" →
Linux development environment → turn ON "Allow Linux to access your
microphone", then restart Linux (right-click the Terminal icon → Shut down
Linux, reopen). Without this, PulseAudio inside the container only shows a
dead auto_null device.
python3 parrot_linux.py # default: base.en, int8, CPU
python3 parrot_linux.py --model small.en # more accurate, ~2.7x slower
python3 parrot_linux.py --no-type # clipboard only, skip xdotool
python3 parrot_linux.py --list-devices # show audio devices and exitFocus any X11 Linux app, hold Right Alt, speak, release. First run
downloads the model (~75 MB base.en / ~250 MB small.en) to
~/.cache/huggingface; afterwards it is fully offline.
To have parrot running whenever the Linux container is up, create
~/.config/systemd/user/parrot.service:
[Unit]
Description=parrot-linux push-to-talk dictation
After=sommelier-x@0.service pipewire.service
Wants=sommelier-x@0.service
[Service]
ExecStart=/usr/bin/python3 %h/parrot-linux/parrot_linux.py
Restart=on-failure
RestartSec=3
# Dictation is interactive: outrank batch dev workloads (linters, builds,
# agent sessions) when the CPU is contended, instead of queueing behind them.
CPUWeight=400
Environment=DISPLAY=:0
Environment=WAYLAND_DISPLAY=wayland-0
[Install]
WantedBy=default.targetthen systemctl --user daemon-reload && systemctl --user enable --now parrot.
- Never run the service and a terminal instance at once — both would type.
- Watch it live:
journalctl --user -fu parrot· stop it:systemctl --user disable --now parrot - The container itself still has to be running: opening any Linux app (VS Code, Terminal) boots it, and parrot comes up with it (~20 s until the model is loaded).
Transcripts are cleaned locally before typing (no AI rewriting — your
wording is preserved): filler words (um, uh, erm, hmm) are stripped,
spacing and sentence capitalization are normalized, and a final period is
added when missing (set ENSURE_FINAL_PUNCT = False in the config block
if you mostly dictate commands). These spoken commands become punctuation:
| Say | Get | Say | Get |
|---|---|---|---|
| comma | , |
question mark | ? |
| period / full stop | . |
exclamation mark | ! |
| colon / semicolon | : / ; |
new line / new paragraph | line / paragraph break |
| underscore | _ (joins words: file_name) |
dash / hyphen | - (joins words: check-in) |
Spoken web addresses are repaired too: Whisper's "www. Linux. Com" becomes
www.linux.com (for TLDs in URL_TLDS: com, net, org, io, dev, edu, gov).
Edit SPOKEN_COMMANDS / FILLER_WORDS / URL_TLDS in the config block
to taste.
To change the hotkey (if Right Alt doesn't reach the container on your
keyboard), edit the config block at the top of parrot_linux.py —
alternatives are listed there. VALIDATE.md has a 3-step checklist that
verifies mic, typing, and hotkey delivery independently.
On an ARM Chromebook (Cortex-X4 + A720, 8 cores), int8, default threading, measured on real dictated speech (24.7 s utterance):
| Model | Transcribe time | Speed vs realtime |
|---|---|---|
| base.en (default) | 2.4 s | ~0.10x |
| distil-small.en | 3.1 s | ~0.13x |
| small.en | 3.5 s | ~0.14x |
All three were near-identical on clean speech. Two caveats learned the hard way:
- Cost is per-utterance, not per-second: Whisper pads audio to a 30 s encoder chunk, so a 3 s dictation costs the same as a 28 s one.
- The table above is a cool machine. After hours of heavy dev
workloads the SoC throttles ~4x and small.en climbs to 12–16 s per
utterance while base.en stays ~5 s — which is why base.en is the
default. With beam=5 + temperature=0 decoding its punctuation matches
small.en on clean speech; use
--model small.enfor maximum accuracy on mumbled/fast speech when the machine is cool.
Note: synthetic pause-free audio (espeak) benchmarks ~10x worse
than real speech — the VAD skips natural pauses. Don't set cpu_threads —
on big.LITTLE SoCs forcing all cores was measured 1.7x slower than the
ctranslate2 default.
For comparison, the original macOS parrot runs whisper-large-v3-turbo on the Apple Neural Engine (~200-300 ms). No NPU is reachable from Crostini, so that model would take tens of seconds per utterance here.
If pactl list sources short shows only auto_null even after turning on
mic sharing and restarting Linux: udev inside the Crostini container cannot
process the VirtIO sound card (uevent writes to /sys are denied), so
PipeWire/WirePlumber's udev-based discovery finds no devices. Declare the
card statically — create ~/.config/pipewire/pipewire.conf.d/10-virtio-snd.conf:
context.objects = [
{ factory = adapter
args = {
factory.name = api.alsa.pcm.source
node.name = "virtio-mic"
media.class = "Audio/Source"
api.alsa.path = "hw:0,0"
audio.format = "S16LE"
audio.rate = 48000
audio.channels = 2
}
}
{ factory = adapter
args = {
factory.name = api.alsa.pcm.sink
node.name = "virtio-speaker"
media.class = "Audio/Sink"
api.alsa.path = "hw:0,0"
audio.format = "S16LE"
audio.rate = 48000
audio.channels = 2
}
}
]
then systemctl --user restart pipewire pipewire-pulse wireplumber.
virtio-mic / virtio-speaker should appear in pactl list sources short.
- X11 apps only. Typing and the global hotkey work only while an X11 app (under XWayland/sommelier) has focus. ChromeOS surfaces — Chrome tabs, the ChromeOS Terminal app, Android apps — can never receive injected text or hotkeys from inside the container. For those, dictate with any Linux window focused, then paste: the clipboard syncs to ChromeOS via sommelier (Ctrl+V, or Ctrl+Shift+V in ChromeOS Terminal).
- Hotkey must be a pass-through key. ChromeOS keeps keys like Launcher and Fn for itself; Right Alt (default), Right Ctrl, Pause, or Menu reach the container.
- VS Code must run under XWayland. If typing does nothing there,
relaunch it with
code --ozone-platform=x11. - English models by default (
*.en); pass a multilingual model id and adjustlanguage=in the code for other languages. - Held-key autorepeat is debounced in software; a very fast double-tap of the hotkey may be treated as one hold.
- The mic opens on the first push-to-talk and stays open between
utterances, keeping a 0.5 s rolling pre-roll so your first syllable isn't
clipped by device wake-up. After
MIC_IDLE_CLOSE_SEC(default 45 s) without dictation the stream closes, so the ChromeOS mic-in-use indicator only shows while you're actually dictating. The first utterance after an idle gap has no pre-roll — press, a tiny beat, then speak.
parrot_linux.py— the whole app, single file, config block at the topVALIDATE.md— 3-step human checklist (mic, xdotool, hotkey)PLAN.md— environment audit results and design decisions
MIT — see LICENSE. Inspired by digimata/parrot; this is an independent Python implementation for ChromeOS Crostini.