Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

parrot-linux

Push-to-talk dictation for Linux apps inside ChromeOS Crostini, inspired by digimata/parrot (macOS).

Hold Right Alt, speak, release. The transcript types itself into the focused X11 app (VS Code, Linux terminals, …) via xdotool and also lands on the clipboard. Transcription is fully local via faster-whisper on the CPU — nothing leaves the machine after the one-time model download.

Install (inside Crostini)

git clone https://github.com/exbald/parrot-linux.git && cd parrot-linux
sudo apt update && sudo apt install -y xdotool xclip wl-clipboard libportaudio2
pip3 install --break-system-packages -r requirements.txt

(Or use a venv instead of --break-system-packages.)

ChromeOS mic access (required): ChromeOS Settings → search "Linux" → Linux development environment → turn ON "Allow Linux to access your microphone", then restart Linux (right-click the Terminal icon → Shut down Linux, reopen). Without this, PulseAudio inside the container only shows a dead auto_null device.

Usage

python3 parrot_linux.py                  # default: base.en, int8, CPU
python3 parrot_linux.py --model small.en # more accurate, ~2.7x slower
python3 parrot_linux.py --no-type        # clipboard only, skip xdotool
python3 parrot_linux.py --list-devices   # show audio devices and exit

Focus any X11 Linux app, hold Right Alt, speak, release. First run downloads the model (~75 MB base.en / ~250 MB small.en) to ~/.cache/huggingface; afterwards it is fully offline.

Run automatically (systemd user service)

To have parrot running whenever the Linux container is up, create ~/.config/systemd/user/parrot.service:

[Unit]
Description=parrot-linux push-to-talk dictation
After=sommelier-x@0.service pipewire.service
Wants=sommelier-x@0.service

[Service]
ExecStart=/usr/bin/python3 %h/parrot-linux/parrot_linux.py
Restart=on-failure
RestartSec=3
# Dictation is interactive: outrank batch dev workloads (linters, builds,
# agent sessions) when the CPU is contended, instead of queueing behind them.
CPUWeight=400
Environment=DISPLAY=:0
Environment=WAYLAND_DISPLAY=wayland-0

[Install]
WantedBy=default.target

then systemctl --user daemon-reload && systemctl --user enable --now parrot.

  • Never run the service and a terminal instance at once — both would type.
  • Watch it live: journalctl --user -fu parrot · stop it: systemctl --user disable --now parrot
  • The container itself still has to be running: opening any Linux app (VS Code, Terminal) boots it, and parrot comes up with it (~20 s until the model is loaded).

Transcript cleanup and spoken commands

Transcripts are cleaned locally before typing (no AI rewriting — your wording is preserved): filler words (um, uh, erm, hmm) are stripped, spacing and sentence capitalization are normalized, and a final period is added when missing (set ENSURE_FINAL_PUNCT = False in the config block if you mostly dictate commands). These spoken commands become punctuation:

Say Get Say Get
comma , question mark ?
period / full stop . exclamation mark !
colon / semicolon : / ; new line / new paragraph line / paragraph break
underscore _ (joins words: file_name) dash / hyphen - (joins words: check-in)

Spoken web addresses are repaired too: Whisper's "www. Linux. Com" becomes www.linux.com (for TLDs in URL_TLDS: com, net, org, io, dev, edu, gov).

Edit SPOKEN_COMMANDS / FILLER_WORDS / URL_TLDS in the config block to taste.

To change the hotkey (if Right Alt doesn't reach the container on your keyboard), edit the config block at the top of parrot_linux.py — alternatives are listed there. VALIDATE.md has a 3-step checklist that verifies mic, typing, and hotkey delivery independently.

Measured performance

On an ARM Chromebook (Cortex-X4 + A720, 8 cores), int8, default threading, measured on real dictated speech (24.7 s utterance):

Model Transcribe time Speed vs realtime
base.en (default) 2.4 s ~0.10x
distil-small.en 3.1 s ~0.13x
small.en 3.5 s ~0.14x

All three were near-identical on clean speech. Two caveats learned the hard way:

  • Cost is per-utterance, not per-second: Whisper pads audio to a 30 s encoder chunk, so a 3 s dictation costs the same as a 28 s one.
  • The table above is a cool machine. After hours of heavy dev workloads the SoC throttles ~4x and small.en climbs to 12–16 s per utterance while base.en stays ~5 s — which is why base.en is the default. With beam=5 + temperature=0 decoding its punctuation matches small.en on clean speech; use --model small.en for maximum accuracy on mumbled/fast speech when the machine is cool.

Note: synthetic pause-free audio (espeak) benchmarks ~10x worse than real speech — the VAD skips natural pauses. Don't set cpu_threads — on big.LITTLE SoCs forcing all cores was measured 1.7x slower than the ctranslate2 default.

For comparison, the original macOS parrot runs whisper-large-v3-turbo on the Apple Neural Engine (~200-300 ms). No NPU is reachable from Crostini, so that model would take tens of seconds per utterance here.

Troubleshooting: mic still auto_null after enabling the toggle

If pactl list sources short shows only auto_null even after turning on mic sharing and restarting Linux: udev inside the Crostini container cannot process the VirtIO sound card (uevent writes to /sys are denied), so PipeWire/WirePlumber's udev-based discovery finds no devices. Declare the card statically — create ~/.config/pipewire/pipewire.conf.d/10-virtio-snd.conf:

context.objects = [
    { factory = adapter
        args = {
            factory.name = api.alsa.pcm.source
            node.name = "virtio-mic"
            media.class = "Audio/Source"
            api.alsa.path = "hw:0,0"
            audio.format = "S16LE"
            audio.rate = 48000
            audio.channels = 2
        }
    }
    { factory = adapter
        args = {
            factory.name = api.alsa.pcm.sink
            node.name = "virtio-speaker"
            media.class = "Audio/Sink"
            api.alsa.path = "hw:0,0"
            audio.format = "S16LE"
            audio.rate = 48000
            audio.channels = 2
        }
    }
]

then systemctl --user restart pipewire pipewire-pulse wireplumber. virtio-mic / virtio-speaker should appear in pactl list sources short.

Known limitations

  • X11 apps only. Typing and the global hotkey work only while an X11 app (under XWayland/sommelier) has focus. ChromeOS surfaces — Chrome tabs, the ChromeOS Terminal app, Android apps — can never receive injected text or hotkeys from inside the container. For those, dictate with any Linux window focused, then paste: the clipboard syncs to ChromeOS via sommelier (Ctrl+V, or Ctrl+Shift+V in ChromeOS Terminal).
  • Hotkey must be a pass-through key. ChromeOS keeps keys like Launcher and Fn for itself; Right Alt (default), Right Ctrl, Pause, or Menu reach the container.
  • VS Code must run under XWayland. If typing does nothing there, relaunch it with code --ozone-platform=x11.
  • English models by default (*.en); pass a multilingual model id and adjust language= in the code for other languages.
  • Held-key autorepeat is debounced in software; a very fast double-tap of the hotkey may be treated as one hold.
  • The mic opens on the first push-to-talk and stays open between utterances, keeping a 0.5 s rolling pre-roll so your first syllable isn't clipped by device wake-up. After MIC_IDLE_CLOSE_SEC (default 45 s) without dictation the stream closes, so the ChromeOS mic-in-use indicator only shows while you're actually dictating. The first utterance after an idle gap has no pre-roll — press, a tiny beat, then speak.

Files

  • parrot_linux.py — the whole app, single file, config block at the top
  • VALIDATE.md — 3-step human checklist (mic, xdotool, hotkey)
  • PLAN.md — environment audit results and design decisions

License

MIT — see LICENSE. Inspired by digimata/parrot; this is an independent Python implementation for ChromeOS Crostini.

About

Push-to-talk dictation for Linux apps in ChromeOS Crostini — hold Right Alt, speak, local Whisper types it into the focused X11 app

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages