A local, private AI articulation coach designed to help users improve their spoken English fluency, confidence, and spontaneity.
Articulate AI implements a fully local Voice-to-Voice (V2V) loop. It captures raw audio, converts it to text, processes it through a conversational AI, and speaks the response back to the user—all without sending data to the cloud.
The backend is a FastAPI server that manages three specialized AI pipelines:
-
Audio Capture
$\rightarrow$ STT:- Captures audio at 44.1kHz.
- Uses
faster-whisper(Small/Medium) for high-accuracy, local speech-to-text transcription. - Implements a "Raw Mode" pipeline to prevent audio corruption and minimize hallucinations.
-
Cognition (LLM):
- Uses
Ollama(gemma4:e4b) as the articulation coach. - Maintains a short-term conversation memory to keep the context of the practice session.
- Uses
-
TTS
$\rightarrow$ Audio Output:- Uses
Piper(ONNX runtime) for ultra-fast, local text-to-speech synthesis. - Outputs fixed-name
.wavfiles (ai_response.wav) to minimize disk clutter.
- Uses
The frontend is a single-file embedded HTML/JS interface designed for zero-latency interaction:
- Reactive UI: Uses a pulse-animation recording button to indicate active listening.
- Asynchronous Flow: Leverages the
Fetch APIto trigger recording start/stop and retrieve AI responses without page reloads. - Real-time Feedback: Displays a live transcript of "What I heard" and a conversation history log.
- Python 3.12
- Ollama: Installed and running with
gemma4:e4bpulled. - Piper: Executable and model files (
.onnxand.json) must be in the root directory.
- Clone the repository
- Set up the virtual environment:
python -m venv .venv .\.venv\Scripts\activate pip install -r requirements.txt
- Run the application:
python app.py
- Access the UI: Open
http://127.0.0.1:8000in your browser.
- Microphone: Uses system default device. To avoid "dimming" or "vanishing" voice, it is recommended to disable "Audio Enhancements" in Windows Sound Settings.
- Audio Flow:
- Root storage for audio:
articulate_audio/ - Only keeps the current
recording_16000.wavandai_response.wavto avoid disk clutter.
- Root storage for audio:
app.py: Core engine (Backend + Embedded Frontend).articulate_audio/: Temporary audio workspace.AGENTS.md: AI agent guidance.requirements.txt: Project dependencies.
Local use only.