Overview
A prototype voice assistant that runs a full conversation loop - listening, understanding, responding, and speaking - without sending anything to the cloud. It listens through the microphone, transcribes what you say using a local speech-recognition model, sends the text to a locally hosted LLM for a reply, and speaks that reply back using an offline text-to-speech engine. Saying something like "goodbye" or "exit" ends the conversation cleanly.
Architecture & Implementation
A continuous loop initializes the speech-to-text, LLM, and text-to-speech subsystems once at startup, then repeatedly captures microphone audio until it detects a complete spoken utterance.
Two interchangeable speech-to-text backends are supported: Whisper, which records audio triggered by amplitude and segments it on silence, and Vosk, which recognizes speech as a continuous stream and finalizes an utterance once a pause threshold is crossed. Both rely on the same volume-based activation and pause-based segmentation heuristics to decide when someone has started and stopped talking.
Once an utterance is transcribed, LangChain sends it to a locally hosted Ollama Llama 3.1 model along with the full dialogue history, and the reply is spoken aloud through Coqui TTS, which picks CUDA or CPU automatically and plays the result back through sounddevice.
Saying an exit phrase - "goodbye", "exit", "quit", or "stop" - ends the loop cleanly, and initialization and response times are printed at runtime to make latency tuning easier.