Hal 9000 Voice Assistant A maker built a voice-activated HAL 9000 assistant using a Raspberry Pi Zero 2 W, a ReSpeaker 2-Mics Pi Hat, and a Moebius Models 1:1 scale HAL 9000 model kit, with all speech processing self-hosted on a homelab server running Proxmox VE on an Intel Core i5-8400 with 16 GB DDR4 and an NVIDIA GeForce GTX 1060 6GB. The system triggers on the wake phrase "Hey HAL" via Porcupine, transcribes speech with Vosk, generates responses with Ollama running the 8-billion-parameter llama3 model, and speaks through Piper using a custom HAL 9000 voice model trained on audio samples from 2001: A Space Odyssey. The HAL 9000 .onnx text-to-speech model and its dataset are published on HuggingFace under the user campwill. | | A voice-activated AI assistant modeled after HAL 9000 from 2001: A Space Odyssey . The assistant is built using a Raspberry Pi Zero 2 W, integrated with dual microphones, a speaker, and status LEDs - all housed within a 1:1 scale HAL 9000 model kit. The system activates on the wake phrase “Hey HAL” using Porcupine https://github.com/Picovoice/porcupine for local wake word detection and processes spoken input using fully self-hosted services for speech-to-text Vosk https://github.com/alphacep/vosk-server , language generation Ollama https://github.com/ollama/ollama , and text-to-speech Piper https://github.com/rhasspy/piper . The text-to-speech voice is custom-trained using samples from the film to closely match HAL’s original tone. A demo of the voice assistant can be viewed here demo . | This project combines the following hardware and software components: - Raspberry Pi Zero 2 W https://www.raspberrypi.com/products/raspberry-pi-zero-2-w/ - The voice assistant's main computer. Runs the Python assistant script and handles communication with STT/LLM/TTS services via local network. - ReSpeaker 2-Mics Pi Hat https://www.seeedstudio.com/ReSpeaker-2-Mics-Pi-HAT.html?srsltid=AfmBOooVBplsE1S27Ix879C-gS0P7OQUHIdzmybpualjSxRoyzHZtlWk - Provides two onboard microphones, a JST 2.0 speaker output, and three programmable LEDs. - adafruit Mono Enclosed Speaker 1W 8 Ohm https://www.adafruit.com/product/5986 - Compatible speaker with JST 2.0 connector, used for audio output. - Moebius Models HAL 9000 1:1 Scale Model Kit https://a.co/d/a20T0uZ - Enclosure used to house hardware, modeled after HAL 9000 from 2001: A Space Odyssey . - Homelab Server – Hosts all compute-heavy services speech-to-text, language generation, and text-to-speech over the LAN. Provides fast, local, offline processing with no reliance on cloud services. - OS: Proxmox VE running Ubuntu Server VM - CPU: Intel Core i5-8400 - RAM: 16 GB DDR4 - GPU: NVIDIA GeForce GTX 1060 6GB used for LLM acceleration and Piper TTS training - Storage: 50 GB SSD allocated to VM - Porcupine https://github.com/Picovoice/porcupine - Used for "Hey Hal" wake word detection, which runs locally on the Raspberry Pi Zero 2 W. - Vosk https://github.com/alphacep/vosk-server - Speech-to-text server used to transcribe recorded voice input into text. - Ollama https://github.com/ollama/ollama - Runs the LLM used for generating responses. - It uses the latest llama3 https://ollama.com/library/llama3 model, featuring 8 billion parameters. - It uses the latest - Piper https://github.com/rhasspy/piper - Text-to-speech engine that converts text into audible speech in real-time. Also used to train a Hal 9000 text-to-speech model using audio samples from the film. - The HAL 9000 .onnx speech-to-text model can be found on my HuggingFace https://huggingface.co/campwill/HAL-9000-Piper-TTS , along with its corresponding dataset https://huggingface.co/datasets/campwill/HAL-9000-Speech . - The HAL 9000 .onnx speech-to-text model can be found on my The voice assistant is driven by a sequence of self-hosted services, coordinated through a Python script app.py running on a Raspberry Pi Zero 2 W. The assistant runs continuously in a listening state, waiting for the wake phrase “Hey HAL.” Wake word detection is handled locally on the Raspberry Pi using Porcupine. When the phrase is recognized, the system begins actively recording voice input until a silence threshold is met. The recorded audio is then sent over the local network to my homelab server, where it is first transcribed by a speech-to-text service Vosk . The transcribed text is then passed to a large language model Ollama , which generates a textual response. This response is then sent to a text-to-speech engine Piper , which synthesizes speech audio. The audio is streamed back to the Raspberry Pi and played through the speaker, enabling a fully self-hosted, offline voice interaction. The process is illustrated below: sequenceDiagram autonumber participant User participant WakeWord as Wake Word Detection