{"slug": "hal-9000-voice-assistant", "title": "Hal 9000 Voice Assistant", "summary": "A maker built a voice-activated HAL 9000 assistant using a Raspberry Pi Zero 2 W, a ReSpeaker 2-Mics Pi Hat, and a Moebius Models 1:1 scale HAL 9000 model kit, with all speech processing self-hosted on a homelab server running Proxmox VE on an Intel Core i5-8400 with 16 GB DDR4 and an NVIDIA GeForce GTX 1060 6GB. The system triggers on the wake phrase \"Hey HAL\" via Porcupine, transcribes speech with Vosk, generates responses with Ollama running the 8-billion-parameter llama3 model, and speaks through Piper using a custom HAL 9000 voice model trained on audio samples from 2001: A Space Odyssey. The HAL 9000 .onnx text-to-speech model and its dataset are published on HuggingFace under the user campwill.", "body_md": "|   |          A voice-activated AI assistant modeled after HAL 9000 from *2001: A Space Odyssey* . The assistant is built using a Raspberry Pi Zero 2 W, integrated with dual microphones, a speaker, and status LEDs - all housed within a 1:1 scale HAL 9000 model kit.          The system activates on the wake phrase “Hey HAL” using [Porcupine](https://github.com/Picovoice/porcupine) for local wake word detection and processes spoken input using fully self-hosted services for speech-to-text ([Vosk](https://github.com/alphacep/vosk-server) ), language generation ([Ollama](https://github.com/ollama/ollama) ), and text-to-speech ([Piper](https://github.com/rhasspy/piper) ). The text-to-speech voice is custom-trained using samples from the film to closely match HAL’s original tone.          A demo of the voice assistant can be viewed [here](#demo) . | \n\nThis project combines the following hardware and software components:\n\n- [**Raspberry Pi Zero 2 W**](https://www.raspberrypi.com/products/raspberry-pi-zero-2-w/) - The voice assistant's main computer. Runs the Python assistant script and handles communication with STT/LLM/TTS services via local network.\n- [**ReSpeaker 2-Mics Pi Hat**](https://www.seeedstudio.com/ReSpeaker-2-Mics-Pi-HAT.html?srsltid=AfmBOooVBplsE1S27Ix879C-gS0P7OQUHIdzmybpualjSxRoyzHZtlWk) - Provides two onboard microphones, a JST 2.0 speaker output, and three programmable LEDs.\n- [**adafruit Mono Enclosed Speaker (1W 8 Ohm)**](https://www.adafruit.com/product/5986) - Compatible speaker with JST 2.0 connector, used for audio output.\n- [**Moebius Models HAL 9000 1:1 Scale Model Kit**](https://a.co/d/a20T0uZ) - Enclosure used to house hardware, modeled after HAL 9000 from*2001: A Space Odyssey* .\n- **Homelab Server** – Hosts all compute-heavy services (speech-to-text, language generation, and text-to-speech) over the LAN. Provides fast, local, offline processing with no reliance on cloud services.\n  - **OS:** Proxmox VE running Ubuntu Server VM\n  - **CPU:** Intel Core i5-8400\n  - **RAM:** 16 GB DDR4\n  - **GPU:** NVIDIA GeForce GTX 1060 6GB (used for LLM acceleration and Piper TTS training)\n  - **Storage:** 50 GB SSD allocated to VM\n\n- [**Porcupine**](https://github.com/Picovoice/porcupine) - Used for \"Hey Hal\" wake word detection, which runs locally on the Raspberry Pi Zero 2 W.\n- [**Vosk**](https://github.com/alphacep/vosk-server) - Speech-to-text server used to transcribe recorded voice input into text.\n- [**Ollama**](https://github.com/ollama/ollama) - Runs the LLM used for generating responses.\n  - It uses the latest [llama3](https://ollama.com/library/llama3) model, featuring 8 billion parameters.\n- It uses the latest \n- [**Piper**](https://github.com/rhasspy/piper) - Text-to-speech engine that converts text into audible speech in real-time. Also used to train a Hal 9000 text-to-speech model using audio samples from the film.\n  - The HAL 9000 .onnx speech-to-text model can be found on my [HuggingFace](https://huggingface.co/campwill/HAL-9000-Piper-TTS) , along with its corresponding[dataset](https://huggingface.co/datasets/campwill/HAL-9000-Speech) .\n- The HAL 9000 .onnx speech-to-text model can be found on my \n\nThe voice assistant is driven by a sequence of self-hosted services, coordinated through a Python script (`app.py`) running on a Raspberry Pi Zero 2 W.\n\nThe assistant runs continuously in a listening state, waiting for the wake phrase “Hey HAL.” Wake word detection is handled locally on the Raspberry Pi using Porcupine. When the phrase is recognized, the system begins actively recording voice input until a silence threshold is met. The recorded audio is then sent over the local network to my homelab server, where it is first transcribed by a speech-to-text service (Vosk). The transcribed text is then passed to a large language model (Ollama), which generates a textual response. This response is then sent to a text-to-speech engine (Piper), which synthesizes speech audio. The audio is streamed back to the Raspberry Pi and played through the speaker, enabling a fully self-hosted, offline voice interaction.\n\nThe process is illustrated below:\n\n```\nsequenceDiagram\n    autonumber\n\n    participant User\n    participant WakeWord as Wake Word Detection<br>(Porcupine)\n    participant Assistant as Voice Assistant \n    participant STT as Speech-to-Text<br>(Vosk)\n    participant LLM as Large Language Model<br>(Ollama)\n    participant TTS as Text-to-Speech<br>(Piper)\n\n    User->>WakeWord: Speak wake word (\"Hey HAL\")\n    WakeWord-->>Assistant: Detect wake word\n    Assistant-->>User: Turn LED on\n    Assistant->>Assistant: Start recording for user input\n    User->>Assistant: Speak voice input\n    Assistant->>Assistant: Detect silence\n    Assistant->>STT: Send audio for transcription\n    STT-->>Assistant: Return transcribed text\n    Assistant->>LLM: Send transcribed text to LLM\n    LLM-->>Assistant: Return LLM text response\n    Assistant->>TTS: Send LLM text response for speech synthesis\n    TTS-->>Assistant: Stream synthesized audio\n    Assistant-->>User: Play response through speaker\n    Assistant-->>User: Turn LED off\n```\n\nBefore setting up the assistant, ensure the following conditions are met:\n\n- A Raspberry Pi Zero 2 W is set up and connected to the same local network as your server.\n- A ReSpeaker 2-Mics Pi Hat is correctly installed and initialized on the Raspberry Pi. Follow the driver setup instructions for the hat on seeed studio's [website](https://wiki.seeedstudio.com/ReSpeaker_2_Mics_Pi_HAT_Raspberry/) .\n- A separate computer or server must be available on the same network to host the required backend services:\n\nThe easiest way to run Vosk is by running the WebSocket server using Docker:\n\n```\ndocker run -d -p 2700:2700 alphacep/kaldi-en:latest\n```\n\nMore information for running the server can be found on the [official documentation](https://alphacephei.com/vosk/server) for Vosk. Ensure Vosk is running the WebSocket server on port 2700.\n\nOllama can be installed using the following command:\n\n```\ncurl -fsSL https://ollama.com/install.sh | sh\n```\n\nBy default, Ollama runs on port 11434. Once Ollama is installed, you will need to run a large language model of your choosing. For this project, I used the latest [llama3](https://ollama.com/library/llama3) model, featuring 8 billion parameters.\n\n```\nollama run llama3:latest\n```\n\nIf you decide to use a different model, you will need to change the `OLLAMA_MODEL` constant in `app.py`.\n\nTo set up the Piper Python HTTP server, I recommend following Thorsten-Voice's tutorial on [YouTube](https://www.youtube.com/watch?app=desktop&v=pLR5AsbCMHs). He provides excellent resources for setting up a Piper TTS environment, as well as training your own voices. More information on Piper can be found on their [GitHub](https://github.com/rhasspy/piper). Ensure Piper is running the HTTP server on port 5000.\n\nIf you'd like to use a custom text-to-speech model, you can run the HTTP server with my custom HAL 9000 model, available on my [HuggingFace](https://huggingface.co/campwill/HAL-9000-Piper-TTS) page. Alternatively, you can train your own model using the corresponding [dataset](https://huggingface.co/datasets/campwill/HAL-9000-Speech).\n\nWith all the services running on their respective ports (Vosk, Ollama, and Piper), the Raspberry Pi client can be set up.\n\n```\nclone https://github.com/campwill/hal-voice-assistant.git\ncd hal-voice-assistant\n```\n\nCreate the virtual environment and install all the required dependencies.\n\n```\npython3 -m venv .venv\nsource .venv/bin/activate\npip install -r requirements.txt\n```\n\nIn `app.py`, set the correct device index for your ReSpeaker 2-Mics Pi Hat:\n\n```\nRESPEAKER_INDEX = 1 # Change to your specific device index\n```\n\nTo find the device index for your audio device:\n\n```\narecord -l\n```\n\nCreate a .env file in the root of the project:\n\n```\nnano .env\nPICOVOICE_KEY=your_picovoice_api_key\nPICOVOICE_MODEL_PATH=/absolute/path/to/your/model.ppn\n```\n\nGet your Picovoice API key and your Porcupine wake word model (.ppn) from [Picovoice Developer Console](https://console.picovoice.ai/login).\n\n```\npython app.py\n```\n\nThe Python script can be ran as a service to start automatically once the Raspberry Pi turns on. More information about running scripts on startup can be found [here](https://www.dexterindustries.com/howto/run-a-program-on-your-raspberry-pi-at-startup/#systemd).\n\n## demo.mp4\n\nFor the physical enclosure of my voice assistant, I used the [Moebius Models HAL 9000 1:1 Scale Model Kit](https://a.co/d/a20T0uZ). This kit arrives as a set of unassembled plastic components. I painted the body with Flat Black for the faceplate and Metallic Aluminum for the frame. For the lens components, I used Elmer’s Glue to secure them without fogging or damaging the clear plastic.\n\nThe speaker grill included in the kit was a solid plastic piece with no perforations. To make it functional, I drilled out all of the holes to allow audio to pass through clearly.\n\nBelow are some pictures of the assembly process.\n\n|   |   |   | \n|   |   |   | \n\nI was able to mount the Raspberry Pi in a position where the LED aligned with HAL 9000’s eye. For now, I used cardboard and tape as a temporary solution (I don't own a 3D printer yet). Below are some pictures of the components mounted inside the model kit.\n\n|   |   | \n\n- I hope to shorten the response time by exploring ways to optimize the Vosk STT pipeline, such as reducing silence detection lag or modifying WebSocket handling.\n- Additionally, instead of sending three separate requests to my homelab server, I may create a unified API endpoint that handles the STT, LLM, and TTS stages in a single request to minimize network overhead.\n- The audio for the ReSpeaker 2-Mics Pi Hat has also been giving me issues, especially after the Raspberry Pi reboots. I will be tinkering with the ReSpeaker 2-Mics Pi Hat's audio functionality to see if I can fix the volume issues.\n- I would also like to implement home assistant.", "url": "https://wpnews.pro/news/hal-9000-voice-assistant", "canonical_source": "https://github.com/campwill/hal-voice-assistant", "published_at": "2026-10-04 09:08:42+00:00", "updated_at": "2026-10-04 09:42:05.598448+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-products"], "entities": ["Raspberry Pi Zero 2 W", "HAL 9000", "Porcupine", "Vosk", "Ollama", "llama3", "Piper", "HuggingFace"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/hal-9000-voice-assistant", "markdown": "https://wpnews.pro/news/hal-9000-voice-assistant.md", "text": "https://wpnews.pro/news/hal-9000-voice-assistant.txt", "jsonld": "https://wpnews.pro/news/hal-9000-voice-assistant.jsonld"}}