Magnus AI: a real-time voice assistant that controls my Windows PC A developer built Magnus AI, a voice-first desktop assistant for Windows that uses Google's Gemini Live API to control apps, adjust system settings, run routines, and analyze screen or webcam input. The project streams 16 kHz mono audio over a WebSocket for roughly 250 ms turn detection with barge-in, exposes self-describing typed tools from an actions/ directory, and uses Win32 APIs, psutil, and PyAutoGUI for system control. The developer plans an offline mode using open-weight Gemma models via Ollama and faster-whisper so voice and screen data stay local. What I Built Magnus AI is a voice-first desktop assistant for Windows. You talk to it, and it can: - Open apps and bring windows to the front - Set volume and brightness - Run one-phrase routines like "Work Mode", "Study Mode" and "Night Mode" - Search the web, play YouTube, and set reminders - Read your clipboard - Look at your screen or webcam and help debug code Demo https://drive.google.com/file/d/137xF8eviOjMR5DFXnfyZKzUN IJyS8Br/view?usp=sharing https://drive.google.com/file/d/137xF8eviOjMR5DFXnfyZKzUN IJyS8Br/view?usp=sharing Code https://github.com/RIDDHIDEV-OPS/Magnus-AI https://github.com/RIDDHIDEV-OPS/Magnus-AI How It Works - Voice loop:microphone audio is resampled to 16 kHz mono and streamed over a WebSocket to the Gemini Live API. Replies return as 24 kHz audio, with about 250 ms turn detection and barge-in so you can interrupt. - Self-describing tools:each file in actions/ exposes one typed function, and its docstrings become function declarations at startup. A new skill is one new file. - Windows control:Win32 APIs, psutil and PyAutoGUI. - Interface:a PyQt6 HUD with an animated orb, waveforms and live CPU/RAM/GPU telemetry. - Memory:preferences are stored in a local JSON file, git-ignored. - Resilience:a fallback ladder across Gemini models when quota is hit. What's Next Magnus currently relies on Google's hosted Gemini model. My next step is an offline mode built on open-weight models Gemma via Ollama plus faster-whisper , so voice and screen data never leave the machine.