Pocket Naturalist
A developer built Pocket Naturalist, an offline field guide that uses a local open-weight vision-language model to identify plants and trees from a single photo and delivers the result as a spoken not…
A developer built Pocket Naturalist, an offline field guide that uses a local open-weight vision-language model to identify plants and trees from a single photo and delivers the result as a spoken not…
A developer built an offline wake-word detector for a Raspberry Pi 4 voice assistant named "Nova" using openWakeWord and Piper, finding that training positives with Spanish TTS voices instead of Engli…
A developer built BirdBuddy, an offline bird call identifier that runs entirely on-device, combining Cornell's BirdNET ONNX model for species classification with Piper neural text-to-speech for spoken…
A developer built Podium, a speech-delivery analysis tool that force-aligns a speaker's reading of a public-domain speech (JFK, Reagan) word-by-word against a reference delivery using torchaudio's MMS…
UrukiApp released Pratevenn, an open-source, fully offline AI speaking partner for Norwegian Bokmål learners that runs locally via Docker on CPU or NVIDIA GPUs with 8GB or more of VRAM. Pratevenn prov…
A developer built FloraTrail AI, a fully offline, screenless voice assistant for hikers that runs on a Raspberry Pi 5 or pocketable Linux/Android device. The system combines whisper.cpp for on-device …
A maker built a voice-activated HAL 9000 assistant using a Raspberry Pi Zero 2 W, a ReSpeaker 2-Mics Pi Hat, and a Moebius Models 1:1 scale HAL 9000 model kit, with all speech processing self-hosted o…
OpenBMB shipped VoxCPM2, a 2B-parameter text-to-speech model that generates speech directly in a continuous latent space with no discrete audio tokens, covering 30 languages from raw text without a ph…
Salesforce Inc. introduced a series of AI agents for sales and technical support teams, most of which are generally available today with the rest launching by year's end. The agents include Hunter, Pi…
Google separated pre-integration verification from merge coordination more than 15 years ago, with its monorepo handling 40,000 commits per day by early 2015, including 24,000 from automated systems, …
Verba, an open-source language-learning app, lets users practice eight languages through conversations with an AI that runs entirely offline on their machine, with no account or server required. The a…
A developer built Dub Any Video, a tool that dubs video courses into another language using local text-to-speech, after finding the best Godot course was in Spanish and dubbing services were too expen…
Solo developer Owen Song released Inflect-Micro-v2, a 9.36-million-parameter English text-to-speech model that runs 6.28× faster than real time on four CPU threads and is licensed under Apache-2.0. Th…
Inflect-v2, two open-weight English TTS models at 3.9M and 9.3M parameters, generate speech multiple times faster than real-time on CPU while delivering quality competitive with larger systems like Ki…
OfflineTTS offers unlimited free, local text-to-speech with 54 voices and no data transmission, while Amazon Polly charges $4 per 1 million characters for 100+ cloud-based voices with full SSML suppor…
Self-hosted text-to-speech (TTS) engines in 2026 allow users to run AI voice servers locally on personal PCs or home servers, eliminating ongoing costs, rate limits, and third-party data sharing. A gu…
OfflineTTS offers unlimited free, on-device text-to-speech with 54 voices and no internet requirement, while Microsoft Azure Speech provides 400+ cloud-based neural voices starting at $4 per million c…
Packt Publishing released 'Learn Robotics Programming' by Danny Staple in May 2026, a 740-page guide teaching readers to build and control AI robots using Raspberry Pi and Python. The book covers moto…
OfflineTTS, a free browser-based text-to-speech and speech-to-text toolkit that runs locally, has launched with support for four TTS engines, 99 STT languages, and workflows for subtitle generation, v…
Timeline Studio, a local-first AI video editor that runs in the browser, combines a CapCut-style multi-track timeline with browser-side AI voiceovers, automatic captions, vision tools, talking-avatar …