Gemma 4 E2B: three jobs on 4 GB
A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…
A practitioner is running Google's Gemma 4 E2B model on a single 4 GB VRAM card to handle screen watching, voice-memo and meeting transcription, and chat simultaneously, consolidating three previously…
A developer forking the open-core Meetily product with AI coding tools reports that implementing speaker diarization requires a full pipeline redesign, as the legacy code for the feature is not wired …
A developer is using AI coding agents to fork the open-core meeting assistant Meetily, replacing its paywalled Pro features with open-source alternatives. The fork, tentatively named LibreMeet, aims t…
Interfaze, a Y Combinator-backed startup, open-sourced diffusion-gemma-asr-small, the first multilingual diffusion-based ASR model. The model transcribes six languages using a single 42M-parameter ada…
AetherCut implements hardware-accelerated video editing entirely in the browser, using WebGL for real-time chroma key compositing and WebAssembly to run OpenAI's Whisper model locally for offline tran…
Researchers introduced Latent Ordinal Prototype Alignment (LOPA) and Semantic-Anchored Layer Routing (SALR), a new framework for spoken language assessment that matches the performance of billion-para…
A developer designed a scalable audio transcription pipeline using Faster-Whisper, a highly optimized implementation of OpenAI's Whisper model. The pipeline focuses on high-throughput GPU inference, b…
OpenAI Whisper has 12 unique checkpoint files, but the Python package exposes 14 local model names due to aliases for 'large' and 'turbo'. Additional confusion arises from runtimes and converted forma…
A developer spent $12,000 and 93 hours building a voice agent using Twilio, OpenAI, and ElevenLabs, only to find the total cost to ship a sellable product would exceed $158,000 upfront plus $100,000 p…
Researchers from IIIT-BGP presented low-resource Bhojpuri-Hindi speech translation systems at IWSLT 2026, including an end-to-end model combining Wav2Vec2 and NLLB-200 with a lightweight adapter, and …
ALTIC released FluidVoice v1.6.0, an open-source AI voice-to-text dictation app for macOS that runs speech recognition locally on the device. The app supports multiple speech models, optional on-devic…
Hanzo Huang released a Docker Compose stack that runs Whisper, Piper, openWakeWord, and Qwen 2.5 1.5B on a Rockchip RK3576 NPU, providing a fully local voice backend for Home Assistant via the Wyoming…
Dotdotduck, an open-source Web Agent SDK that turns existing websites into AI-native sites by operating the DOM, was released on Hacker News. The SDK offers a palette UI, proactive offers, dwell-based…
Developer wassgha released OpenDex, an open-source desktop app that turns any LLM into a hands-free voice assistant with a cinematic interface. The app supports fully offline operation on Mac with App…
Pinch released a five-metric benchmark comparing real-time speech translation systems including DeepL, Soniox, GPT-RT, Hibiki, Palabra, and its own Relay-1, evaluating translation quality, intelligibi…
Yap, a free and open-source offline voice dictation tool for Mac, Windows, and Linux, has been released as an alternative to Wispr Flow and SuperWhisper. It runs locally using Whisper for transcriptio…
A developer built Apiarium, a unified API gateway that abstracts multiple AI providers (OpenAI, Anthropic, and more) behind a single endpoint with normalized error handling and credit-based pricing. T…
A developer created AudioTrace, an open-source library that extracts structured signals from voice agent call recordings. The library combines classical signal processing for acoustic measurements and…
Rudrite Research released a 47-minute documentary explaining how modern text-to-speech works, from sound waves as numbers to neural audio codecs and talking LLMs. The film covers key milestones like W…
A developer building a passive AI meeting assistant discovered that their speech recognition model was hallucinating the phrase 'Thank you' during applause, laughter, and silence. The model, trained o…