mst3k-anything
Mst3k-anything, an open-source tool that automatically generates Mystery Science Theater 3000-style robot heckling for any video, is now available, using yt-dlp, ffmpeg, sherpa-onnx with Parakeet 110M…
Mst3k-anything, an open-source tool that automatically generates Mystery Science Theater 3000-style robot heckling for any video, is now available, using yt-dlp, ffmpeg, sherpa-onnx with Parakeet 110M…
NVIDIA's Nemotron 3.5 Lightning, a 30B-A3B hybrid Mamba MoE agent model released two weeks ago, has been extended to process images and speech without any training, thanks to a geometry coincidence wi…
NVIDIA's Nemotron-3-Nano-Omni-30B model, when distributed as a GGUF, often lacks audio and video functionality because popular GGUF repos ship mmproj files containing only the vision tower, omitting t…
VocalCode, a push-to-talk dictation tool for AI coding agents, launches with a $4.99 one-time price and 30-day free trial, processing speech entirely on-device with zero audio uploaded. The tool, whic…
Fireworks, an AI inference platform, does not support voice AI models such as Parakeet, Kokoro, and Qwen ASR because voice workloads require different optimization strategies than text LLMs, according…
DictaFlow's development testing shows Parakeet offers near-instant local dictation on iPhone and Mac with a 483 MB download footprint, while Whisper provides better vocabulary, accent handling, and la…
EnviousWispr, a free open-source AI dictation app for macOS, runs entirely on-device using Whisper and Parakeet speech-to-text models on Apple Silicon, with no cloud, account, or subscription required…
NVIDIA Parakeet is now available as a speech-to-text engine for Telnyx Voice AI, enabling self-hosted multilingual transcription with automatic language detection for 25 European languages on Telnyx i…
Piotr has released StageWhisper Lite, a free, on-device macOS app for meeting transcription, summaries, and action items using local models like parakeet and gemma 4 or BYO AI via OpenClaw or Hermes A…
Apple's new SpeechAnalyzer engine, introduced with iOS 26 and macOS 26, beats OpenAI's Whisper Small on clean English read speech with a 2.12% word error rate on LibriSpeech test-clean versus Whisper …
A developer forking the open-core Meetily product with AI coding tools reports that implementing speaker diarization requires a full pipeline redesign, as the legacy code for the feature is not wired …
Hugging Face and Cerebras have partnered to create a real-time voice AI pipeline using Google DeepMind's Gemma 4, Nvidia's Parakeet, and Alibaba's Qwen3TTS, achieving low-latency speech-to-speech inte…
ALTIC released FluidVoice v1.6.0, an open-source AI voice-to-text dictation app for macOS that runs speech recognition locally on the device. The app supports multiple speech models, optional on-devic…
Privatewhisper.ai launched a private AI voice dictation tool that runs on-device, in a hardware-secured enclave, or via anonymous cloud processing, with no account required and audio never stored or t…
The MLLP-VRAIN research group submitted a cascaded system for the IWSLT 2026 Simultaneous Speech Translation task, using Parakeet and Qwen 3.5 models with adaptive policies. Their system achieved a +5…
MimicScribe offers an AI notetaker that runs audio processing locally on a Mac, allowing users to choose which cloud provider handles text features, aligning with company data policies that restrict u…
A developer has released MimicScribe, a macOS menu bar app that transcribes meetings with approximately 97% accurate on-device speaker identification. The tool offers real-time talking points, a keybo…
A developer benchmarked two speech-to-text models on a CPU-only Azure VM and found that using an AI agent to plan and execute the task cut costs by 62%, from $1.96 to $0.74. The agent, Neo, achieved t…
Hitoku Draft, an open-source, voice-first AI assistant that runs entirely locally, now includes transcription with voice editing and context-aware features that read the user's screen, documents, and …