GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates OpenAI released GPT Transcribe and GPT Live Transcribe, two new speech recognition models available through its API, with GPT Transcribe achieving a 3.31 percent word error rate on the AA-WER benchmark, a 0.7 percentage point improvement over its predecessor GPT-4o Transcribe. Pricing dropped 25 percent to $0.0045 per minute, but OpenAI still trails competitors ElevenLabs Scribe v2 (2.3 percent), Google Gemini 3 Pro (2.9 percent), and Mistral Voxtral Small (3 percent) on error rates. GPT Transcribe improves on its predecessor but can't catch ElevenLabs, Google, or Mistral on error rates OpenAI has released GPT Transcribe and GPT Live Transcribe, two new speech recognition models available through its API. GPT Transcribe handles pre-recorded audio files, processing them about 34 times faster than real time. GPT Live Transcribe is built for real-time streaming with low latency. According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe hits a word error rate of 3.31 percent. That's a 0.7 percentage point improvement over its year-old predecessor GPT-4o Transcribe https://the-decoder.com/openai-releases-new-ai-voice-models-with-customizable-speaking-styles/ . Pricing drops 25 percent at the same time, landing at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages. In the AA-WER ranking https://artificialanalysis.ai/speech-to-text/non-streaming , OpenAI still sits behind several competitors. ElevenLabs Scribe v2 https://the-decoder.com/elevenlabs-and-google-dominate-artificial-analysis-updated-speech-to-text-benchmark/ leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro https://the-decoder.com/analysts-say-google-now-leads-the-ai-performance-race-with-gemini-3-pro/ at 2.9 percent and Mistral's Voxtral Small https://the-decoder.com/mistral-unveils-voxtral-an-open-source-speech-model-with-lower-costs-than-proprietary-rivals/ at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2 https://the-decoder.com/voxtral-transcribe-2-offers-speech-recognition-at-0-003-per-minute/ , starting at just $0.003 per minute. Full details are in OpenAI's Transcription Guide https://developers.openai.com/api/docs/guides/transcription . The new transcription models complement OpenAI's recently announced Realtime model generation https://the-decoder.com/openais-new-voice-model-brings-gpt-5-level-reasoning-to-real-time-conversations/ , which also includes the real-time transcription model GPT-Realtime-Whisper. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now OpenAI https://developers.openai.com/api/docs/changelog