cd /news/artificial-intelligence/gemini-3-5-transcribe-vs-openais-gpt… · home › topics › artificial-intelligence › article
[ARTICLE · art-141034] src=kdnuggets.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Gemini 3.5 Transcribe vs OpenAI’s GPT-Transcribe

Google released Gemini 3.5 Transcribe on August 26, 2026, a two-model transcription line that Google says delivers a 70% improvement in time-to-final-transcription over Chirp 3, with a 4.0% word error rate for streaming and 2.6% for non-streaming as measured by Artificial Analysis. The release follows OpenAI's GPT-Transcribe by four weeks; OpenAI's model, launched July 28, 2026, roughly halves whisper-1's word error rate on Common Voice across 22 languages from 40.37% to 19.27% and costs $0.0045 per minute for file transcription and $0.017 per minute of session audio for streaming. Gemini 3.5 Transcribe ships as gemini-3.5-transcribe-live for streaming and gemini-3.5-transcribe for pre-recorded audio, while GPT-Transcribe's streaming sibling is gpt-live-transcribe.

by read6 min views3 publishedSep 28, 2026
Gemini 3.5 Transcribe vs OpenAI’s GPT-Transcribe
Image: Kdnuggets (auto-discovered)

Here's how each got to where it is, a real use case and working code for both, and a side-by-side on the numbers that actually matter.

Google shipped Gemini 3.5 Transcribe on August 26, 2026, and the timing makes it a genuinely useful comparison. OpenAI had released its own current flagship transcription model, GPT-Transcribe, just four weeks earlier, on July 28, 2026. Two labs, two new transcription models, released close enough together that comparing them actually means something right now instead of stacking one model generation against another.

Both companies split their offering the same way too — one model built for real-time streaming, one built for pre-recorded audio — which makes the comparison unusually apples-to-apples. Here's how each got to where it is, a real use case and working code for both, and a side-by-side on the numbers that actually matter.

Gemini 3.5 Transcribe #

Gemini 3.5 Transcribe replaces Chirp 3, Google's previous transcription model, and the improvement Google is leaning on hardest is speed: a 70% improvement in time-to-final-transcription over Chirp 3, alongside better accuracy. It ships as two distinct model IDs rather than one general-purpose endpoint: gemini-3.5-transcribe-live for continuous, sub-second-latency streaming through the Live API, and gemini-3.5-transcribe for pre-recorded audio, meetings, call logs, and similar, through the Interactions API.

The real numbers, as measured by Artificial Analysis and cited directly in Google's announcement: a 4.0% word error rate (WER) for streaming use and 2.6% for non-streaming. On the FLEURS multilingual benchmark specifically, Google reports 5.50% WER streaming and 5.04% non-streaming — worth noting as a separate, harder benchmark rather than mixing the two numbers together.

Beyond raw accuracy, the pre-recorded model includes built-in multi-speaker attribution (reliably up to three speakers, with more listed as experimental) and word-level timestamps out of the box, no separate model needed. It also supports over 85 languages, recognizes custom vocabulary, and can delegate follow-up tasks like image generation or file analysis to other Gemini models via function calling, currently live in the Gemini app on macOS.

OpenAI's GPT-Transcribe #

Whisper was OpenAI's original open transcription model, superseded by gpt-4o-transcribe in March 2025, OpenAI's first transcription model actually built on the GPT-4o architecture rather than Whisper's older approach. GPT-Transcribe, released July 28, 2026, is the next step in that same line, and OpenAI now recommends it ahead of whisper-1, gpt-4o-transcribe, and gpt-4o-mini-transcribe for transcribing recorded speech in its original language. Like Gemini, it splits into a streaming sibling, gpt-live-transcribe, for continuous, low-latency sessions.

The numbers: on OpenAI's own launch benchmark against Common Voice across 22 languages, GPT-Transcribe roughly halves whisper-1's word error rate, from 40.37% down to 19.27%, while costing 25% less per minute than its predecessor. Pricing lands at $0.0045 per minute for file transcription and $0.017 per minute of session audio for the streaming variant. It accepts keyword hints and multiple language hints to help with domain-specific terms and code-switching, and reports which languages it detected in the audio. The honest gap worth naming directly: plain GPT-Transcribe doesn't do speaker diarization or word-level timestamps — those still require the separate gpt-4o-transcribe-diarize model or, for timestamps specifically, the older whisper-1.

Let's take a quick look at some use cases.

Using Gemini 3.5 Transcribe for a Multi-Speaker Meeting #

Consider a real scenario where the built-in diarization actually earns its keep: transcribing a recorded three-person meeting and getting back who said what, not just a wall of undifferentiated text.

from google import genai

client = genai.Client(api_key="YOUR_GOOGLE_API_KEY")

with open("meeting_recording.mp3", "rb") as f:
    audio_bytes = f.read()

response = client.models.generate_content(
    model="gemini-3.5-transcribe",
    contents=[
        {"text": "Transcribe this meeting with speaker labels and timestamps."},
        {"inline_data": {"mime_type": "audio/mp3", "data": audio_bytes}},
    ],
)
print(response.text)

The request sends the raw audio bytes alongside a plain-language instruction, since gemini-3.5-transcribe is built specifically to produce speaker-attributed, timestamped output without needing a separate diarization step or model. For a real meeting, that means the returned transcript already distinguishes Speaker 1, Speaker 2, and Speaker 3 with timestamps attached — output a post-call analytics pipeline could consume directly.

Using GPT-Transcribe for Live Captioning #

Here's a scenario suited to streaming: real-time captions for a live event, where latency matters more than diarization.

import asyncio
import websockets
import json

async def stream_captions(audio_chunks):
    uri = "wss://api.openai.com/v1/realtime?intent=transcription"
    headers = {"Authorization": "Bearer YOUR_OPENAI_API_KEY"}

    async with websockets.connect(uri, extra_headers=headers) as ws:
        await ws.send(json.dumps({
            "type": "transcription_session.update",
            "session": {"input_audio_transcription": {"model": "gpt-live-transcribe"}},
        }))

        for chunk in audio_chunks:
            await ws.send(json.dumps({
                "type": "input_audio_buffer.append",
                "audio": chunk,
            }))
            message = await ws.recv()
            event = json.loads(message)
            if event.get("type") == "conversation.item.input_audio_transcription.delta":
                print(event["delta"], end="", flush=True)

This opens a persistent WebSocket connection rather than sending one request per audio clip, which is the whole point of a streaming model. Partial transcription text arrives as delta events while the speaker is still talking, not after the recording ends. Each audio chunk gets appended to an ongoing buffer, and gpt-live-transcribe returns incremental text as it becomes confident enough to commit — exactly the behavior a live-captioning display needs to stay in sync with the speaker.

Comparison Table #

# Gemini 3.5 Transcribe OpenAI GPT-Transcribe
Release date August 26, 2026 July 28, 2026
Predecessor Chirp 3 gpt-4o-transcribe
Streaming model gemini-3.5-transcribe-live gpt-live-transcribe
File/pre-recorded model gemini-3.5-transcribe gpt-transcribe
Word error rate 4.0% streaming / 2.6% non-streaming (Artificial Analysis) ~19.27% on Common Voice, down from whisper-1's 40.37%
Language support 85+ languages Keyword and language hints across 22+ benchmarked languages
Built-in speaker diarization Yes, up to 3 speakers reliably No, requires separate gpt-4o-transcribe-diarize
Word-level timestamps Yes, built in No, requires whisper-1
Streaming pricing Not published per-minute as of this writing $0.017 per minute of session audio
File pricing Not published per-minute as of this writing $0.0045 per minute

Wrapping Up #

Gemini 3.5 Transcribe's built-in diarization and timestamps make it the stronger pick the moment your use case is a meeting, a call log, or anything with multiple speakers you need told apart — that capability alone saves an entire second model call OpenAI's stack still requires.

GPT-Transcribe earns its place on the other end: a cheaper, faster-to-integrate option when the job is straightforward single-speaker transcription or live captioning, and you don't need attribution at all.

[Shittu Olumide](https://www.linkedin.com/in/olumide-shittu/) is a software engineer and technical writer passionate about leveraging cutting-edge technologies to craft compelling narratives, with a keen eye for detail and a knack for simplifying complex concepts. You can also find Shittu on Twitter.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-3-5-transcrib…] indexed:0 read:6min 2026-09-28 · —