Google shipped Gemini 3.5 Transcribe on August 26, 2026, and the timing makes it a genuinely useful comparison. OpenAI had released its own current flagship transcription model, GPT-Transcribe, just four weeks earlier, on July 28, 2026. Two labs, two new transcription models, released close enough together that comparing them actually means something right now instead of stacking one model generation against another.
Both companies split their offering the same way too — one model built for real-time streaming, one built for pre-recorded audio — which makes the comparison unusually apples-to-apples. Here’s how each got to where it is, a real use case and working code for both, and a side-by-side on the numbers that actually matter.
Contents
Gemini 3.5 Transcribe
Gemini 3.5 Transcribe replaces Chirp 3, Google’s previous transcription model, and the improvement Google is leaning on hardest is speed: a 70% improvement in time-to-final-transcription over Chirp 3, alongside better accuracy. It ships as two distinct model IDs rather than one general-purpose endpoint: gemini-3.5-transcribe-live for continuous, sub-second-latency streaming through the Live API, and gemini-3.5-transcribe for pre-recorded audio, meetings, call logs, and similar, through the Interactions API.
The real numbers, as measured by Artificial Analysis and cited directly in Google’s announcement: a 4.0% word error rate (WER) for streaming use and 2.6% for non-streaming. On the FLEURS multilingual benchmark specifically, Google reports 5.50% WER streaming and 5.04% non-streaming — worth noting as a separate, harder benchmark rather than mixing the two numbers together.
Beyond raw accuracy, the pre-recorded model includes built-in multi-speaker attribution (reliably up to three speakers, with more listed as experimental) and word-level timestamps out of the box, no separate model needed. It also supports over 85 languages, recognizes custom vocabulary, and can delegate follow-up tasks like image generation or file analysis to other Gemini models via function calling, currently live in the Gemini app on macOS.
OpenAI’s GPT-Transcribe
Whisper was OpenAI’s original open transcription model, superseded by gpt-4o-transcribe in March 2025, OpenAI’s first transcription model actually built on the GPT-4o architecture rather than Whisper’s older approach. GPT-Transcribe, released July 28, 2026, is the next step in that same line, and OpenAI now recommends it ahead of whisper-1, gpt-4o-transcribe, and gpt-4o-mini-transcribe for transcribing recorded speech in its original language. Like Gemini, it splits into a streaming sibling, gpt-live-transcribe, for continuous, low-latency sessions.
The numbers: on OpenAI’s own launch benchmark against Common Voice across 22 languages, GPT-Transcribe roughly halves whisper-1’s word error rate, from 40.37% down to 19.27%, while costing 25% less per minute than its predecessor. Pricing lands at $0.0045 per minute for file transcription and $0.017 per minute of session audio for the streaming variant. It accepts keyword hints and multiple language hints to help with domain-specific terms and code-switching, and reports which languages it detected in the audio. The honest gap worth naming directly: plain GPT-Transcribe doesn’t do speaker diarization or word-level timestamps — those still require the separate gpt-4o-transcribe-diarize model or, for timestamps specifically, the older whisper-1.
Let’s take a quick look at some use cases.
Using Gemini 3.5 Transcribe for a Multi-Speaker Meeting
Consider a real scenario where the built-in diarization actually earns its keep: transcribing a recorded three-person meeting and getting back who said what, not just a wall of undifferentiated text.
from google import genai
client = genai.Client(api_key="YOUR_GOOGLE_API_KEY")
with open("meeting_recording.mp3", "rb") as f:
audio_bytes = f.read()
response = client.models.generate_content(
model="gemini-3.5-transcribe",
contents=[
{"text": "Transcribe this meeting with speaker labels and timestamps."},
{"inline_data": {"mime_type": "audio/mp3", "data": audio_bytes}},
],
)
print(response.text)
The request sends the raw audio bytes alongside a plain-language instruction, since gemini-3.5-transcribe is built specifically to produce speaker-attributed, timestamped output without needing a separate diarization step or model. For a real meeting, that means the returned transcript already distinguishes Speaker 1, Speaker 2, and Speaker 3 with timestamps attached — output a post-call analytics pipeline could consume directly.
Using GPT-Transcribe for Live Captioning
Here’s a scenario suited to streaming: real-time captions for a live event, where latency matters more than diarization.
import asyncio
import websockets
import json
async def stream_captions(audio_chunks):
uri = "wss://api.openai.com/v1/realtime?intent=transcription"
headers = {"Authorization": "Bearer YOUR_OPENAI_API_KEY"}
async with websockets.connect(uri, extra_headers=headers) as ws:
await ws.send(json.dumps({
"type": "transcription_session.update",
"session": {"input_audio_transcription": {"model": "gpt-live-transcribe"}},
}))
for chunk in audio_chunks:
await ws.send(json.dumps({
"type": "input_audio_buffer.append",
"audio": chunk,
}))
message = await ws.recv()
event = json.loads(message)
if event.get("type") == "conversation.item.input_audio_transcription.delta":
print(event["delta"], end="", flush=True)
This opens a persistent WebSocket connection rather than sending one request per audio clip, which is the whole point of a streaming model. Partial transcription text arrives as delta events while the speaker is still talking, not after the recording ends. Each audio chunk gets appended to an ongoing buffer, and gpt-live-transcribe returns incremental text as it becomes confident enough to commit — exactly the behavior a live-captioning display needs to stay in sync with the speaker.
Comparison Table
| # | Gemini 3.5 Transcribe | OpenAI GPT-Transcribe |
|---|---|---|
| Release date | August 26, 2026 | July 28, 2026 |
| Predecessor | Chirp 3 | gpt-4o-transcribe |
| Streaming model | gemini-3.5-transcribe-live |
gpt-live-transcribe |
| File/pre-recorded model | gemini-3.5-transcribe |
gpt-transcribe |
| Word error rate | 4.0% streaming / 2.6% non-streaming (Artificial Analysis) | ~19.27% on Common Voice, down from whisper-1’s 40.37% |
| Language support | 85+ languages | Keyword and language hints across 22+ benchmarked languages |
| Built-in speaker diarization | Yes, up to 3 speakers reliably | No, requires separate gpt-4o-transcribe-diarize |
| Word-level timestamps | Yes, built in | No, requires whisper-1 |
| Streaming pricing | Not published per-minute as of this writing | $0.017 per minute of session audio |
| File pricing | Not published per-minute as of this writing | $0.0045 per minute |
Wrapping Up
Gemini 3.5 Transcribe’s built-in diarization and timestamps make it the stronger pick the moment your use case is a meeting, a call log, or anything with multiple speakers you need told apart — that capability alone saves an entire second model call OpenAI’s stack still requires.
GPT-Transcribe earns its place on the other end: a cheaper, faster-to-integrate option when the job is straightforward single-speaker transcription or live captioning, and you don’t need attribution at all.
Shittu Olumide is a software engineer and technical writer passionate about leveraging cutting-edge technologies to craft compelling narratives, with a keen eye for detail and a knack for simplifying complex concepts. You can also find Shittu on Twitter.
