cd /news/ai-agents/one-broadcast-several-jobs · home › topics › ai-agents › article
[ARTICLE · art-144198] src=sageox.ai ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

One broadcast, several jobs

SageOx is building a real-time AI coworker on Media over QUIC (MoQ) that delivers context to teams within 1–3 seconds during live conversations, with Zoom Scribe handling transcription, according to the company's account of its architecture. SageOx investor Gokul Rajaram said AI-native companies are the hardest customers to win because they build tools themselves and discard them quickly, and SageOx reports its own workflow runs 40x faster with memory now the bottleneck. The company counted phrases from 114 recorded morning standups to justify surfacing yesterday's work while the standup is happening.

by read8 min views5 publishedSep 22, 2026
One broadcast, several jobs
Image: Sageox (auto-discovered)

Why 40x might be the floor

AI made us 40x faster. Memory became the bottleneck.

SageOx lets AI coworkers help while the conversation is happening.

Ryan put it simply: “There are zero switching costs.” MoQ lets us keep choosing the best tools. One live conversation can reach coworkers, AI coworkers, storage, transcription, and whatever we try next. Here is why we built that on MoQ, and why Zoom Scribe won the transcription job.

Why does Multiplayer AI learn so fast? #

The models are better than they were a year ago. That is not the whole 40x.

Austin Vernon explains why some fields improve exponentially:

Feedback is constant, inexpensive, high signal, and instantaneous. You are making hole, or you aren't.

When a decision gets made — in a meeting room, during a walk, or at a cafe — the rest of the team and their AI coworkers can see it in real time. For nerds: tail -f, not a transcript a few hours later.

A couple of months ago, I was visiting a customer site in San Francisco. The team back in Seattle was fixing bugs and rolling out fixes while I was onboarding the customer.

Here is the plumbing.

40x is a problem at the 10 a.m. standup #

We still run a normal morning standup. We like seeing one another. One problem: at 40x, yesterday becomes a blur. It's hard to remember every important thing you need to tell your colleagues.

The AI vampire impact isn't just a hunch. We record our own standups, so I went and counted. Here is every phrase somebody reached for while trying to remember what they had done the day before, across 114 mornings.

Yesterday's work and conversations are already captured. So we're building an AI coworker that brings the right context up on our giant TV while everyone is standing up.

That's why the words have to arrive during the conversation. We have 1–3 seconds to turn them, along with previous context, into insights, decisions, or murals. An AI coworker that hears the question while it is being asked can answer it.

Don't pick a winner. Build a learning loop. #

Gokul Rajaram — my friend, mentor, IITK senior, and a SageOx investor —

[writes](https://www.linkedin.com/posts/gokulrajaram1_the-new-kingmakers-its-become-clear-that-share-7506022173434773504-zxUd/)
this about AI-native companies:

AI-native companies are the hardest customers to win. They have the engineers to build it themselves. They have the taste to know when something is mediocre. They try every new tool the week it launches and rip it out the week after.

We try new tools in the real system. Benchmarks cannot tell us whether they work while people are still talking.

Do the words arrive in time to help? Does the quietest person get heard? What survives a dropped connection? A strong average score can hide the failure that matters most to our customers.

We also call the builders. They know what still breaks, usually before the docs say so. Then we run it ourselves. That is how MoQ and Zoom Scribe earned their place. They have to keep earning it.

Thinking in tools, thinking in bets #

For a long time the honest version of every meeting product was the top panel: record, upload, wait, read. The input to everything downstream was a finished file. However good your model is, it cannot help a conversation whose audio it has not been given yet.

That design also made every new tool a migration. The recorder, storage path, and transcription provider were tangled together. Trying a better model meant changing the system around it.

Media over QUIC untangled them. A publisher announces a named broadcast. A relay forwards media objects. Each tool subscribes to the track it needs and does one job:

Subscriber Its job What it ignores
archiver audio → WebM → S3, with a durability receipt words, speakers
voiceprint audio → embeddings → who is talking words, storage
live bridge audio → Zoom Scribe Live → words identity, storage
the next idea a live mural, an AI coworker that notices a question all of it

Adding a new way to understand the audio stopped requiring another recorder. We can run two transcription providers against the same conversation, add voice identification without changing capture, or try the next good idea without betting the whole system on it. We can change our minds without a migration.

## End-to-end, one frame at a time

Cloudflare's [explainer](https://blog.cloudflare.com/moq/#end-to-end-data-flow) walks through the generic version. Ours has one twist

worth pointing out.

Steps 1–5 are a barrier. The archiver subscribes to a broadcast that does not exist yet, and the relay holds that subscription open; only once it is acknowledged does api-go hand the browser a credential to publish. We built it because the archiver used to attach 70 to 240 milliseconds after the browser started talking, and the first group of audio was simply gone. You do not win that race by being faster. You win it by deleting it.

Steps 7–9 are the fan-out — one publish, two independent subscribers. The publisher does not have to know which tools are listening. The one capture lane on the left is doing real work in that picture: browser, desktop, mobile, and OxDot all publish the same way, over whichever of the two transports the network allows. Nothing downstream needs to know which one showed up.

Steps 10–13 are two answers on two different clocks. Words come back while the talking continues. Durability comes back separately, on a control plane, and it is the only thing that licenses the client to drop its buffered copy. A socket reporting connected is not evidence that anything was saved — we learned that from a transport that reported connected for 42 seconds after it had stopped delivering anything.

“The difference between theory and practice is smaller in theory than in practice.” The explanation above sounds obvious laid out this way. But we built

the system step by step, over many conversations with Alan Frindell. I would not trade my weekly one-hour meetings with Alan for anything. I have that privilege because I spent two years sitting next to him at Meta while we

built out Proxygen together.

Rip it out until it works. Zoom worked. #

AWS Transcribe was too late to help a live conversation. ElevenLabs missed too many words. Deepgram looked good until real usage showed it could collapse without warning. We left AWS behind, rejected ElevenLabs, and pulled Deepgram from production. Zoom Scribe won nine of eleven controlled tests and kept its quality across the set. So Zoom is what we ship.

Deepgram was the hardest one to drop. Its live socket responded in about two seconds, a tenfold win over Zoom's old chunked lane, and we shipped it to all of production. Then we watched it for a month. Its word retention against Zoom's concurrent output ranged from parity down to 20%. Not reliably worse. Sometimes worse, on some conversations, unpredictably.

That is much harder to live with than being consistently mediocre. A system that is reliably 80% good is one you can design around. A system that is occasionally 20% good reads fine on every day you happen to check it. And when diarization collapses it does not fail evenly — it steamrolls the quietest person in the room, which is exactly who a hivemind exists to hear.

So we measured it properly: eleven identical five-minute feeds.

Provider Word error rate
Zoom Scribe 14.23%
Deepgram 18.87%
ElevenLabs 20.90%

These were real SageOx team conversations, with the technical vocabulary and room dynamics our product has to handle. Your mileage may vary.

Zoom Speech ranks among the top models on the

[HuggingFace Open ASR leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard),
and [Scribe Live](https://www.zoom.com/en/products/ai-services/scribe-api/) runs

25¢ an hour.

Consistent quality, fair price, no cliff. That was the whole evaluation.

The original bet on Zoom was a leap of faith in Zhenbin Xu, who sat across the table from me at an AI community dinner in Seattle. Since then, he has told me what works, pulled in the engineer who owns the missing Opus support, and been honest about the roadmap. That is why we can bet our product on theirs.

Zoom Live went GA on July 30. We had it in a PR by August 6, with one wart: Zoom took only PCM, our wire speaks Opus, so we decoded and re-chopped fifty packets a second.

Today Zhenbin emailed: "We have added opus codec…try it now." Four minutes later Zoom returned a perfect transcript of our raw packets. Two hours later the translator was gone from production. The docs still say pcm16.

From the inside, 40x is mostly this: someone tells you, and it's done.

One hard part, so this doesn't sound easy #

Zoom Live gives us words. In our integration its finals carry no speaker labels, so voiceprint stays our speaker authority — and those two evidence streams have to meet again on a single timeline.

Packet identity is exactly what a pipeline destroys casually: media conversion on the way to WebM and MP3, sample positions derived from jittery microphone clocks. The relay is forbidden by design from preserving any of it. So the endpoints carry it themselves — client sequence ranges ride every stream, and the archived media embeds an explicit sample-to-sequence map.

That is the end-to-end argument: Saltzer, Reed and Clark wrote it down in 1984.

Why Seattle #

In April 2025, Yucheng Low and I pitched Hugging Face's Julien Chaumond on building an AI cloud in Seattle. Hugging Face had the developer mindshare; Seattle had the people who built AWS, Azure, GCP, and OCI. Our last line:

“Bet on the cloudy city.” Same reason here. The useful answer had not been written down. It was in someone's head, across a table in Seattle.

── more in #ai-agents 4 stories · sorted by recency
── more on @sageox 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/one-broadcast-severa…] indexed:0 read:8min 2026-09-22 · —