{"slug": "one-broadcast-several-jobs", "title": "One broadcast, several jobs", "summary": "SageOx is building a real-time AI coworker on Media over QUIC (MoQ) that delivers context to teams within 1–3 seconds during live conversations, with Zoom Scribe handling transcription, according to the company's account of its architecture. SageOx investor Gokul Rajaram said AI-native companies are the hardest customers to win because they build tools themselves and discard them quickly, and SageOx reports its own workflow runs 40x faster with memory now the bottleneck. The company counted phrases from 114 recorded morning standups to justify surfacing yesterday's work while the standup is happening.", "body_md": "# One broadcast, several jobs\n\nWhy 40x might be the floor\n\n**AI made us 40x faster. Memory became the bottleneck.**\n\n**SageOx lets AI coworkers help while the conversation is happening.**\n\n[Ryan](https://sageox.ai/about/ryan-snodgrass) put it simply: “There are zero switching costs.” MoQ lets us keep choosing the best tools.\n\nOne live conversation can reach coworkers, AI coworkers, storage, transcription, and whatever we try next. Here is why we built that on MoQ, and why Zoom Scribe won the transcription job.\n\n## Why does Multiplayer AI learn so fast?\n\nThe models are better than they were a year ago. That is not the whole 40x.\n\n[Austin Vernon](https://www.austinvernon.site/blog/drillinglearningcurve.html)\nexplains why some fields improve exponentially:\n\nFeedback is constant, inexpensive, high signal, and instantaneous. You are making hole, or you aren't.\n\nWhen a decision gets made — in a meeting room, during a walk, or at a cafe —\nthe rest of the team and their AI coworkers can see it in real time. For nerds:\n`tail -f`, not a transcript a few hours later.\n\nA couple of months ago, I was visiting a customer site in San Francisco. The team back in Seattle was fixing bugs and rolling out fixes while I was onboarding the customer.\n\nHere is the plumbing.\n\n## 40x is a problem at the 10 a.m. standup\n\nWe still run a normal morning standup. We like seeing one another. One problem: at 40x, yesterday becomes a blur. It's hard to remember every important thing you need to tell your colleagues.\n\nThe [AI vampire impact](https://steve-yegge.medium.com/the-ai-vampire-eda6e4f07163)\nisn't just a hunch. We record our own standups, so I went and counted. Here is\nevery phrase somebody reached for while trying to remember what they had done\nthe day before, across 114 mornings.\n\nYesterday's work and conversations are already captured. So we're building an\nAI coworker that brings the right context up on our giant TV **while everyone\nis standing up**.\n\nThat's why the words have to arrive during the conversation. We have 1–3 seconds to turn them, along with previous context, into insights, decisions, or murals. An AI coworker that hears the question while it is being asked can answer it.\n\n## Don't pick a winner. Build a learning loop.\n\nGokul Rajaram — my friend, mentor, IITK senior, and a SageOx investor —\n[writes](https://www.linkedin.com/posts/gokulrajaram1_the-new-kingmakers-its-become-clear-that-share-7506022173434773504-zxUd/)\nthis about AI-native companies:\n\nAI-native companies are the hardest customers to win. They have the engineers to build it themselves. They have the taste to know when something is mediocre. They try every new tool the week it launches and rip it out the week after.\n\nWe try new tools in the real system. Benchmarks cannot tell us whether they work while people are still talking.\n\nDo the words arrive in time to help? Does the quietest person get heard? What survives a dropped connection? A strong average score can hide the failure that matters most to our customers.\n\nWe also call the builders. They know what still breaks, usually before the docs say so. Then we run it ourselves. That is how MoQ and Zoom Scribe earned their place. They have to keep earning it.\n\n## Thinking in tools, thinking in bets\n\nFor a long time the honest version of every meeting product was the top panel:\nrecord, upload, wait, read. The input to everything downstream was a *finished\nfile*. However good your model is, it cannot help a conversation whose audio it\nhas not been given yet.\n\nThat design also made every new tool a migration. The recorder, storage path, and transcription provider were tangled together. Trying a better model meant changing the system around it.\n\nMedia over QUIC untangled them. A publisher announces a named broadcast. A relay forwards media objects. Each tool subscribes to the track it needs and does one job:\n\n| Subscriber | Its job | What it ignores | \n|---|---|---|\n| archiver | audio → WebM → S3, with a durability receipt | words, speakers | \n| voiceprint | audio → embeddings → who is talking | words, storage | \n| live bridge | audio → Zoom Scribe Live → words | identity, storage | \n| the next idea | a live mural, an AI coworker that notices a question | all of it | \n\nAdding a new way to *understand* the audio stopped requiring another *recorder*.\nWe can run two transcription providers against the same conversation, add\nvoice identification without changing capture, or try the next good idea\nwithout betting the whole system on it. We can change our minds without a\nmigration.\n\n## End-to-end, one frame at a time\n\nCloudflare's [explainer](https://blog.cloudflare.com/moq/#end-to-end-data-flow) walks through the generic version. Ours has one twist\nworth pointing out.\n\n**Steps 1–5 are a barrier.** The archiver subscribes to a broadcast *that does\nnot exist yet*, and the relay holds that subscription open; only once it is\nacknowledged does api-go hand the browser a credential to publish. We built it\nbecause the archiver used to attach 70 to 240 milliseconds *after* the browser\nstarted talking, and the first group of audio was simply gone. You do not win\nthat race by being faster. You win it by deleting it.\n\n**Steps 7–9 are the fan-out** — one publish, two independent subscribers. The\npublisher does not have to know which tools are listening. The one capture lane\non the left is doing real work in that picture: browser, desktop, mobile, and\nOxDot all publish the same way, over whichever of the two transports the\nnetwork allows. Nothing downstream needs to know which one showed up.\n\n**Steps 10–13 are two answers on two different clocks.** Words come back while\nthe talking continues. Durability comes back separately, on a control plane,\nand it is the only thing that licenses the client to drop its buffered copy. A\nsocket reporting `connected` is not evidence that anything was saved — we\nlearned that from a transport that reported `connected` for 42 seconds after it\nhad stopped delivering anything.\n\n“The difference between theory and practice is smaller in theory than in\npractice.” The explanation above sounds obvious laid out this way. But we built\nthe system step by step, over many conversations with [Alan Frindell](https://sageox.ai/blog/infra-as-unlock-moq-realtime-presence).\n\nI would not trade my weekly one-hour meetings with Alan for anything. I have\nthat privilege because I spent two years sitting next to him at Meta while we\nbuilt out [Proxygen](https://github.com/facebook/proxygen) together.\n\n## Rip it out until it works. Zoom worked.\n\nAWS Transcribe was too late to help a live conversation. ElevenLabs missed too many words. Deepgram looked good until real usage showed it could collapse without warning. We left AWS behind, rejected ElevenLabs, and pulled Deepgram from production. Zoom Scribe won nine of eleven controlled tests and kept its quality across the set. So Zoom is what we ship.\n\nDeepgram was the hardest one to drop. Its live socket responded in about two\nseconds, a tenfold win over Zoom's old chunked lane, and we shipped it to all\nof production. Then we watched it for a month. Its word retention against\nZoom's concurrent output ranged from parity down to **20%**. Not reliably\nworse. *Sometimes* worse, on some conversations, unpredictably.\n\nThat is much harder to live with than being consistently mediocre. A system that is reliably 80% good is one you can design around. A system that is occasionally 20% good reads fine on every day you happen to check it. And when diarization collapses it does not fail evenly — it steamrolls the quietest person in the room, which is exactly who a hivemind exists to hear.\n\nSo we measured it properly: eleven identical five-minute feeds.\n\n| Provider | Word error rate | \n|---|---|\n| **Zoom Scribe** | **14.23%** | \n| Deepgram | 18.87% | \n| ElevenLabs | 20.90% | \n\nThese were real SageOx team conversations, with the technical vocabulary and room dynamics our product has to handle. Your mileage may vary.\n\nZoom Speech ranks among the top models on the\n[HuggingFace Open ASR leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard),\nand [Scribe Live](https://www.zoom.com/en/products/ai-services/scribe-api/) runs\n25¢ an hour.\n\nConsistent quality, fair price, no cliff. That was the whole evaluation.\n\nThe original bet on Zoom was a leap of faith in\n[Zhenbin Xu](https://sageox.ai/blog/scribe-api-powering-sageox), who sat across the table from me\nat an AI community dinner in Seattle. Since then, he has told me what works,\npulled in the engineer who owns the missing Opus support, and been honest about\nthe roadmap. That is why we can bet our product on theirs.\n\nZoom Live went GA on July 30. We had it in a PR by August 6, with one wart: Zoom took only PCM, our wire speaks Opus, so we decoded and re-chopped fifty packets a second.\n\nToday Zhenbin emailed: \"We have added opus codec…try it now.\"\nFour minutes later Zoom returned a perfect transcript of our raw packets. Two\nhours later the translator was gone from production. The docs still say `pcm16`.\n\nFrom the inside, 40x is mostly this: someone tells you, and it's done.\n\n## One hard part, so this doesn't sound easy\n\nZoom Live gives us words. In our integration its finals carry no speaker labels, so voiceprint stays our speaker authority — and those two evidence streams have to meet again on a single timeline.\n\nPacket identity is exactly what a pipeline destroys casually: media conversion on the way to WebM and MP3, sample positions derived from jittery microphone clocks. The relay is forbidden by design from preserving any of it. So the endpoints carry it themselves — client sequence ranges ride every stream, and the archived media embeds an explicit sample-to-sequence map.\n\nThat is the end-to-end argument:\n[Saltzer, Reed and Clark wrote it down in 1984](https://web.mit.edu/Saltzer/www/publications/endtoend/endtoend.pdf).\n\n## Why Seattle\n\nIn April 2025, Yucheng Low and I pitched Hugging Face's Julien Chaumond on\nbuilding an AI cloud in Seattle. Hugging Face had the developer mindshare;\nSeattle had the people who built AWS, Azure, GCP, and OCI. Our last line:\n[“Bet on the cloudy city.”](https://www.linkedin.com/posts/ajitbanerjee_showboxmarket-startups-entrepreneurship-activity-7489103513839390720-7MlS)\n\nSame reason here. The useful answer had not been written down. It was in\nsomeone's head, [across a table in Seattle](https://sageox.ai/blog/scribe-api-powering-sageox).", "url": "https://wpnews.pro/news/one-broadcast-several-jobs", "canonical_source": "https://sageox.ai/blog/one-broadcast-several-jobs", "published_at": "2026-09-22 00:00:00+00:00", "updated_at": "2026-10-02 23:07:39.579864+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "artificial-intelligence", "natural-language-processing"], "entities": ["SageOx", "Media over QUIC", "Zoom Scribe", "Ryan Snodgrass", "Austin Vernon", "Gokul Rajaram", "Steve Yegge"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/one-broadcast-several-jobs", "markdown": "https://wpnews.pro/news/one-broadcast-several-jobs.md", "text": "https://wpnews.pro/news/one-broadcast-several-jobs.txt", "jsonld": "https://wpnews.pro/news/one-broadcast-several-jobs.jsonld"}}