Add AI meeting summaries to any app with Zoom AI Services in 15 minutes A developer demonstrated how to add AI meeting summaries and action items to any application using Zoom AI Services, Zoom's developer API platform, with roughly 60 lines of Python and two REST calls. The tutorial covers the Scribe transcription endpoint (Fast mode for files up to 5 minutes, Batch for up to 6 hours, Live over WebSocket), which supports 9 locales, speaker diarization, and word-level timestamps via query parameters, then feeds the speaker-labeled transcript to a Summarizer. The author stresses that API keys are server-to-server only and must never ship in client-side code. If you have meeting recordings standups, sales calls, support calls and you want summaries + action items out of them without running your own speech models, two REST calls are all it takes: Both are part of Zoom AI Services , Zoom's developer API platform. The developer portal lives at zoom.ai https://www.zoom.ai/ verified October 2026 — that's Zoom's own portal, not a third party . No SDK needed — it's plain REST. Total code: about 60 lines of Python. requests pip install requests . That's it for credentials. One important rule: API keys are server-to-server only. Never put the key in client-side code or ship it in a mobile/web app bundle. All calls in this tutorial run from a backend. Not ready to create a key? The portal has an interactive playground with sample data — no setup required. I'd suggest poking it for two minutes so you know what the responses look like before wiring up the script. Scribe has three modes: Live real-time transcription over a secure WebSocket at wss://api.zoom.us/v2/aiservices/scribe/live , Fast synchronous — POST your audio bytes, the response returns when transcription completes , and Batch async jobs API for long files, with webhook callbacks . For this tutorial we use Fast mode , the simplest. Fast mode handles files up to 5 minutes ; longer recordings go through Batch up to 6 hours per file . Input is the raw audio bytes: WAV, Opus, and µ-law encodings audio/wav , etc. via Content-Type . 9 locales: en-US , de-DE , es , fr-FR , it-IT , ja-JP , ko-KR , pt-BR , zh-CN . Endpoint: POST https://api.zoom.us/v2/aiservices/scribe/transcribe Auth: one header — x-api-key: