{"slug": "add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes", "title": "Add AI meeting summaries to any app with Zoom AI Services in 15 minutes", "summary": "A developer demonstrated how to add AI meeting summaries and action items to any application using Zoom AI Services, Zoom's developer API platform, with roughly 60 lines of Python and two REST calls. The tutorial covers the Scribe transcription endpoint (Fast mode for files up to 5 minutes, Batch for up to 6 hours, Live over WebSocket), which supports 9 locales, speaker diarization, and word-level timestamps via query parameters, then feeds the speaker-labeled transcript to a Summarizer. The author stresses that API keys are server-to-server only and must never ship in client-side code.", "body_md": "If you have meeting recordings (standups, sales calls, support calls) and you want summaries + action items out of them without running your own speech models, two REST calls are all it takes:\n\nBoth are part of **Zoom AI Services**, Zoom's developer API platform. The developer portal lives at [zoom.ai](https://www.zoom.ai/) (verified October 2026 — that's Zoom's own portal, not a third party). No SDK needed — it's plain REST.\n\nTotal code: about 60 lines of Python.\n\n`requests` (`pip install requests`).\nThat's it for credentials. One important rule: **API keys are server-to-server only. Never put the key in client-side code or ship it in a mobile/web app bundle.** All calls in this tutorial run from a backend.\n\nNot ready to create a key? The portal has an **interactive playground with sample data** — no setup required. I'd suggest poking it for two minutes so you know what the responses look like before wiring up the script.\n\nScribe has three modes: **Live** (real-time transcription over a secure WebSocket at `wss://api.zoom.us/v2/aiservices/scribe/live`), **Fast** (synchronous — POST your audio bytes, the response returns when transcription completes), and **Batch** (async jobs API for long files, with webhook callbacks). For this tutorial we use **Fast mode**, the simplest. Fast mode handles files up to **5 minutes**; longer recordings go through Batch (up to 6 hours per file).\n\nInput is the raw audio bytes: WAV, Opus, and µ-law encodings (`audio/wav`, etc. via `Content-Type`). 9 locales: `en-US`, `de-DE`, `es`, `fr-FR`, `it-IT`, `ja-JP`, `ko-KR`, `pt-BR`, `zh-CN`.\n\n**Endpoint:** `POST https://api.zoom.us/v2/aiservices/scribe/transcribe`\n\n**Auth:** one header — `x-api-key: <your API key>`. That's the whole auth story per the portal docs.\n\nOptions go as **query parameters** (`model`, `language`, `diarization`, `text_polishing`, `profanity_filter`, `word_time_offsets`, `channel_separation`, `custom_terms`):\n\n```\ncurl -X POST \"https://api.zoom.us/v2/aiservices/scribe/transcribe?model=zoom-scribe-en&language=en-US&diarization=true&text_polishing=true\" \\\n  -H \"x-api-key: $ZOOM_API_KEY\" \\\n  -H \"Content-Type: audio/wav\" \\\n  --data-binary @meeting.wav\n```\n\nSetting `diarization=true` adds a `speaker` label (`Speaker 1`, `Speaker 2`, …) to each segment — worth it for meetings. Set `Content-Type` to match your file.\n\n**Response (200)** — synchronous; it returns after transcription completes:\n\n```\n{\n  \"request_id\": \"req_123\",\n  \"duration_sec\": 181.04,\n  \"result\": {\n    \"text_display\": \"So, let's have a good conversation, my friend. Do you think...\",\n    \"segments\": [\n      {\n        \"id\": \"seg_1\",\n        \"start\": 0.0,\n        \"end\": 5.2,\n        \"channel\": 1,\n        \"speaker\": \"Speaker 1\",\n        \"text\": \"So, let's have a good conversation, my friend.\",\n        \"words\": []\n      }\n    ]\n  },\n  \"model\": \"zoom-scribe-en\",\n  \"usage\": { \"input_units\": 181.04, \"unit_type\": \"seconds\" }\n}\n```\n\nWith `word_time_offsets=true` the `words` arrays carry per-word timestamps; `channel_separation=true` splits stereo channels.\n\nIn Python:\n\n``` python\nimport os, requests\n\nAPI_BASE = \"https://api.zoom.us/v2/aiservices\"\nAPI_KEY = os.environ[\"ZOOM_API_KEY\"]\n\ndef transcribe(audio_path: str, language: str = \"en-US\") -> dict:\n    with open(audio_path, \"rb\") as f:\n        audio_bytes = f.read()\n    resp = requests.post(\n        f\"{API_BASE}/scribe/transcribe\",\n        headers={\"x-api-key\": API_KEY, \"Content-Type\": \"audio/wav\"},\n        params={\"model\": \"zoom-scribe-en\", \"language\": language,\n                \"diarization\": \"true\", \"text_polishing\": \"true\"},\n        data=audio_bytes,\n        timeout=300,\n    )\n    resp.raise_for_status()\n    return resp.json()[\"result\"]\n```\n\n**Tip:** feed the Summarizer a speaker-labeled transcript — it's much better at action-item ownership. Build it from the segments:\n\n``` php\ndef segments_to_transcript(result: dict) -> str:\n    lines = []\n    for seg in result.get(\"segments\", []):\n        speaker = seg.get(\"speaker\") or \"Speaker\"\n        lines.append(f\"{speaker}: {seg.get('text', '').strip()}\")\n    text = \"\\n\".join(lines)\n    return text or result.get(\"text_display\", \"\")\n```\n\nThe Summarizer API takes transcript text (from *any* system — it doesn't have to come from Scribe) and returns rendered, display-ready text in `result.text`. Fast mode accepts one inline transcript (up to 96 KB) and returns synchronously.\n\n**Endpoint:** `POST https://api.zoom.us/v2/aiservices/summarizer/summarize`\n\nThe request is flat JSON — no `config` wrapper. `transcript` is a **JSON-encoded string** of speaker/text turns:\n\n```\ncurl -X POST https://api.zoom.us/v2/aiservices/summarizer/summarize \\\n  -H \"x-api-key: $ZOOM_API_KEY\" \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"summary_type\": \"CONVERSATION\",\n    \"task\": \"full_summary\",\n    \"transcript\": \"[{\\\"speaker\\\": \\\"Speaker 1\\\", \\\"text\\\": \\\"Let us finalize the timeline.\\\"}, {\\\"speaker\\\": \\\"Speaker 2\\\", \\\"text\\\": \\\"We need one more week for validation.\\\"}]\",\n    \"language\": \"en-US\"\n  }'\n```\n\nFour tasks to choose from, set in `task`:\n\n| `task` | What `result.text` contains | \n|---|---|\n| `recap` | A concise `Recap:` paragraph | \n| `summary` | A detailed `Summary:` section | \n| `action_items` | An `Action Items:` section grouped by owner | \n| `full_summary` | All three — recap + summary + action items | \n\n`summary_type` is `\"CONVERSATION\"` (default; also `\"GENERIC\"`, `\"MEETING_NOTES\"`), `language` is a BCP-47 output locale (14 supported, `en-US` default).\n\n**Response (200)** — the exact shape, from the official docs:\n\n```\n{\n  \"request_id\": \"req_123\",\n  \"task\": \"full_summary\",\n  \"result\": {\n    \"text\": \"Recap:\\n\\nThe team agreed to delay the launch by one week.\\n\\nSummary:\\n\\n...\"\n  }\n}\n```\n\nFor `full_summary`, `result.text` is rendered text, ready to display — recap, summary, and action items grouped by owner.\n\n``` php\nimport json\n\ndef summarize(transcript: str, task: str = \"full_summary\") -> str:\n    turns = []\n    for line in transcript.splitlines():\n        if \":\" in line:\n            speaker, text = line.split(\":\", 1)\n            turns.append({\"speaker\": speaker.strip(), \"text\": text.strip()})\n    resp = requests.post(\n        f\"{API_BASE}/summarizer/summarize\",\n        headers={\"x-api-key\": API_KEY, \"Content-Type\": \"application/json\"},\n        json={\"summary_type\": \"CONVERSATION\", \"task\": task,\n              \"transcript\": json.dumps(turns), \"language\": \"en-US\"},\n        timeout=120,\n    )\n    resp.raise_for_status()\n    data = resp.json()\n    result = data.get(\"result\") or {}\n    return result.get(\"text\") or data.get(\"text\") or data.get(\"summary\") or \"\"\n```\n\nOne script: audio file in, meeting summary out. Save as `summarize_meeting.py`:\n\n``` php\n#!/usr/bin/env python3\n\"\"\"Audio file -> transcript -> meeting summary, via Zoom AI Services (Scribe + Summarizer).\n\nUsage:\n    ZOOM_API_KEY=<your key> python summarize_meeting.py meeting.wav\n\nThe key is server-side only -- never expose it to clients.\n\"\"\"\nimport json, os, sys, requests\n\nAPI_BASE = \"https://api.zoom.us/v2/aiservices\"\nAPI_KEY = os.environ[\"ZOOM_API_KEY\"]\n\ndef transcribe(audio_path: str, language: str = \"en-US\") -> dict:\n    with open(audio_path, \"rb\") as f:\n        audio_bytes = f.read()\n    resp = requests.post(\n        f\"{API_BASE}/scribe/transcribe\",\n        headers={\"x-api-key\": API_KEY, \"Content-Type\": \"audio/wav\"},\n        params={\"model\": \"zoom-scribe-en\", \"language\": language,\n                \"diarization\": \"true\", \"text_polishing\": \"true\"},\n        data=audio_bytes,\n        timeout=300,\n    )\n    resp.raise_for_status()\n    return resp.json()[\"result\"]\n\ndef segments_to_transcript(result: dict) -> str:\n    lines = [\n        f\"{seg.get('speaker') or 'Speaker'}: {seg.get('text', '').strip()}\"\n        for seg in result.get(\"segments\", [])\n    ]\n    return \"\\n\".join(lines) or result.get(\"text_display\", \"\")\n\ndef summarize(transcript: str, task: str = \"full_summary\") -> str:\n    turns = [{\"speaker\": s.strip(), \"text\": t.strip()}\n             for s, _, t in (line.partition(\":\") for line in transcript.splitlines()) if t.strip()]\n    resp = requests.post(\n        f\"{API_BASE}/summarizer/summarize\",\n        headers={\"x-api-key\": API_KEY, \"Content-Type\": \"application/json\"},\n        json={\"summary_type\": \"CONVERSATION\", \"task\": task,\n              \"transcript\": json.dumps(turns), \"language\": \"en-US\"},\n        timeout=120,\n    )\n    resp.raise_for_status()\n    data = resp.json()\n    result = data.get(\"result\") or {}\n    return result.get(\"text\") or data.get(\"text\") or data.get(\"summary\") or \"\"\n\ndef main() -> None:\n    audio_path = sys.argv[1]\n    print(\"Transcribing...\", flush=True)\n    result = transcribe(audio_path)\n    transcript = segments_to_transcript(result)\n    print(f\"Transcript ready ({len(transcript)} chars). Summarizing...\\n\", flush=True)\n    print(summarize(transcript))\n\nif __name__ == \"__main__\":\n    main()\n```\n\nRun it:\n\n```\nZOOM_API_KEY=your_key_here python summarize_meeting.py meeting.wav\n```\n\nThat's the whole pipeline: audio bytes → diarized transcript → recap + summary + action items. (Fast mode caps at 5 minutes of audio; longer files go through the Batch jobs API.)\n\n**On pricing:** Zoom AI Services are usage-based via prepaid credits. I don't have verified per-minute rates, so check the current numbers at [zoom.us/pricing/developer](https://zoom.us/pricing/developer) before you estimate costs.\n\n**401 Unauthorized.** Most common causes: the key is wrong or expired (remember, it's shown only once at creation — regenerate it in the portal if you lost it), or the wrong header name. Per the portal docs, auth is the `x-api-key` header only.\n\n**Unsupported format / fetch errors.** Scribe Fast mode takes the raw audio bytes (`--data-binary`), not a URL — set `Content-Type` to match (`audio/wav`, `audio/mpeg`, …). Keep Fast-mode files under 5 minutes; longer audio goes through Batch.\n\n**429 rate limits.** Back off and retry with exponential backoff. Exact rate limits and quotas depend on your plan — check the portal rather than guessing.\n\n**Diarization quality.** Heavy background noise, overlapping speech, and very short clips all degrade speaker separation. If labels matter to you (they should — the Summarizer groups action items by owner), a quick pass with a noise-reduction tool before upload pays off.\n\n**Transcript too long for Summarizer Fast mode.** Inline input caps at **96 KB**. For longer meetings, split the transcript into chunks and summarize each, or move to Batch mode.\n\n`GET /aiservices/summarizer/jobs/{jobId}` or `.../scribe/jobs/{jobId}`, or get webhook callbacks. Webhooks are signed — verify with the `x-zm-signature` (HMAC-SHA256, `sha256=` prefix) and `x-zm-request-timestamp` headers against your secret.`wss://api.zoom.us/v2/aiservices/scribe/live` — one completed segment per detected speech turn.\n*Disclosure: this article was drafted with AI assistance; all API calls and code were checked against the official Zoom AI Services portal docs (October 2026), and the Scribe flow was run live in the portal playground.*", "url": "https://wpnews.pro/news/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes", "canonical_source": "https://dev.to/dummy_chen_4cfe24c5fe88fc/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes-6ke", "published_at": "2026-10-05 23:32:28+00:00", "updated_at": "2026-10-05 23:47:22.167392+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools", "ai-products"], "entities": ["Zoom", "Zoom AI Services", "Zoom Scribe", "Zoom Summarizer"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes", "markdown": "https://wpnews.pro/news/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes.md", "text": "https://wpnews.pro/news/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes.txt", "jsonld": "https://wpnews.pro/news/add-ai-meeting-summaries-to-any-app-with-zoom-ai-services-in-15-minutes.jsonld"}}