# Build a Voice Assistant with the OpenAI Realtime API and WebRTC

> Source: <https://sourcefeed.dev/a/build-a-voice-assistant-with-the-openai-realtime-api-and-webrtc>
> Published: 2026-07-28 17:40:34+00:00

# Build a Voice Assistant with the OpenAI Realtime API and WebRTC

Wire a browser to gpt-realtime-2.1 over WebRTC for interruptible voice chat with function calling.

[Rachel Goldstein](https://sourcefeed.dev/u/rachel_goldstein)

## What you'll build

A browser-based voice assistant you can talk to — and interrupt mid-sentence — built on the [OpenAI Realtime API](https://developers.openai.com/api/docs/guides/realtime) over [WebRTC](https://webrtc.org), with a function call the model invokes to fetch your actual local time. You'll end with a zero-dependency [Node.js](https://nodejs.org) server that mints ephemeral tokens and a static page that streams audio both ways.

## Prerequisites

- Node.js 20 or later (you need the built-in global
`fetch`

; tested with Node 22 LTS) - An OpenAI account with billing enabled and an API key exported as
`OPENAI_API_KEY`

— Realtime sessions bill per audio token, so expect a few cents for this tutorial - A current desktop browser (Chrome, Edge, Firefox, or Safari) and a working microphone
- Verified July 2026 against the GA Realtime API and the
`gpt-realtime-2.1`

model

No npm packages required — the server uses only Node built-ins.

## 1. Scaffold the project

```
mkdir voice-assistant && cd voice-assistant
mkdir public
printf '{ "name": "voice-assistant", "type": "module" }\n' > package.json
export OPENAI_API_KEY="sk-..."   # your real key
```

The `package.json`

exists only to enable ES modules; there's nothing to install.

## 2. Mint ephemeral tokens on the server

Your API key must never reach the browser. Instead, the server calls `POST /v1/realtime/client_secrets`

to create a short-lived client secret (an `ek_...`

string, 10-minute default TTL) that the browser uses to connect directly to OpenAI:

```
sequenceDiagram
    participant B as Browser
    participant S as Your Node server
    participant O as OpenAI Realtime API
    B->>S: GET /token
    S->>O: POST /v1/realtime/client_secrets (API key)
    O-->>S: ephemeral key (ek_...)
    S-->>B: ephemeral key
    B->>O: POST /v1/realtime/calls (SDP offer + ek_...)
    O-->>B: SDP answer
    B<<->>O: WebRTC: mic audio ⇄ model audio + "oai-events" data channel
```

Create `server.js`

. It serves the static files and exposes `/token`

. The session config — model, voice, instructions, and the tool definition — lives here, baked into the client secret at creation time, so the browser can't tamper with it:

``` js
import { createServer } from "node:http";
import { readFile } from "node:fs/promises";

const SESSION_CONFIG = {
  session: {
    type: "realtime",
    model: "gpt-realtime-2.1",
    instructions:
      "You are a concise, friendly voice assistant. " +
      "When the user asks about the time, call get_current_time and " +
      "answer using its result.",
    audio: { output: { voice: "marin" } },
    tools: [
      {
        type: "function",
        name: "get_current_time",
        description: "Get the user's current local time and timezone.",
        parameters: { type: "object", properties: {}, required: [] },
      },
    ],
    tool_choice: "auto",
  },
};

const server = createServer(async (req, res) => {
  if (req.url === "/token") {
    const r = await fetch("https://api.openai.com/v1/realtime/client_secrets", {
      method: "POST",
      headers: {
        Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(SESSION_CONFIG),
    });
    res.writeHead(r.status, { "Content-Type": "application/json" });
    res.end(JSON.stringify(await r.json()));
    return;
  }
  const path = req.url === "/" ? "/index.html" : req.url;
  try {
    const file = await readFile(`./public${path}`);
    const type = path.endsWith(".js") ? "text/javascript" : "text/html";
    res.writeHead(200, { "Content-Type": type });
    res.end(file);
  } catch {
    res.writeHead(404);
    res.end("Not found");
  }
});

server.listen(3000, () =>
  console.log("Voice assistant running at http://localhost:3000")
);
```

## 3. Build the page

Create `public/index.html`

. The Connect button matters beyond UX: starting from a click satisfies browser autoplay policy, so the model's audio is allowed to play.

```
<!DOCTYPE html>
<html lang="en">
  <head>
    <meta charset="utf-8" />
    <title>Voice Assistant</title>
  </head>
  <body>
    <button id="connect">Connect</button>
    <p id="status">Disconnected</p>
    <pre id="transcript"></pre>
    <script type="module" src="/app.js"></script>
  </body>
</html>
```

## 4. Connect the browser over WebRTC

Create `public/app.js`

. The flow: fetch an ephemeral key, build an `RTCPeerConnection`

with your mic as the outgoing track and an `<audio>`

element for the model's voice, open the `oai-events`

data channel for JSON events, then exchange SDP with `POST https://api.openai.com/v1/realtime/calls`

.

``` js
const connectBtn = document.getElementById("connect");
const statusEl = document.getElementById("status");
const transcriptEl = document.getElementById("transcript");

let pc, dc;

connectBtn.addEventListener("click", async () => {
  connectBtn.disabled = true;
  statusEl.textContent = "Connecting…";

  const tokenResponse = await fetch("/token");
  const { value: EPHEMERAL_KEY } = await tokenResponse.json();

  pc = new RTCPeerConnection();

  const audioEl = document.createElement("audio");
  audioEl.autoplay = true;
  pc.ontrack = (e) => (audioEl.srcObject = e.streams[0]);

  const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
  pc.addTrack(mic.getTracks()[0]);

  dc = pc.createDataChannel("oai-events");
  dc.addEventListener("open", () => {
    statusEl.textContent = "Connected — start talking";
  });
  dc.addEventListener("message", (e) => handleServerEvent(JSON.parse(e.data)));

  const offer = await pc.createOffer();
  await pc.setLocalDescription(offer);

  const sdpResponse = await fetch("https://api.openai.com/v1/realtime/calls", {
    method: "POST",
    body: offer.sdp,
    headers: {
      Authorization: `Bearer ${EPHEMERAL_KEY}`,
      "Content-Type": "application/sdp",
    },
  });
  await pc.setRemoteDescription({ type: "answer", sdp: await sdpResponse.text() });
});
```

That's the entire transport. Turn-taking and barge-in need no extra code — the API's server-side voice activity detection is on by default, and over WebRTC it tracks playback for you, so speaking over the assistant cuts it off cleanly.

## 5. Handle events and function calls

Append this to `public/app.js`

. Every server event arrives as JSON on the data channel. Two matter here: transcript deltas (to render what the assistant is saying) and `response.done`

, whose `response.output`

array contains any `function_call`

items. For each call you run the function locally, send the result back as a `function_call_output`

conversation item — echoing the `call_id`

— then send `response.create`

so the model speaks the answer:

```
function handleServerEvent(event) {
  console.log(event.type);
  switch (event.type) {
    case "response.output_audio_transcript.delta":
      transcriptEl.textContent += event.delta;
      break;
    case "response.output_audio_transcript.done":
      transcriptEl.textContent += "\n\n";
      break;
    case "response.done":
      for (const item of event.response.output ?? []) {
        if (item.type === "function_call") runTool(item);
      }
      break;
    case "error":
      console.error(event.error);
      break;
  }
}

function runTool(item) {
  let output = { error: `Unknown tool: ${item.name}` };
  if (item.name === "get_current_time") {
    output = {
      time: new Date().toLocaleTimeString(),
      timezone: Intl.DateTimeFormat().resolvedOptions().timeZone,
    };
  }
  dc.send(
    JSON.stringify({
      type: "conversation.item.create",
      item: {
        type: "function_call_output",
        call_id: item.call_id,
        output: JSON.stringify(output),
      },
    })
  );
  dc.send(JSON.stringify({ type: "response.create" }));
}
```

## Verify it works

Start the server from the shell where you exported the key:

```
node server.js
```

Expected terminal output:

```
Voice assistant running at http://localhost:3000
```

Open `http://localhost:3000`

, click **Connect**, and allow microphone access. The status line flips to "Connected — start talking." Say: *"Hey — what time is it right now?"* You should hear a spoken reply quoting your actual local time and timezone, and the transcript appears in the `<pre>`

element. The browser console logs the event flow:

```
session.created
input_audio_buffer.speech_started
input_audio_buffer.speech_stopped
response.created
response.output_audio_transcript.delta   (repeated)
response.done                            (contains the function_call item)
response.done                            (the spoken answer, after the tool result)
```

Now interrupt it: ask a question, and while it's answering, start talking. Playback stops immediately and it responds to your new input.

## Troubleshooting

** Incorrect API key provided: undefined** (401 from

`/token`

) — the server process can't see `OPENAI_API_KEY`

, so it sent `Bearer undefined`

. Export the variable in the same shell that runs `node server.js`

, or prefix the command: `OPENAI_API_KEY="sk-..." node server.js`

.** TypeError: Cannot read properties of undefined (reading 'getUserMedia')** — you opened the page over an insecure origin (e.g.

`http://192.168.1.20:3000`

from another machine). `navigator.mediaDevices`

only exists in secure contexts; `http://localhost`

counts, LAN IPs don't. Use localhost or put the server behind HTTPS.** NotAllowedError: Permission denied** from

`getUserMedia`

— microphone access was blocked, now or previously. Click the camera/mic icon in the address bar, reset the permission for the site, and reload.**401 on the POST https://api.openai.com/v1/realtime/calls request** — the ephemeral key expired before the SDP exchange. Client secrets default to a 10-minute TTL, so this happens when a tab sits open between minting and connecting. The code above fetches a fresh token inside the click handler; keep that ordering if you refactor.

## Next steps

- Add real tools with arguments — the
`parameters`

field is standard JSON Schema, so a`get_weather(city)`

function is a few lines plus a fetch in`runTool`

- Enable input transcription in the session's
`audio.input`

config to render the user's side of the conversation too - Read the
[Realtime function calling guide](https://developers.openai.com/api/docs/guides/realtime-function-calling)for multi-call turns, and the[WebRTC guide](https://developers.openai.com/api/docs/guides/realtime-webrtc)for the unified SDP interface that keeps your server in the connection path - For production, look at the
[Agents SDK voice agents quickstart](https://openai.github.io/openai-agents-js/guides/voice-agents/quickstart/), which wraps this transport with sessions, handoffs, and guardrails

## Sources & further reading

-
[Realtime API with WebRTC](https://developers.openai.com/api/docs/guides/realtime-webrtc)— developers.openai.com -
[Realtime API function calling](https://developers.openai.com/api/docs/guides/realtime-function-calling)— developers.openai.com -
[Create client secret - API Reference](https://developers.openai.com/api/reference/resources/realtime/subresources/client_secrets/methods/create)— developers.openai.com -
[Realtime conversations](https://developers.openai.com/api/docs/guides/realtime-conversations)— developers.openai.com -
[Models - OpenAI API](https://developers.openai.com/api/docs/models)— developers.openai.com

[Rachel Goldstein](https://sourcefeed.dev/u/rachel_goldstein)· Dev Tools Editor

Rachel has been embedded in the developer tooling ecosystem for nearly eight years, covering everything from IDE wars and package-manager drama to the quiet rise of AI-assisted coding. She has a soft spot for open-source maintainers and an unhealthy number of terminal emulators installed on a single laptop.

## Discussion 0

No comments yet

Be the first to weigh in.
