Both apps run on the Waveshare ESP32-S3 Touch AMOLED 2.06, the watch-shaped board from our first post, and Amuse also runs on its 1.8″ sibling. You swipe to pick a character, tap Start, and talk. The source for both is on GitHub:
Spark #
Spark is a voice agent. The board opens a WebSocket to OpenAI's Realtime API
(`gpt-realtime-2.1`) itself, streams microphone audio up and plays the reply as it
arrives. There is no server or phone in between. Audio is full duplex, so you can interrupt a reply by
speaking, the way you would with a person.
It has four characters, each with its own voice, persona and animation drawn on a Canvas:
-
KITT , a protective, dry-witted car, with a red scanner while it listens and three columns of voice lamps while it speaks.
-
Nova , an astronomer on an orbital observatory, with twinkling stars and planetary orbits.
-
Echo , a deep-sea navigator, with a sonar sweep and a speech waveform.
-
Flora , a mindfulness companion, with a growing leaf and a flower.
A character is one folder: a `character.ts` with its voice and instructions, a `View.tsx` component, a stylesheet and an optional Canvas animation. Add a folder and the build picks it up; there is nothing else to register. The conversation runs in a Worker, which GeaStack compiles to native code alongside the UI.
Amuse #
Amuse puts a talking face on the screen. Its cast is Zuck, an awkward tech nerd; Alfred, a thoughtful
scholar; Felipe, an imaginative artist; Jojo, a lively optimist; and Todd, an easygoing frog. Each
is a [LemonSlice](https://lemonslice.com) video avatar with an
[ElevenLabs](https://elevenlabs.io) voice. When you tap Start, the board creates a hosted
session, and the character's face, lip-synced to its voice, plays on the AMOLED as you talk.
Amuse needs a small bridge running on a computer on the same network. The bridge joins the hosted session's call and relays the microphone, audio and video between it and the board. Several boards can connect to one bridge at once, each with its own session. On the board, echo cancellation runs natively, so the microphone stays open while the character speaks without hearing itself.
Swiping to a different character ends the current conversation. Each character keeps its configuration and portraits in its own folder, prepared for both supported screen sizes.
Try them #
Each repository's README covers setup. In short: `npm ci`, `npx gea setup` to
register the board and install ESP-IDF, fill in `.env` with your Wi-Fi and API
credentials, then build and flash with npm.
npm ci
npx gea setup
npm run build
npm run flash
Some things to know before you start:
Your API keys go into the firmware. Keep.env and any binaries you
build private. #
Amuse needs the bridge (Python 3.11+) and LemonSlice credentials to hold a conversation.
Both are meant to be forked. Change a persona, add a character, or point them at a different board
with `npx gea setup`. If you build something with them, we'd like to see it.