Shaving Latency Off Real-Time Speech Translation: What Actually Worked in Our Flutter App Owll Translator, a real-time interpretation app, cut perceived latency in its Flutter client by prewarming session connections, treating preloaded session grants as single-use, and streaming partial recognition results through a client-side clause splitter before cloud text-to-speech. The team reports that caching and reusing session grants caused silent failures where audio flowed but no translation returned, and notes it still lacks trustworthy end-to-end mic-to-ear measurements. In a chat app, a slow reply is mildly annoying. In a spoken conversation, it breaks the exchange. Silence feels longer than a spinner does, and people judge the delay from the moment they stop speaking, not from when a request reaches a server. We build Owll Translator https://translator.owll.ai , a real-time interpretation app. You wear earbuds, speak normally, and the app translates live. It can also speak the translation in a clone of your own voice. Most of our latency work over the past months hasn't been about making any single model faster. It has been about perceived latency : hiding the unavoidable waits, cutting the avoidable ones, and getting sound into the listener's ear as early as we responsibly can. This post covers what worked, what didn't, and what we still don't know. One caveat before we start: we don't have trustworthy end-to-end mic-to-ear numbers yet, so there are no millisecond claims here. The last section explains why and what we're doing about it. We run more than one backend path, and the server picks one per session. Here is a simplified view of both: php a Client-orchestrated path ---------------------------- Phone mic - native recorder PCM 16 kHz, 16-bit, mono - local VAD prob threshold 0.5 + noise suppression - cloud speech SDK: continuous recognition + translation push stream; "recognizing" partials, "recognized" finals - client clause splitter complete clauses only - cloud TTS cloned voice - synthesis queue - playback queue - local player - earbuds b Server-orchestrated RTC path -------------------------------- Phone mic - WebRTC audio track LiveKit - server agent: ASR - translate - TTS - translated audio track ------------ earbuds - captions / events via data channel - UI For path a we use the Microsoft Azure Speech SDK for continuous recognition and translation. For path b we use LiveKit on top of WebRTC. Voice synthesis and cloning go through cloud providers. We support Azure neural voices with speaker embeddings, as well as Cartesia and MiniMax. Most of the techniques below apply to path a , where the client has the most control. A few apply to both. The first delay a user notices is between tapping the button and the app actually listening. In the naive version, that gap is a chain of network round trips: fetch an auth token, create a conversation session, receive connection details, connect, then start capture. None of those steps depend on what the user is about to say, so we moved them earlier. While the user is on the home screen, we: prepareConnection ahead of time and enable pre-connect audio when the session starts, so capture can begin before the room is fully joined. A simplified version of the preloader: // Simplified for illustration. class SessionPreloader { Future