cd /news/ai-tools/shaving-latency-off-real-time-speech… · home › topics › ai-tools › article
[ARTICLE · art-148049] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Shaving Latency Off Real-Time Speech Translation: What Actually Worked in Our Flutter App

Owll Translator, a real-time interpretation app, cut perceived latency in its Flutter client by prewarming session connections, treating preloaded session grants as single-use, and streaming partial recognition results through a client-side clause splitter before cloud text-to-speech. The team reports that caching and reusing session grants caused silent failures where audio flowed but no translation returned, and notes it still lacks trustworthy end-to-end mic-to-ear measurements.

by read11 min views3 publishedOct 9, 2026

In a chat app, a slow reply is mildly annoying. In a spoken conversation, it breaks the exchange. Silence feels longer than a spinner does, and people judge the delay from the moment they stop speaking, not from when a request reaches a server.

We build Owll Translator, a real-time interpretation app. You wear earbuds, speak normally, and the app translates live. It can also speak the translation in a clone of your own voice. Most of our latency work over the past months hasn't been about making any single model faster. It has been about perceived latency: hiding the unavoidable waits, cutting the avoidable ones, and getting sound into the listener's ear as early as we responsibly can.

This post covers what worked, what didn't, and what we still don't know. One caveat before we start: we don't have trustworthy end-to-end mic-to-ear numbers yet, so there are no millisecond claims here. The last section explains why and what we're doing about it.

We run more than one backend path, and the server picks one per session. Here is a simplified view of both:

 (a) Client-orchestrated path
 ----------------------------
 Phone mic
   -> native recorder (PCM 16 kHz, 16-bit, mono)
   -> local VAD (prob threshold 0.5) + noise suppression
   -> cloud speech SDK: continuous recognition + translation
        (push stream; "recognizing" partials, "recognized" finals)
   -> client clause splitter (complete clauses only)
   -> cloud TTS (cloned voice) -> synthesis queue
   -> playback queue -> local player
   -> earbuds

 (b) Server-orchestrated RTC path
 --------------------------------
 Phone mic
   -> WebRTC audio track (LiveKit)
   -> server agent: ASR -> translate -> TTS
   -> translated audio track  ------------> earbuds
   -> captions / events via data channel -> UI

For path (a) we use the Microsoft Azure Speech SDK for continuous recognition and translation. For path (b) we use LiveKit on top of WebRTC. Voice synthesis and cloning go through cloud providers. We support Azure neural voices with speaker embeddings, as well as Cartesia and MiniMax.

Most of the techniques below apply to path (a), where the client has the most control. A few apply to both.

The first delay a user notices is between tapping the button and the app actually listening. In the naive version, that gap is a chain of network round trips: fetch an auth token, create a conversation session, receive connection details, connect, then start capture.

None of those steps depend on what the user is about to say, so we moved them earlier. While the user is on the home screen, we:

prepareConnection() ahead of time and enable pre-connect audio when the session starts, so capture can begin before the room is fully joined. A simplified version of the pre:

// Simplified for illustration.
class SessionPre {
  Future<SessionGrant>? _inFlight;
  SessionGrant? _ready;

  Future<SessionGrant> warm() {
    // Dedupe: concurrent callers share one request.
    return _inFlight ??= _api.createSession().then((grant) {
      _ready = grant;
      _inFlight = null;
      return grant;
    });
  }

  /// Grants are single-use. Hand it out once, then forget it.
  SessionGrant? take() {
    final grant = _ready;
    _ready = null;
    return grant;
  }
}

The take() method is the result of a bug we caused ourselves. Our first version cached the session response and reused it. Sometimes that produced the worst kind of failure: the UI showed "connected", audio flowed, and no translation ever came back. The cached grant was stale, and the backend treated it as already used. We now treat preloaded grants as single-use. They're consumed once and invalidated, and we recorded that rule in an architecture decision record so nobody "optimizes" it back later.

Lesson: prewarming is free latency only if the thing you prewarm is safe to hold. Anything with session semantics needs an explicit lifecycle.

Prewarming narrows the startup gap but doesn't close it. The recognizer may still be initializing when the user starts talking, and people tend to start talking the moment they tap. If capture only starts once the SDK reports it is ready, the first syllables are gone, and recognition of a clipped first word is often wrong, which forces the user to repeat the sentence. That is far worse than a short delay.

So the recorder starts immediately and pushes PCM continuously. If the speech SDK isn't ready yet, we buffer the audio and flush it as soon as the push stream exists:

// Simplified for illustration.
void onPcmFrame(Uint8List frame) {
  final stream = _pushStream;
  if (stream == null) {
    _pending.add(frame); // SDK not ready yet: keep it
    return;
  }
  if (_pending.isNotEmpty) {
    for (final f in _pending) stream.write(f);
    _pending.clear();
  }
  stream.write(frame);
}

The recognizer gets a slightly late but complete stream instead of a punctual one that's missing its first words. Local VAD and noise suppression run before the network, which keeps silence and background noise from making the recognizer do extra work.

This was the biggest perceived improvement, and also the change with the most trade-offs.

The speech SDK emits two kinds of events: recognizing (partial, may still change) and recognized (final for that segment). The simple approach is to wait for recognized and send the whole translated segment to TTS. That's correct, but the listener hears nothing until the speaker finishes an entire thought.

Instead, when the user wears earbuds, we act on partials. Each time a partial translation arrives, we split it at sentence-ending punctuation (。!?.!?) and submit only the clauses that are complete. The trailing fragment stays unsent. A small tracker remembers how many sentences we already submitted for this segment, so a growing partial doesn't trigger repeats. When the final result arrives, we submit whatever remains.

// Simplified for illustration.
final _enders = RegExp(r'(?<=[。!?.!?])');

void onTranslation(String text, {required bool isFinal}) {
  final parts = text.split(_enders);
  // On a partial, the last piece may be an unfinished clause.
  final complete = isFinal ? parts : parts.sublist(0, parts.length - 1);

  for (var i = _submittedCount; i < complete.length; i++) {
    final clause = complete[i].trim();
    if (clause.isNotEmpty) _tts.enqueue(clause);
  }
  _submittedCount = complete.length;

  if (isFinal) _submittedCount = 0; // next segment starts fresh
}

We also enable the SDK's stable-partial setting, which reduces how often already-emitted partial text gets rewritten, and optionally semantic segmentation. Both make it safer to act on partials.

Why only with earbuds? Without them, the translation plays through the phone speaker, right next to the mic that's still listening. Playing early clauses there means the app partly hears and recognizes its own output. In speaker mode we wait for the final result. It's slower, but it doesn't feed back on itself.

After clause-level submission, the next visible gap was between clauses. Clause 1 finished playing, and then there was silence while clause 2 was synthesized.

The fix was to split one "speak this" queue into two:

While clause N is playing, clause N+1 is already synthesizing. When N finishes, N+1 is usually ready to go.

There's one more rule that matters in real conversations: stale audio is worse than missing audio. If the speaker talks quickly for a long time, the playback queue can fall behind until the listener hears translations of things said well before. So the playback queue has a cap. When it reaches 32 items, we keep only the newest 2 and drop the rest. Losing some content is a real cost, but an interpreter who is a minute behind isn't useful either.

// Simplified for illustration.
void enqueuePlayback(AudioClip clip) {
  _playback.add(clip);
  if (_playback.length >= 32) {
    final keep = _playback.sublist(_playback.length - 2);
    _playback
      ..clear()
      ..addAll(keep);
  }
  _pumpPlayback();
}

Two smaller wins in the same layer:

Voice cloning sounds expensive, and it would be if it were done per utterance. It isn't.

The user records a sample once during setup. An asynchronous job creates the voice with the provider, and we store the resulting reference: a speaker profile or voice id, plus the region it lives in. During a live session, synthesis just references that voice. With Azure, that means a speaker embedding referenced through SSML. With other providers, it's a voice_id parameter.

// Simplified for illustration.
final request = TtsRequest(
  text: clause,
  language: targetLang,
  voice: user.clonedVoice, // created once, offline, at setup
);

From the hot path's point of view, a cloned voice costs the same as a stock voice.

This isn't strictly latency, but it shapes the whole earbuds experience.

When an iOS app records audio while Bluetooth earbuds are connected, the system may route the input to the earbuds' mic. To do that, Bluetooth switches from A2DP (high-quality, output-only) to the SCO/HFP headset profile, which carries narrowband audio in both directions. Your translated speech, possibly in your own cloned voice, suddenly sounds like a phone call from 2005, and the earbud mic usually picks up worse audio than the phone's mic anyway.

Our fix: on iOS we pin capture to the phone's built-in mic and use the earbuds for output only. We re-apply that preference every time the audio route changes, because connecting, disconnecting, or switching devices can reset it. Earbuds stay in A2DP, and the output stays high quality.

On route changes we also wait about 100 ms after setting the output route before enabling the remote audio track in the RTC path. That adds a small delay, deliberately. We had cases where enabling the track too early sent the first audio to the wrong device. Here we chose correctness over speed.

Some of this is still open, and we'd rather say so.

Early audio can't be taken back. Clause-level playback acts on partial recognition. Stable partials help, but recognition sometimes revises text after we've already spoken a clause. Once audio has reached someone's ear, it can't be retracted. We accept occasional small inconsistencies in exchange for responsiveness, but it is a real trade-off and not a solved problem.

Segmentation behavior shifts. Where the recognizer decides a segment ends affects how our clause splitter behaves. Semantic segmentation helps in some languages and conversation styles and hurts in others, and provider behavior changes over time. We keep it optional and keep re-testing.

Streaming transport isn't streaming playback. Our TTS responses arrive as server-sent event chunks, but today the client collects all the chunks and plays the clip once synthesis is complete. We do this for integrity (no half-clips on network hiccups) and so the complete clip can go into the cache. Playing the first chunk as it arrives is the obvious next improvement, and it will need a different approach to caching and error handling.

Echo handling is a compromise. Some local playback paths capture during playback and resume after a configured guard of about 800 ms, so the app doesn't transcribe its own voice. This works, but it means the app is briefly deaf. Separately, we removed the iOS hardware voice-processing I/O unit because it interfered with system volume control of the translated audio, which users really did notice. We now rely on WebRTC's software echo cancellation instead. We're watching whether that trade holds up in noisy, real-world rooms. We have not eliminated echo. We've managed it.

"Connected" doesn't mean "ready". A transport-level connection says nothing about whether the translation agent on the other end is ready to process audio. We've had to separate those states in both code and UI. Reconnection uses bounded retries with backoff, an explicit reconnecting state, and a configuration re-sync after reconnecting, because a session that silently reconnects with stale config is worse than one that visibly fails.

Right now we instrument the parts we control directly:

That catches regressions in individual stages, but it isn't the number users feel: the time from when the speaker's mouth stops to when the listener's ear starts hearing the translation. We don't have reliable mic-to-ear p50/p95 numbers yet, and we don't want to guess. Measuring it properly means correlating timestamps across the device, the cloud services, and the audio output path, including Bluetooth output buffering, which the app can't see directly. That's our next instrumentation project. Until then, the changes in this post were judged by stage-level timings and side-by-side listening, not a headline number.

If there's one theme here, it's that perceived latency is mostly a scheduling problem: do session work before the user asks, never drop audio while you're getting ready, release output in the smallest units that are safe to act on, and keep the next unit preparing while the current one plays.

If you want to see how this feels in practice, Owll Translator is where this work ships.

We're curious how others handle the "can't retract spoken audio" problem. If you've built voice agents or live captioning, do you act on partial results, wait for finals, or do something in between? Let us know in the comments.

── more in #ai-tools 4 stories · sorted by recency
── more on @owll translator 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/shaving-latency-off-…] indexed:0 read:11min 2026-10-09 · —