# Inworld TTS in Famulor: How to Test the Voice

> Source: <https://www.famulor.io/blog/inworld-tts-famulor-voice-agent-guide>
> Published: 2026-09-01 11:08:00+00:00

### Summarize Content With:

The [Famulor update dated
August 30, 2026](https://docs.famulor.io/changelog) adds Inworld as a new voice provider for AI
assistants. The changelog says two options are available and highlights
pronunciation and intonation for names, addresses, and numbers. That
matters in phone workflows where an attractive voice is not enough. A
mispronounced street name or an ambiguous sequence of digits can make
the next process step unusable.

The right question is therefore not whether the new voice is
universally “faster” or “better.” Test Inworld with the phrases, data,
and conversation patterns that your assistant actually uses. One current
product boundary is equally important: in Famulor, Inworld voices are
presently intended for assistants configured with **one
language**.

**Key takeaways**

- Inworld has been available as a Famulor voice provider since August 30, 2026.
- Famulor positions the new option particularly for precise delivery of names, addresses, and numbers.
- The current Famulor integration is intended for assistants configured with one language.
- Inworld API multilingual capability must not be confused with current multilingual configuration in Famulor.
- Select a voice through a reproducible listening test, not a single latency or quality claim.

## What exactly changed in Famulor?

The changelog introduces Inworld as a voice provider and describes two selectable variants with different trade-offs. It does not publish a mapping to specific model IDs. The labels shown in your workspace are therefore the authoritative reference for the options actually available to you.

You can preview voices in the [Famulor Inworld voice
library](https://www.famulor.io/voices/inworld). The general [Models and
voices documentation](https://docs.famulor.io/assistants/models-and-voices) recommends choosing a voice first and then
changing one compatible control at a time. The controls shown depend on
the selected model. Important scenarios should be retested after any
voice or model change.

This update does not replace speech recognition or the language
model. In a pipeline, STT converts the caller’s speech into text, the
LLM decides what to say or do, and TTS produces the audible reply.
Inworld affects the **voice output** in this architecture.
A failure to understand a name therefore needs a different diagnosis
from a failure to pronounce that name.

## Where is Inworld most relevant?

The new option is especially worth evaluating when spoken output needs to carry concrete data reliably.

| Conversation scenario | Typical test data | What to listen for | 
|---|---|---|
| [Appointment](/use-cases/appointment-booking-faqs) confirmation | date, time, weekday | clear emphasis and useful pauses | 
| Address confirmation | street, house number, city, postcode | correct word boundaries and digit grouping | 
| Name confirmation | personal, company, and product names | pronunciation, accent, and consistency | 
| Reference number | booking, ticket, or case number | no missing characters and sensible grouping | 
| Callback details | phone number and time window | understandable pace and reliable repetition | 

This does not make Inworld the automatic best choice for every assistant. An emotional service greeting, a brief status call, and the delivery of complex product codes impose different requirements. Use the same test set for Inworld and your current voice. That compares the real workflow instead of unrelated demos.

When a specialist term must first be transcribed correctly and then
spoken correctly, separate input from output. The existing [pronunciation
dictionary guide for AI voice agents](https://www.famulor.io/blog/custom-vocabulary-for-ai-voice-agents-pronunciation-guide) covers TTS output. For input,
the speech-recognition glossary is the relevant control.

## The key boundary: provider multilingual support is not platform multilingual support

Inworld’s public [model documentation](https://docs.inworld.ai/tts/tts-models)
currently describes Realtime TTS-2 and Realtime TTS-2 Flash. The
provider documents broad language and locale coverage for that model
family. Its [language
documentation](https://docs.inworld.ai/tts/capabilities/multilingual) lists more than 200 languages and locales.

Famulor still has a narrower current product rule: its changelog says
Inworld voices are presently for assistants that speak **one
language**. When additional languages are configured, Famulor
currently directs users to [ElevenLabs](https://www.famulor.io/alternatives/elevenlabs-alternative) or Cartesia, and the voice picker
accounts for that boundary.

These statements are not contradictory. They describe a model and a specific platform integration at different layers. A provider can expose capabilities that a platform has not enabled in every mode. Do not plan a multilingual rollout from the Inworld API documentation alone. Check the configuration in your Famulor workspace and test every required language in the intended assistant.

## The Hacker News signal: fewer milliseconds do not guarantee a better voice

In a [Hacker
News discussion dated August 21, 2026](https://news.ycombinator.com/item?id=49389952) about very fast TTS output,
practitioners emphasized time-to-first-audio. Several comments also
described a quality boundary: cadence, expression, voice quality, and
the full voice pipeline may matter more than optimizing a single
component.

This is a **community signal**, not a benchmark for
Inworld or Famulor. It does point to the questions a useful test must
answer. How quickly does the reply begin in a real phone call? Does the
voice stay stable during longer sentences? Are digit sequences easy to
understand? Does interruption handling remain clean? A provider’s
server-side latency metric cannot answer all of these because telephony,
turn detection, the LLM, and audio transport also contribute to
perceived response time.

## A reproducible five-step test plan

### 1. Build a fixed test set

Use 15 to 25 short utterances drawn from real conversation patterns. Cover names, streets, cities, phone numbers, times, decimals, abbreviations, and at least one longer explanatory sentence. Do not use real personal data when the test does not require it.

### 2. Change one variable only

Keep the prompt, STT, LLM, conversation flow, and phone connection unchanged. First switch only the voice or Inworld option. This makes differences more plausibly attributable to TTS output.

### 3. Test both preview and the real audio path

A text preview reveals pronunciation and voice character, but not the complete phone experience. Add a real test call or voice simulation. Review reply onset, volume, pauses, interruptions, and longer dialogue turns.

### 4. Score with a simple matrix

Mark each case as “correct,” “understandable with limitations,” or “unacceptable,” followed by a short observation. Avoid decimal scores that imply more precision than a small listener group and test set can support.

### 5. Repeat critical cases

One successful run is not enough. Repeat important names, addresses, and numbers in several sentences. Then adjust one compatible control or dictionary entry and rerun the same test set.

## How to choose without keyword or feature hype

Choose Inworld when your single-language configuration delivers
critical phrases more consistently and the complete call path fits the
workflow. Keep your existing voice when it better serves brand
character, multilingual operation, or a specialized speaking style. The
general guide to [choosing
a TTS provider](https://www.famulor.io/blog/how-to-choose-the-right-text-to-speech-tts-provider-for-your-ai-voice-agent) supports the broader architecture decision. This
article intentionally focuses on the new Inworld product update and its
rollout test.

Before a broad switch, review your fallback configuration as well.
According to the [Famulor
model reference](https://docs.famulor.io/assistants/models-and-voices), compatible fallback chains can move to an
appropriate alternative when the primary voice service errors, provided
the relevant plan feature is included. Confirm availability and visible
options directly under **Settings → Plan** and in the
assistant editor.

## Conclusion: precise speech needs a precise test

Inworld expands Famulor’s voice selection with an option that the changelog particularly recommends for names, addresses, and numbers. The strongest use case is not replacing every voice. It is improving critical output in a single-language assistant. Separate provider capability from the current Famulor integration, compare one variable at a time, and evaluate the complete phone experience. That turns a new model into a traceable product decision.

*Sarah Müller writes about Voice AI, telephony, and practical
automation at Famulor. Product details in this article were checked
against the linked documentation as of August 31, 2026.*

Writer at Famulor
