Mastering Predictable Voice AI: Building Enterprise Agents with Felona Voice & TypeScript A developer released Felona Voice, an open-source TypeScript framework for building deterministic voice agents that replaces per-turn LLM inference with pre-computed Joint Embedding Vectors and stateful VoiceGraph transition graphs. The project claims intent matching in roughly 5ms by comparing utterance embeddings against action-specific vectors in memory, aiming to cut latency, cost and hallucination risk for structured interactions like bookings and check-ins. In the rapidly evolving landscape of artificial intelligence, voice agents have emerged as a critical interface for customer interactions, support, and automation. Yet, the promise of seamless, intelligent voice experiences often collides with the harsh realities of latency, cost, and unpredictability when relying solely on large language models LLMs for core conversational logic. Traditional LLM-driven voice agents, while powerful for open-ended conversations, frequently struggle with the demands of structured, intent-driven interactions. Imagine a banking bot needing to confirm a transaction, or a healthcare assistant scheduling an appointment. In these scenarios, milliseconds matter, accuracy is paramount, and hallucinations are unacceptable. This is precisely where Felona Voice https://github.com/mohitjoer/felona voice https://github.com/mohitjoer/felona voice steps in, offering a revolutionary, open-source framework for building ultra-low-latency, deterministic voice agents in TypeScript. Before diving into Felona Voice's innovation, let's understand the pain points of current voice AI architectures that lean heavily on LLMs for every conversational turn: Felona Voice rethinks the core architecture of voice agents, moving beyond the limitations of auto-regressive LLM token generation for structured interactions. Its core innovation lies in the combination of Joint Embedding Vectors JEV and stateful conversational transition graphs, dubbed VoiceGraph . Instead of sending an entire prompt to an LLM for every decision, Felona Voice leverages pre-computed Joint Embedding Vectors. When a user speaks, their utterance is transcribed and immediately converted into an embedding. This embedding is then compared against a set of action-specific JEVs in sub-10ms ~5ms , directly within memory. This similarity matching instantly determines the user's intent and triggers the appropriate action defined in your VoiceGraph. Key Advantages of this approach: Felona Voice is built for developers, by developers. Its fluent TypeScript builder API makes crafting sophisticated voice agents intuitive and powerful. You define your agent's personality, available actions, and fallback behaviors with clear, concise code. js import { createAgent } from "felona-voice"; // Define a simple voice agent for a concierge service const conciergeAgent = createAgent "ConciergeBot" .system "You are an intelligent voice concierge for a luxury hotel. Your goal is to assist guests with bookings and information." .action "book table", "Book a restaurant reservation for a guest", async ctx = { // In a real application, this would interact with a booking system console.log Booking request received for: ${ctx.utterance} ; return "Certainly, I can help with that. For what time and how many people?"; } .action "check in", "Assist with guest check-in process", async ctx = { console.log Check-in request received for: ${ctx.utterance} ; return "Welcome Do you have a reservation name or number?"; } .action "ask weather", "Provide local weather information", async ctx = { // Simulate fetching weather data console.log Weather inquiry: ${ctx.utterance} ; return "The current weather in our location is sunny with a temperature of 25 degrees Celsius."; } .fallback "I'm sorry, I didn't quite catch that. Could you please rephrase or tell me how I can assist you today?" ; // Simulate an interaction async function runInteraction { console.log "User: Can I book a table for dinner?" ; let reply = await conciergeAgent.interact "Can I book a table for dinner?" ; console.log Agent: ${reply} ; // Output: Certainly, I can help with that. For what time and how many people? console.log "User: What's the weather like?" ; reply = await conciergeAgent.interact "What's the weather like?" ; console.log Agent: ${reply} ; // Output: The current weather in our location is sunny with a temperature of 25 degrees Celsius. console.log "User: Tell me a joke." ; reply = await conciergeAgent.interact "Tell me a joke." ; console.log Agent: ${reply} ; // Output: I'm sorry, I didn't quite catch that. Could you please rephrase or tell me how I can assist you today? } runInteraction ; This example demonstrates how easily you can define intents book table , check in , ask weather and their corresponding actions. The fallback mechanism ensures graceful handling of out-of-scope requests. Furthermore, Felona Voice offers pluggable audio pipelines for seamless integration with services like WebSockets, WebRTC, Deepgram, Whisper, ElevenLabs, and Cartesia, giving you full control over your audio stack. Crucially, for local development and testing, Felona Voice allows for deterministic routing without needing any external API keys, making the development loop incredibly fast and efficient. Let's put the benefits into perspective with a direct comparison: | Metric | Traditional Voice Agent LLM Loop | Felona Voice JEV + VoiceGraph | |---|---|---| | Intent Decision Latency | 850ms – 1,800ms | ~5ms Sub-10ms | | Inference Cost / Turn | $0.02 – $0.06+ / turn | $0.00 / turn | | Hallucination Risk | High probabilistic text tokens | 0% deterministic transition graph | | Network Dependency | Requires constant cloud LLM API | Local/In-memory embedding matching | Consider an enterprise handling 500,000 customer calls per month, with each call averaging 5 conversational turns. This translates to 2.5 million conversational turns monthly. This isn't just an optimization; it's a fundamental shift that enables massive scalability without proportional increases in operational costs for your core conversational intelligence. For enterprise applications, predictability isn't a luxury; it's a necessity. In sectors like finance, healthcare, or critical customer support, every interaction must be reliable, compliant, and consistent. The deterministic nature of Felona Voice's JEV and VoiceGraph architecture provides: Felona Voice empowers you to build the next generation of voice agents: fast, reliable, cost-effective, and entirely predictable. Move beyond the chaos of LLM-centric voice AI and embrace a framework designed for enterprise-grade performance and developer satisfaction. Ready to build your ultra-low-latency voice agent? 🌟 Star the repository on GitHub: github.com/mohitjoer/felona voice https://github.com/mohitjoer/felona voice 📦 Install via npm: npm install felona-voice 📖 Explore full documentation: felona-voice.mohitjoe.tech/docs https://felona-voice.mohitjoe.tech/docs This article was originally published on felona-voice.mohitjoe.tech https://felona-voice.mohitjoe.tech/blog/predictable-voice-ai-enterprise-agents-felona-typescript .