#
Building Shiksha: An AI English Coach for Indian Learners
For many Indian learners, the biggest barrier to speaking English fluently isn't a lack of vocabulary or grammar rules learned in school—it's speaking anxiety and the fear of making mistakes in front of peers or teachers. Over the past 10 days, as part of the 10 Days of Voice Agents — Voice for Bharat Edition under the Learning & Literacy track, I built Shiksha: an interactive, real-time AI English Communication Coach designed to provide friendly, judgment-free spoken practice.
#
🌟 Why Voice?
Text chatbots don't build spoken confidence. Reading and typing are passive activities, whereas real-world conversations require instant auditory processing, cognitive framing, and spoken articulation.
Shiksha gives learners a low-latency, empathetic voice partner that understands Hinglish (code-mixed Hindi and English), allowing them to practice daily presentations, grammar rules, and workplace conversations without embarrassment.
#
🏗️ High-Level Architecture
User Speech (WebRTC / SIP) ──► LiveKit Audio Ingest
│
▼
Speech-to-Text (STT) │
▼
LLM + Tools (agent.py + db.py) │
▼
Murf Falcon (Ultra-Low Latency TTS) │
▼
Audio Output ◄────────────── WebRTC Audio Sink
#
🚀 Key Features Built Over the 10 Days
Ultra-Low Latency Indian Voice: Powered by Murf Falcon TTS, Shiksha delivers natural, culturally resonant Indian English voice output with near-instant response times. #
Persistent Conversational Memory (SQLite): Retains learner names, historical presentation goals, and specific practice needs across calls (agent_memory.db
). #
Curriculum-Driven Vocabulary Tools: Dynamically fetches context-specific vocabulary drills from exercises.json
and evaluates sentences live. #
Outbound Daily Practice Telephony (LiveKit SIP): Initiates automated daily check-in calls straight to a learner's phone. #
Human-in-the-Loop Escalation & Privacy Guardrails: Detects severe learner frustration or explicit requests for human mentors, requests explicit permission, and logs sanitized support tickets with clear reference IDs. #
Call Analytics Dashboard: A real-time Next.js dashboard displaying aggregated metrics (Total Calls, Successful Drills, Incomplete Calls) with zero personal transcripts exposed. #
Multi-Agent Specialist Handoff: Dynamically transitions the call from Shiksha (general coach) to Arjun (Grammar Specialist with a distinct male voice persona) for complex syntactic queries without dropping the WebRTC session.
#
🛠️ Hardest Technical Challenges & Fixes
- Hindi/Devanagari Pronunciation Glitches in TTS
Issue: Romanized Hindi text caused phonetic glitches in English voice models. #
Fix: Structured the system prompt to output pure Hindi terms in native Devanagari script (नमस्ते!
), allowing Murf Falcon to pronounce localized nuances cleanly.
- Next.js Dashboard Real-Time Cache vs. SQLite
Issue: Call logs updated in SQLite, but the Next.js /dashboard
served cached numbers. #
Fix: Enforced dynamic rendering with export const dynamic = "force-dynamic"
and export const revalidate = 0
at the top of the dashboard page.
- Context Preservation During Specialist Handoff
Issue: Switching agents risked losing conversational context, requiring the user to repeat themselves. #
Fix: Implemented dynamic prompt-state switching in the same LiveKit session loop, passing the handoff_reason
and recent turns directly into Arjun's context.
#
💻 How to Run the Project Locally
- Clone Repository & Setup Backend