Building a Voice-Controlled Robot with Kotlin and LLMs A developer detailed a tutorial on building a voice-controlled robot using Kotlin and large language models. The design captures speech on Android, parses it with an LLM into structured commands, and validates them through a deterministic safety layer before execution on a robot. The approach emphasizes keeping the LLM out of the safety-critical control loop. Natural-language interfaces can make robots easier to operate. Instead of selecting individual buttons, an operator can say commands such as asking a robot to move, inspect an area, or report its status. In this tutorial, we will design an Android application that captures speech, sends the text through an LLM-based command parser, and converts the result into validated robot actions. User Voice ↓ Android Speech Recognition ↓ Command Text ↓ LLM Intent Parser ↓ Structured Robot Command ↓ Safety Validator ↓ Robot Gateway ↓ ROS 2 / Robot The LLM should not directly control motors. It should produce structured intent that a deterministic safety layer validates. Android provides speech-recognition APIs that can be used to convert spoken commands into text. A simplified flow is: fun onSpeechResult text: String { viewModel.processCommand text } The ViewModel can then pass the text to the command-processing layer. Define a strict command structure: sealed interface RobotCommand { data object Stop : RobotCommand data class Move val direction: String, val distanceMeters: Double : RobotCommand data class Rotate val degrees: Double : RobotCommand } Using a structured representation prevents the robotics layer from receiving arbitrary natural-language instructions. The LLM can transform: "Move forward two meters" into structured data such as: { "command": "move", "direction": "forward", "distance meters": 2 } Use structured output or a schema-constrained response where the selected LLM/API supports it. Never send the LLM result directly to the robot. Validate: LLM Output ↓ Schema Validation ↓ Range Validation ↓ Robot State Check ↓ Safety Policy ↓ Execution For example: fun validate command: RobotCommand : Boolean { return when command { RobotCommand.Stop - true is RobotCommand.Move - command.distanceMeters in 0.0..5.0 is RobotCommand.Rotate - command.degrees in -180.0..180.0 } } The exact limits should be determined by the robot's capabilities and safety requirements. After validation, the Android app sends a structured command: { "type": "move", "direction": "forward", "distanceMeters": 2.0 } The gateway converts this command into the appropriate ROS 2 service, action, or topic. The robot should return status information: { "state": "executing", "battery": 84, "position": { "x": 2.1, "y": 4.3 } } The Android application can show this in a Compose dashboard. Natural language can be ambiguous. For example: "Go over there." The system should not guess what "there" means. Instead, the application can request clarification: "I need a destination before I can move the robot." This is especially important for physical actions. The system becomes more powerful when voice and vision are combined. For example: User: "Follow the person wearing a red shirt." Voice → Intent Vision → Person Detection ↓ Robot Navigation The LLM can coordinate high-level intent while deterministic robotics components handle perception and navigation. Speech recognition and LLM inference can be deployed in different ways: Android | +-- Local speech recognition | +-- Cloud LLM or: Android | Local/Edge AI | Robot / Jetson Choose the architecture based on latency, privacy, connectivity, and hardware constraints. A recommended control hierarchy is: Natural Language ↓ LLM ↓ Structured Intent ↓ Deterministic Planner ↓ Safety Controller ↓ Robot The LLM should remain outside the final safety-critical control loop. Test commands using a simulated robot before physical deployment. Include: Combining Android, Kotlin, speech recognition, LLMs, and robotics creates a natural interface for Physical AI systems. The key design principle is to use AI for interpretation and high-level planning while deterministic software remains responsible for validation and safe physical execution. This architecture can be extended toward multimodal robot agents that combine voice, vision, maps, sensors, and autonomous task planning. SDK Flutter: https://github.com/v-modal/vmodal sdk flutter https://github.com/v-modal/vmodal sdk flutter SDK Android: https://github.com/v-modal/vmodal sdk android https://github.com/v-modal/vmodal sdk android Discord: https://discord.gg/K72z28KUx https://discord.gg/K72z28KUx