Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution A paper titled "Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution" proposes speculative execution to hide tool latency in on-device cascaded voice agents, which typically serialize automatic speech recognition, large language model inference, and external tool execution so that tool latency is incurred only after the user finishes speaking and the LLM identifies the required tool calls. The approach targets the delay that arises from that serialized pipeline. Tool-augmented speech assistants typically serialize automatic speech recognition, large language model inference, and external tool execution. As a result, tool latency is incurred only after the user has finished speaking and the LLM has identified the required tool calls. We present speculative t