Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selectio
Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution