MintAct: A Unified Visual Agent for Digital Environments
Researchers submitted MintAct, a family of vision-language models trained at 2B, 4B, and 8B scales that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use…