{"slug": "imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation", "title": "IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation", "summary": "Researchers introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable semantic maps for Object Goal Navigation (ObjectNav). The framework requires no training or modification of the underlying navigation policy and consistently improves object grounding and end-to-end navigation gains across benchmarks OVON and HSSD-rare, a new benchmark featuring semantically specific subcategories. The study highlights that translating localization gains to navigation performance depends critically on downstream detection quality, revealing a key systems bottleneck in long-tail embodied navigation.", "body_md": "arXiv:2607.25106v1 Announce Type: new\nAbstract: Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable maps. Retrieved images are encoded using a vision-language model, matched against the semantic map to produce similarity maps, and aggregated to yield context-aware localization. Notably, this requires no training or modification of the underlying navigation policy. To explicitly evaluate long-tail behavior, we present HSSD-rare, a new ObjectNav benchmark built on Habitat Synthetic Scenes and featuring semantically specific subcategories. Across both OVON and HSSD-rare, image-conditioned queries consistently improve object grounding and yield end-to-end navigation gains. Further analysis reveals that translating localization gains to navigation performance depends critically on downstream detection quality, highlighting a key systems bottleneck in long-tail embodied navigation.", "url": "https://wpnews.pro/news/imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation", "canonical_source": "https://arxiv.org/abs/2607.25106", "published_at": "2026-07-29 04:00:00+00:00", "updated_at": "2026-07-29 04:22:17.952698+00:00", "lang": "en", "topics": ["artificial-intelligence", "computer-vision", "large-language-models", "robotics", "ai-research"], "entities": ["IMPRINT", "Object Goal Navigation", "Habitat Synthetic Scenes", "HSSD-rare", "OVON"], "alternates": {"html": "https://wpnews.pro/news/imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation", "markdown": "https://wpnews.pro/news/imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation.md", "text": "https://wpnews.pro/news/imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation.txt", "jsonld": "https://wpnews.pro/news/imprint-image-conditioned-query-enrichment-for-long-tail-object-goal-navigation.jsonld"}}