cd /news/computer-vision/miner-multi-crop-inference-time-enha… · home topics computer-vision article
[ARTICLE · art-138786] src=arxiv.org ↗ pub= topic=computer-vision verified=true sentiment=↑ positive

MINER: Multi-crop INference-time Enhancement for Rare-Object Retrieval with Frozen Dual Encoders

MINER, a training-free inference framework from researchers publishing on arXiv (2609.27142v1), improves text-to-image retrieval with frozen dual encoders by augmenting a single global image embedding with a small bank of region-level embeddings plus hubness-correcting similarity rescoring. Experiments on CLIP, SigLIP, and SigLIP 2 show MINER improves retrieval on every backbone, both on the new ROCS benchmark — built from high-clutter subsets of Flickr30K and MS COCO re-captioned to name a single low-salience object — and on standard splits, with gains attributed primarily to broader spatial coverage rather than precise crop placement. Code is available at github.com/aalquwayfili/MINER and the ROCS dataset at huggingface.co/datasets/aalquwayfili/ROCS.

by read1 min views2 publishedSep 24, 2026

arXiv:2609.27142v1 Announce Type: new Abstract: Text-to-image retrieval with frozen dual encoders degrades when the query names a small, visually subordinate object in a cluttered scene: a single global image embedding underrepresents the localized visual evidence. We present MINER, a training-free inference framework that augments a frozen dual encoder's global image embedding with a small bank of region-level embeddings and a hubness-correcting similarity rescoring, recovering visual evidence that global pooling underweights. To evaluate this setting, we introduce ROCS, a benchmark built from high-clutter subsets of Flickr30K and MS COCO whose images are re-captioned to name a single low-salience object. Experiments on CLIP, SigLIP, and SigLIP 2 show that MINER improves retrieval on every backbone, on ROCS and on the standard splits. Analyses show that these gains come primarily from broader spatial coverage rather than precise crop placement, revealing a simple and general way to recover localized evidence from frozen representations. Code: https://github.com/aalquwayfili/MINER. Dataset: https://huggingface.co/datasets/aalquwayfili/ROCS.

── more in #computer-vision 4 stories · sorted by recency
── more on @miner 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/miner-multi-crop-inf…] indexed:0 read:1min 2026-09-24 ·