{"slug": "retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and", "title": "RetrievalFormer: A Dual-Encoder Transformer for Efficient Approximate Nearest Neighbor Retrieval and Cold-Item Recommendation", "summary": "A new dual-encoder transformer model, RetrievalFormer, can score new items from features alone, keeping a shared search-and-recommendation index open to cold items without retraining, according to an arXiv paper (2608.24079v1). In a public log, 38.6% of held-out query-search impressions involved items never previously shown or visited, and the feature-based tower matched warm performance (0.9595 vs 0.9510 Recall@20) against 99 sampled negatives. On MovieLens-1M, warm accuracy trails the strongest retrained baseline by 5.2% Recall@20 and 11.4% NDCG@20, but under strict zero-leakage cold-start evaluation, the content tower achieves 0.172 ± 0.006 Recall@20, 1.4× the strongest dedicated method and 3× a training-free floor.", "body_md": "arXiv:2608.24079v1 Announce Type: cross\nAbstract: A shared search-and-recommendation index must score new items from features alone because search has no exploration slot. In a public log covering both surfaces over one catalog, $38.6\\%$ of held-out query-search impressions show an item never previously shown or visited. For user-cold engagements, the feature-based tower serves this demand without measurable loss against $99$ sampled negatives ($0.9595$ Recall@20 versus $0.9510$ warm). A lexical baseline reaches similar parity, while a full-catalog check remains statistically undecided. Dual-encoder retrieval therefore keeps the index \\emph{open} to new items, unlike an ID-softmax recommender that requires retraining. We price this openness on recommendation against six sequential baselines, each retrained and tuned through five rounds on corrected targets. A float32 timestamp bug had reordered leave-one-out targets for $19.7\\%$ of users. On MovieLens-1M, warm accuracy trails the strongest retrained baseline by $5.2\\%$ Recall@20 and $11.4\\%$ NDCG@20. On MIND, the gap narrows to $0.8$--$3.6\\%$ relative to the five strongest baselines, though the model ranks sixth of seven. Under strict zero-leakage cold-start evaluation, the content tower achieves $0.172 \\pm 0.006$ Recall@20, $1.4\\times$ the strongest retrained dedicated method ($0.124 \\pm 0.007$) and $3\\times$ a training-free floor, without cold-specific training. Exact full-softmax training raises Recall@20 by $54\\%$ on MIND-small and $6.9\\%$ on MovieLens-1M over sampled InfoNCE, but recomputes the full catalog each step and exhausts accelerator memory at $240$K items. Approximate nearest-neighbor search explains none of the remaining gap, serving cost does not regress against ID-softmax retrieval, and a history-window sweep explains half the post-recipe remainder. Exact-quality training at catalog scale remains the open problem.", "url": "https://wpnews.pro/news/retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and", "canonical_source": "https://www.machinebrief.com/news/retrievalformer-a-dual-encoder-transformer-for-efficient-app-hqxe", "published_at": "2026-08-26 04:00:00+00:00", "updated_at": "2026-08-26 07:14:07.499998+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-products"], "entities": ["RetrievalFormer", "arXiv", "MovieLens-1M", "MIND", "InfoNCE"], "alternates": {"html": "https://wpnews.pro/news/retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and", "markdown": "https://wpnews.pro/news/retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and.md", "text": "https://wpnews.pro/news/retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and.txt", "jsonld": "https://wpnews.pro/news/retrievalformer-a-dual-encoder-transformer-for-efficient-approximate-nearest-and.jsonld"}}