Bag of Decisions Reranker Doug Turnbull's "bag of decisions" reranking approach beat both BM25 and single-question Jev reranking on two e-commerce datasets, reaching a mean NDCG of 0.6098 on Wayfair WANDS and 0.360 on Amazon ESCI. The technique prompts an LLM to generate a per-query rubric of yes/no relevance questions, then sums Jev's probability of "yes" for each question into a final rerank score. On WANDS, BM25 scored 0.5407 and Jev reranking 0.574; on ESCI, BM25 scored 0.289 and Jev reranking 0.314. What if instead of search queries being interpreted as a “bag of terms” we could see them as a “bag of decisions”? We know Jev acts as frontier quality reranker https://hevmind.com/writing/jev-as-a-reranker/ . As a fun experiment, I wanted to try another technique: Don’t just ask Jev “Is this relevant for {query}?”. Instead, let an LLM generate the query’s criteria: a rubric to assess relevance for just this query. We prompt an LLM with “generate yes / no question where the affirmative indicates relevance”. For the query desk for kids it might produce: 1 - Is the document about desks designed for children? 2 - Does the document offer products or information on kids’ desks? 3 - Does it mention child-sized or age-appropriate desk dimensions or ergonomics? 4 - Does it reference desks for homework, study, or activities for kids? 5 - Are features specific to children’s desks e.g., height-adjustable for kids, safety/rounded edges, storage fo r school supplies discussed? 6 - Does it include buying guides, reviews, or comparisons of kids’ desks? 7 - Does it provide pricing or availability for children’s desks? 8 - Does it mention desks for specific child age ranges e.g., toddlers, elementary, teens as kids’ desks? 9 - Are brands or models of children’s desks listed? 10 - Does it cover where to buy or how to choose a desk for kids? Each of these can be turned into a noul question. P yes of each can be summed to score overall relevance: document = """ This adjustable study desk is perfect for kids ages 5–12. Its height grows with your child. """ response = client.decide state=document, questions={ "question 1": Noul instructions="Is this document about desks designed for children?" , "question 2": Noul instructions="Does the document offer products or information on kids’ desks?" ... }, Instead of a query being represented as a bag of terms, we have a bag of questions . Since Jev gives the probability of ‘yes’, we can simply sum the probability of each question being true as the rerank score: total probability = 0.0 for question id, answer in response.answers.items : probability = answer.noul print f"{question id}: {probability:.2f}" total probability += probability Now total probability can be our reranker’s final score. We can add that to BM25, or just take it directly. In other words: How does it stack up? I ran an experiment comparing BM25 vs two variants gory details here https://github.com/softwaredoug/search-experiments/blob/main/research/decisions/ecom.md : 1. Jev reranking: i.e. “Is this document relevant for this query?” 2. Bag of question reranking - multi-question flow described above. I then evaluated on two e-commerce datasets Wayfair WANDS and Amazon ESCI. How did it fair? Does this work? It seems to. Here’s some initial results: | Dataset | Strategy name | Mean NDCG | Median NDCG | |---|---|---|---| | WANDS | bm25 | 0.5407 | 0.475 | | WANDS | jev reranker | 0.574 | 0.561 | | WANDS | bag of decisions | 0.6098137673305334 | 0.561 | | ESCI | bm25 | 0.289 | 0.171 | | ESCI | jev reranker | 0.314 | 0.205 | | ESCI | bag of decisions | 0.360 | 0.341 | Or graphically We’re essentially doing per-query feature creation. Use an LLM as inspiration of what’s important to consider for this query . Then use Jev to measure whether we’re covering that criteria. We might wonder how we might modify the LLM call to provide more incisive + diverse questions then the ones above IMO some are repetitive It’s a fun approach. We seem to be able to train more and more decision models, possible even ones focused on specific domains https://x.com/MaziyarPanahi/status/2108180624372605186 . At least in this one toy experiment, it seems to be a promising approach worth deeper study. Enjoy Enjoy softwaredoug in training course form Starting in October Signup here - https://maven.com/softwaredoug/cheat-at-search https://maven.com/softwaredoug/cheat-at-search I hope you join me at Cheat at Search with Agents https://maven.com/softwaredoug/cheat-at-search to learn to use agents in search, build better RAG, and use LLMs in query understanding.