{"slug": "bag-of-decisions-reranker", "title": "Bag of Decisions Reranker", "summary": "Doug Turnbull's \"bag of decisions\" reranking approach beat both BM25 and single-question Jev reranking on two e-commerce datasets, reaching a mean NDCG of 0.6098 on Wayfair WANDS and 0.360 on Amazon ESCI. The technique prompts an LLM to generate a per-query rubric of yes/no relevance questions, then sums Jev's probability of \"yes\" for each question into a final rerank score. On WANDS, BM25 scored 0.5407 and Jev reranking 0.574; on ESCI, BM25 scored 0.289 and Jev reranking 0.314.", "body_md": "What if instead of search queries being interpreted as a “bag of terms” we could see them as a “bag of decisions”?\n\nWe know Jev acts as [frontier quality reranker](https://hevmind.com/writing/jev-as-a-reranker/). As a fun experiment, I wanted to try another technique: Don’t just ask Jev “Is this relevant for {query}?”. Instead, let an LLM generate the query’s criteria: a rubric to assess relevance for just this query.\n\nWe prompt an LLM with “generate yes / no question where the affirmative indicates relevance”. For the query `desk for kids` it might produce:\n\n```\n1 -  Is the document about desks designed for children?\n2 -  Does the document offer products or information on kids’ desks?\n3 -  Does it mention child-sized or age-appropriate desk dimensions or ergonomics?\n4 -  Does it reference desks for homework, study, or activities for kids?\n5 -  Are features specific to children’s desks (e.g., height-adjustable for kids, safety/rounded edges, storage fo\nr school supplies) discussed?\n6 -  Does it include buying guides, reviews, or comparisons of kids’ desks?\n7 -  Does it provide pricing or availability for children’s desks?\n8 -  Does it mention desks for specific child age ranges (e.g., toddlers, elementary, teens) as kids’ desks?\n9 -  Are brands or models of children’s desks listed?\n10 -  Does it cover where to buy or how to choose a desk for kids?\n```\n\nEach of these can be turned into a `noul` question. `P(yes)` of each can be summed to score overall relevance:\n\n```\ndocument = \"\"\"\nThis adjustable study desk is perfect for kids ages 5–12.\nIts height grows with your child.\n\"\"\"\n\nresponse = client.decide(\n    state=document,\n    questions={\n        \"question_1\": Noul(\n            instructions=\"Is this document about desks designed for children?\"\n        ),\n        \"question_2\": Noul(\n            instructions=\"Does the document offer products or information on kids’ desks?\"\n        )\n        ...\n    },\n)\n```\n\nInstead of a query being represented as a bag of terms, we have a **bag of questions**. Since Jev gives the probability of ‘yes’, we can simply sum the probability of each question being true as the rerank score:\n\n```\ntotal_probability = 0.0\n\nfor question_id, answer in response.answers.items():\n    probability = answer.noul\n    print(f\"{question_id}: {probability:.2f}\")\n    total_probability += probability\n```\n\nNow `total_probability` can be our reranker’s final score. We can add that to BM25, or just take it directly. In other words:\n\n## How does it stack up?\n\nI ran an experiment comparing BM25 vs two variants ([gory details here](https://github.com/softwaredoug/search-experiments/blob/main/research/decisions/ecom.md)):\n\n1. Jev reranking: i.e. “Is this document relevant for this query?”\n2. Bag of question reranking - multi-question flow described above.\n\nI then evaluated on two e-commerce datasets Wayfair WANDS and Amazon ESCI. How did it fair?\n\nDoes this work? It seems to. Here’s some initial results:\n\n| Dataset | Strategy name | Mean NDCG | Median NDCG | \n|---|---|---|---|\n| WANDS | `bm25` | 0.5407 | 0.475 | \n| WANDS | `jev_reranker` | 0.574 | 0.561 | \n| WANDS | `bag_of_decisions` | 0.6098137673305334 | 0.561 | \n| ESCI | `bm25` | 0.289 | 0.171 | \n| ESCI | `jev_reranker` | 0.314 | 0.205 | \n| ESCI | `bag_of_decisions` | 0.360 | 0.341 | \n\nOr graphically\n\nWe’re essentially doing per-query feature creation. Use an LLM as inspiration of what’s important to consider for *this query*. Then use Jev to measure whether we’re covering that criteria. We might wonder how we might modify the LLM call to provide more incisive + diverse questions then the ones above (IMO some are repetitive)\n\nIt’s a fun approach. We seem to be able to train more and more decision models, possible even [ones focused on specific domains](https://x.com/MaziyarPanahi/status/2108180624372605186). At least in this one toy experiment, it seems to be a promising approach worth deeper study. Enjoy!\n\n### Enjoy softwaredoug in training course form!\n\n#### Starting in October!\n\nSignup here -\n[https://maven.com/softwaredoug/cheat-at-search](https://maven.com/softwaredoug/cheat-at-search)\n\nI hope you join me at [Cheat at Search with Agents](https://maven.com/softwaredoug/cheat-at-search) to learn to use agents in search, build better RAG, and use LLMs in query understanding.", "url": "https://wpnews.pro/news/bag-of-decisions-reranker", "canonical_source": "http://softwaredoug.com/blog/2026/10/08/bag-of-decisions.html", "published_at": "2026-10-08 00:00:00+00:00", "updated_at": "2026-10-08 18:50:05.189393+00:00", "lang": "en", "topics": ["ai-search", "large-language-models", "ai-research", "natural-language-processing"], "entities": ["Doug Turnbull", "Jev", "Wayfair WANDS", "Amazon ESCI", "BM25"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/bag-of-decisions-reranker", "markdown": "https://wpnews.pro/news/bag-of-decisions-reranker.md", "text": "https://wpnews.pro/news/bag-of-decisions-reranker.txt", "jsonld": "https://wpnews.pro/news/bag-of-decisions-reranker.jsonld"}}