cd /news/ai-search/bag-of-decisions-reranker · home › topics › ai-search › article
[ARTICLE · art-147768] src=softwaredoug.com ↗ pub= topic=ai-search verified=true sentiment=↑ positive

Bag of Decisions Reranker

Doug Turnbull's "bag of decisions" reranking approach beat both BM25 and single-question Jev reranking on two e-commerce datasets, reaching a mean NDCG of 0.6098 on Wayfair WANDS and 0.360 on Amazon ESCI. The technique prompts an LLM to generate a per-query rubric of yes/no relevance questions, then sums Jev's probability of "yes" for each question into a final rerank score. On WANDS, BM25 scored 0.5407 and Jev reranking 0.574; on ESCI, BM25 scored 0.289 and Jev reranking 0.314.

by read3 min views1 publishedOct 8, 2026
Bag of Decisions Reranker
Image: Softwaredoug (auto-discovered)

What if instead of search queries being interpreted as a “bag of terms” we could see them as a “bag of decisions”?

We know Jev acts as frontier quality reranker. As a fun experiment, I wanted to try another technique: Don’t just ask Jev “Is this relevant for {query}?”. Instead, let an LLM generate the query’s criteria: a rubric to assess relevance for just this query.

We prompt an LLM with “generate yes / no question where the affirmative indicates relevance”. For the query desk for kids it might produce:

1 -  Is the document about desks designed for children?
2 -  Does the document offer products or information on kids’ desks?
3 -  Does it mention child-sized or age-appropriate desk dimensions or ergonomics?
4 -  Does it reference desks for homework, study, or activities for kids?
5 -  Are features specific to children’s desks (e.g., height-adjustable for kids, safety/rounded edges, storage fo
r school supplies) discussed?
6 -  Does it include buying guides, reviews, or comparisons of kids’ desks?
7 -  Does it provide pricing or availability for children’s desks?
8 -  Does it mention desks for specific child age ranges (e.g., toddlers, elementary, teens) as kids’ desks?
9 -  Are brands or models of children’s desks listed?
10 -  Does it cover where to buy or how to choose a desk for kids?

Each of these can be turned into a noul question. P(yes) of each can be summed to score overall relevance:

document = """
This adjustable study desk is perfect for kids ages 5–12.
Its height grows with your child.
"""

response = client.decide(
    state=document,
    questions={
        "question_1": Noul(
            instructions="Is this document about desks designed for children?"
        ),
        "question_2": Noul(
            instructions="Does the document offer products or information on kids’ desks?"
        )
        ...
    },
)

Instead of a query being represented as a bag of terms, we have a bag of questions. Since Jev gives the probability of ‘yes’, we can simply sum the probability of each question being true as the rerank score:

total_probability = 0.0

for question_id, answer in response.answers.items():
    probability = answer.noul
    print(f"{question_id}: {probability:.2f}")
    total_probability += probability

Now total_probability can be our reranker’s final score. We can add that to BM25, or just take it directly. In other words:

How does it stack up? #

I ran an experiment comparing BM25 vs two variants (gory details here):

  1. Jev reranking: i.e. “Is this document relevant for this query?”
  2. Bag of question reranking - multi-question flow described above.

I then evaluated on two e-commerce datasets Wayfair WANDS and Amazon ESCI. How did it fair?

Does this work? It seems to. Here’s some initial results:

Dataset Strategy name Mean NDCG Median NDCG
WANDS bm25 0.5407 0.475
WANDS jev_reranker 0.574 0.561
WANDS bag_of_decisions 0.6098137673305334 0.561
ESCI bm25 0.289 0.171
ESCI jev_reranker 0.314 0.205
ESCI bag_of_decisions 0.360 0.341

Or graphically

We’re essentially doing per-query feature creation. Use an LLM as inspiration of what’s important to consider for this query. Then use Jev to measure whether we’re covering that criteria. We might wonder how we might modify the LLM call to provide more incisive + diverse questions then the ones above (IMO some are repetitive)

It’s a fun approach. We seem to be able to train more and more decision models, possible even ones focused on specific domains. At least in this one toy experiment, it seems to be a promising approach worth deeper study. Enjoy!

Enjoy softwaredoug in training course form!

Starting in October!

Signup here - https://maven.com/softwaredoug/cheat-at-search

I hope you join me at Cheat at Search with Agents to learn to use agents in search, build better RAG, and use LLMs in query understanding.

── more in #ai-search 4 stories · sorted by recency
── more on @doug turnbull 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/bag-of-decisions-rer…] indexed:0 read:3min 2026-10-08 · —