{"slug": "check-twice-cut-once-with-llm-search-relevance-eval", "title": "Check twice, cut once with LLM search relevance eval", "summary": "Doug Turnbull's local_llm_judge experiment on the WANDS furniture e-commerce search dataset found that forcing an LLM to pick a more relevant product yielded 75.08% precision at 100% recall over 1000 pairs, while allowing a \"Neither - not confident\" answer raised precision to 85.38% but cut recall to 17.10%. Turnbull added a sanity check that swaps the left- and right-hand products and re-queries the LLM to catch position bias, since the WANDS dataset provides only absolute 0-2 human relevance labels rather than direct pairwise labels.", "body_md": "In my [previous post](https://softwaredoug.com/blog/2025/01/13/llm-for-judgment-lists) ([git repo](https://github.com/softwaredoug/local-llm-judge)) I ask an LLM to compare the relevance of two products from the [WANDS furniture e-commerce search dataset](https://github.com/wayfair/WANDS). I then compare the agreement with the preference of human search raters, hoping if I got close to agreement, I might be onto creating a good pairwise LLM search relevance evaluator without so many humans.\n\nI note that when letting the LLM chicken-out and say “I dont know”, precision improves. But naturally reduces the recall - the percentage of labeled product pairs - quite a bit.\n\nSo, as an example, we might start with the following, simple, forced decision, asking an LLM to tell us which chair is more relevant to the `leather chairs` query:\n\n```\nSystem: You are a helpful assistant evaluating search relevance of furniture products.\n\nPrompt: \n    Which of these products is more relevant to the furniture e-commerce search query:\n\n    Query: leather chairs\n\n    Product LHS: fashion casual work beauty salon task chair\n    Product RHS: angelyn cotton upholstered parsons chair in gray/white\n\n    Respond with just 'LHS' or 'RHS'\n\nResponse: LHS\n```\n\nIf humans also say “LHS” - the salon task chair - is more relevant, then this is a win!\n\nOf course, in the WANDS case, as with many datasets we don’t have direct pairwise labels. Instead we have absolute labels. We’re really just comparing which human label is higher. That is, on a 0-2 scale, LHS might be labeled by humans a 2 for completely relevant, RHS a 0 for absolutely horrid result for `leather chair`, etc. So we say humans see LHS as more relevant.\n\nThe above prompt, over 1000 pairs, gives:\n\n```\npoetry run python -m local_llm_judge.main --verbose --eval-fn name\n\n...\n\nPrecision: 75.08% | Recall: 100% (N=1000)\n```\n\nBut we can also allow the LLM to NOT label a pair due to lack of information:\n\n```\n    Neither product is more relevant to the query, unless given compelling evidence.\n    \n    Which of these product names (if either) is more relevant to the furniture e-commerce search query:\n    \n    Query: leather chairs\n    \n    Product LHS name: fashion casual work beauty salon task chair\n        (remaining product attributes omited)\n    Or\n    Product RHS name: angelyn cotton upholstered parsons chair in gray/white\n        (remaining product attributes omited)\n    Or\n    Neither / Need more product attributes\n    \n    Only respond 'LHS' or 'RHS' if you are confident in your decision\n    \n    Respond with just 'LHS - I am confident', 'RHS - I am confident', or 'Neither - not confident'\n    with no other text. Respond 'Neither' if not enough evidence.\n\nResponse:\n\n   Neither -  not confident\n```\n\nDoing this 1000 times gives:\n\n```\npoetry run python -m local_llm_judge.main --verbose --eval-fn name_allow_neither\n\n...\n\nPrecision: 85.38% | Recall: 17.10% (N=1000)\n```\n\nOut of the 17.10% labeled, 85.38% agreed with humans.\n\n## Adding an important sanity check - checking both ways\n\nTurns out, it’s important to check twice. Put the LHS product on the RHS, and double check that you the same result. In other words ask first: is ‘salon chair’ more relevant than ‘parsons chair’ for query `leather chair`? Then reset, swap, and ask that is `parsons chair` more relevant than `salon chair`? We can then account for any biases an LLM might have to attending to the first or second product listed.\n\nWe wrap the prompt (`eval_fn` below) in a python function that double checks:\n\n``` python\ndef check_both_ways(query, product_lhs, product_rhs, eval_fn):  # eval_fn is a function wrapping the prompt\n    \"\"\"Get pairwise preference from LLM, but double check by swapping LHS and RHS to confirm consistent result.\"\"\"\n    decision1 = eval_fn(query, product_lhs, product_rhs)  # This just calls the LLM with the prompt LHS to RHS\n    decision2 = eval_fn(query, product_rhs, product_lhs)  # Now check RHS first...\n\n    if decision1 == 'LHS' and decision2 == 'RHS':\n        return 'LHS'\n    elif decision1 == 'RHS' and decision2 == 'LHS':\n        return 'RHS'\n    return 'Neither'\n```\n\nAdding a command line switch check twice:\n\n``` bash\n$ poetry run python -m local_llm_judge.main --verbose --eval-fn name --check-both-ways\n```\n\nThis dramatically improves performance. When the above is run, precision goes up appreciability with a high degree of product coverage:\n\n```\nPrecision: 87.99% | Recall: 65.80% (N=1000)\n```\n\nEven more interesting, when combining with allowing ‘I dont know’, we get the highest precision. Though with a significant reduction in recall:\n\n``` bash\n$ poetry run python -m local_llm_judge.main --verbose --eval-fn name_allow_neither --check-both-ways\n\n...\n\nPrecision: 90.76% | Recall: 11.90% (N=1000)\n```\n\nSo to summarize where my efforts stand, just using product name, we can build this confusion matrix. Showing Precision / Recall for each approach:\n\n|  | Dont check | `--check-both-ways` | \n|---|---|---|\n| Force | 75.08% / 100% | 87.99% / 65 % | \n| Allow Neither | 85.38% / 17.10% | 90.76% / 11.90% | \n\nSo depending on your use case, you should pick the appropriate solution. Want a high degree of coverage and can tolerate a lot of mistakes. Use the force / dont double check. Want to tolerate very few mistakes and only flag results with real issues? Double check and allow the LLM to say it doesn’t know (ie ‘allow neither’).\n\nI tried creating evaluators looking only at a few other fields. Here’s showing the LLM only product “class” (as in classification). Example product classes:\n\n```\nProduct LHS class: Beds\nProduct RHS class: Kids Beds\nProduct LHS class: Coffee & Cocktail Tables\nProduct RHS class: Outdoor Fireplaces\n```\n\nRun with these 4 variants:\n\n```\npoetry run python -m local_llm_judge.main --verbose --eval-fn classs\npoetry run python -m local_llm_judge.main --verbose --eval-fn class_allow_neither\npoetry run python -m local_llm_judge.main --verbose --eval-fn classs --check-both-ways\npoetry run python -m local_llm_judge.main --verbose --eval-fn class_allow_neither --check-both-ways\n```\n\n|  | Dont check | `--check-both-ways` | \n|---|---|---|\n| Force | 70.5% / 100% | 87.76% / 58.0% | \n| Allow Neither | 87.01% / 17.70% | 84.47% / 10.3% | \n\nAnd repeating for the product’s full categorization hierarchy (like `Outdoor furniture > Seating > Adirondak Chairs...`)\n\n|  | Dont check | `--check-both-ways` | \n|---|---|---|\n| Force | 74.6% / 100% | 86.1% / 69.70% | \n| Allow Neither | 85.71% / 18.20% | 89.91% / 10.8% | \n\nFinally, noisiest of them all, product description:\n\n|  | Dont check | `--check-both-ways` | \n|---|---|---|\n| Force | 70.31% / 98.70% | 76.58% / 72.60% | \n| Allow Neither | 79.21% / 10.10% | 83.02% / 5.3% | \n\n(Note with product description even ‘forcing’ a decision the LLM still sometimes said it couldn’t tell, completely unprompted!)\n\n### Enjoy softwaredoug in training course form!\n\n#### Starting in October!\n\nSignup here -\n[https://maven.com/softwaredoug/cheat-at-search](https://maven.com/softwaredoug/cheat-at-search)\n\nI hope you join me at [Cheat at Search with Agents](https://maven.com/softwaredoug/cheat-at-search) to learn to use agents in search, build better RAG, and use LLMs in query understanding.", "url": "https://wpnews.pro/news/check-twice-cut-once-with-llm-search-relevance-eval", "canonical_source": "https://softwaredoug.com/blog/2025/01/19/llm-as-judge-both-ways", "published_at": "2026-09-28 23:18:07+00:00", "updated_at": "2026-09-28 23:47:47.889407+00:00", "lang": "en", "topics": ["ai-research", "large-language-models", "ai-tools", "ai-search"], "entities": ["Doug Turnbull", "local_llm_judge", "WANDS", "Wayfair"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/check-twice-cut-once-with-llm-search-relevance-eval", "markdown": "https://wpnews.pro/news/check-twice-cut-once-with-llm-search-relevance-eval.md", "text": "https://wpnews.pro/news/check-twice-cut-once-with-llm-search-relevance-eval.txt", "jsonld": "https://wpnews.pro/news/check-twice-cut-once-with-llm-search-relevance-eval.jsonld"}}