Relevance Is Not Preference: What We Learned Training ZooWork-ShopRanker ZooWork released ShopRanker, a family of e-commerce rerankers at 0.6B, 4B and 8B parameters, plus ShopRank-Bench, a contamination-limited preference benchmark built from private search traffic, according to the company's training writeup. The team reported that a panel of reasoning LLMs from three different families kept only 45.7% of the hard pairs it judged, that a 0.6B student distilled from an aligned 8B teacher scored 78.3 on hard development pairs versus 70.7 for the same student taught by a DeepSeek/Gemma panel, and that rank-16 LoRA on attention matched full fine-tuning at about 1% of trainable parameters. The released 0.6B and 4B models are fit to 135k scores from the aligned 8B, and the LoRA update is merged before deployment so each ShopRanker runs at its base model's latency. Relevance Is Not Preference: What We Learned Training ZooWork-ShopRanker How we aligned open rerankers to shopping preference: LLM-judge labels, a pairwise loss, distillation, LoRA, and what alignment does not fix. Open rerankers are very good at one question: is this document about the query? They were trained for web retrieval, where that question is most of the job. Shopping search asks a different one. Between two products that are both about “jumpsuit under $100, boho”, the better one is the one that is under $100 . A topical-relevance model has no reason to prefer it, and in our data it usually doesn’t. We recently released ZooWork-ShopRanker https://zoowork.ai/blog/shopranker-open-source-ecommerce-reranker/ , a family of e-commerce rerankers at 0.6B, 4B and 8B parameters, together with ShopRank-Bench, a contamination-limited preference benchmark built from private search traffic. The launch post https://zoowork.ai/blog/shopranker-open-source-ecommerce-reranker/ explains what the models do. This post is about how we trained them, which choices mattered, which didn’t, and where alignment stops helping. The full details, tables and significance tests are in the paper https://arxiv.org/abs/2609.31002 . In short: - Clean labels are the scarce resource, not compute. Real traffic gives us authentic queries and candidates but no labels. A panel of reasoning LLMs from three different families, judging both presentation orders under a constraint-first protocol, kept only 45.7% of the hard pairs it saw, and that filter was worth it. - For a scoring model, DPO https://arxiv.org/abs/2305.18290 is just a pairwise logistic loss. A reranker emits its reward directly, so the reference policy drops out. LoRA and a general-relevance replay mix do the anchoring a reference would otherwise do. - On-policy mining hits a clean-label ceiling. Mining the pairs the model finds hardest helped once, then stalled: the gold-label yield of a mining batch fell from 26% to 6%, because the hardest pairs became ones the judges themselves tie on. - Distillation beats direct alignment for the small models, and label volume beats teacher strength. The released 0.6B and 4B are fit to 135k scores from the aligned 8B. In controlled development runs, a 0.6B distilled from an aligned reranker teacher beat the same student aligned on judged pairs, and beat one taught by a costlier, more precise LLM panel 78.3 with an earlier aligned 4B teacher, versus 70.7 with the DeepSeek/Gemma panel teacher, on hard development pairs . - Rank-16 LoRA on attention matches full fine-tuning at about 1% of the trainable parameters. - Some things alignment does not install. Training on judged preference does not teach a model to follow a hand-designed attribute hierarchy which turns out to be the wrong target anyway , and it does not teach it to honor a stated budget. A small amount of programmatically labeled supervision does. - None of it costs anything at serving time. The LoRA update is merged before deployment, so each ShopRanker runs at its base model’s latency. Setup Models and the scoring readout All three ShopRankers are LoRA https://arxiv.org/abs/2106.09685 adapters on the Qwen3-Reranker https://arxiv.org/abs/2506.05176 family 0.6B, 4B, 8B . These are decoder models used as rerankers: the query and the product text go into a fixed prompt, and instead of generating an answer we read the logits of two answer tokens in a single forward pass. The relevance score is their difference: f q, d = logit "yes" | q, d − logit "no" | q, d No text is generated. This is the same “score the answer vocabulary instead of decoding it” pattern used by setwise rerankers https://arxiv.org/abs/2310.09497 and by our Instinct decision models https://zoowork.ai/blog/instinct-customer-service-decision-models/ , and it is what makes it affordable to evaluate every model, including a 27B LLM reference, on 10,511 pairs on one cost axis. We chose this family for practical reasons. One architecture spans all three sizes, so the aligned 8B can teach the smaller two. The bases are already strong: zero-shot, Qwen3-Reranker-8B matches Jina-Reranker-m0 https://huggingface.co/jinaai/jina-reranker-m0 on structured product text. And the yes/no score is exactly the scalar the pairwise loss acts on. What “better” means Preference here follows a two-tier rule taken from the judging protocol. A candidate must first satisfy the query’s hard constraints: the intended product type, a stated budget, the intended recipient, and any explicit exclusion. Only among the survivors does softer evidence style, color, fit, quality decide. The ordering is query-conditional. The hard constraints are whatever a given query states, not a fixed global ranking of attributes, and that distinction turns out to matter see the section on what alignment doesn’t install what-alignment-doesnt-install . Open rerankers fail exactly on the first tier. On the 1,843 gold-tier benchmark pairs, splitting by whether the query states an explicit constraint shows the gap clearly: Figure 1. Accuracy on gold-tier pairs, split by whether the query states an explicit constraint 684 constrained, 1,159 unconstrained; different pairs, so this is an association, not a controlled effect . Open baselines lose 14 to 20 points on constrained queries; ShopRanker is flat. The dashed line is chance. The effect is sharpest for budgets. On the 338 gold pairs that name a price cap, BM25 falls to 47.9%, below chance, the cross-encoders lose 26 to 34 points, and ShopRanker-4B and -8B reach 97.9%. Evaluation hygiene ShopRank-Bench contains 10,511 pairs from 2,991 queries in Gensmo’s private search traffic, each released in both a structured attribute format and a model-rendered natural-language format. Three details shaped how we read every result: 1. Pairs are adjacent in the production ranking. Randomly sampled pairs are dominated by easy contrasts that hide model differences. Products a deployed system already ranks next to each other are where a reranker has to earn its keep. 2. Labels are tiered by agreement. Gold pairs had all three judge families commit to the same product, silver two, bronze one. Bronze is 40% of the benchmark, and every system is 20 to 26 points weaker on it than on gold. An aggregate score is therefore dominated by the weakest labels, so we always report the tiers separately. 3. Pairs are not independent. 10,511 pairs come from 2,991 queries, so all intervals are query-clustered bootstrap intervals, and all model comparisons are paired McNemar tests https://doi.org/10.1007/BF02295996 on the same pairs, re-checked under a clustered bootstrap with Holm–Bonferroni https://www.jstor.org/stable/4615733 correction. Labels: a judge panel as a preference oracle Search logs give you queries and the products a system surfaced, but not which of two products a shopper should have seen first. Clicks are biased by position and presentation, and there are no clean pairwise labels in them. We used reasoning LLMs as the oracle and spent most of our design effort on making their labels trustworthy: - Three families, not one judge. Qwen3.5-122B, Gemma-4-31B and DeepSeek-V4-Pro each assess every benchmark pair. Inter-family agreement, not a single judge’s self-reported confidence, is the credibility signal. - Both orders. Each judge sees both candidate orders. A judge only counts as committing when both orders pick the same product; otherwise it counts as abstaining. - Constraint-first protocol. The judge must state the query’s hard constraints, check each product against each one, weigh the remaining evidence, and only then decide or declare a tie. - No conflicts. A pair is kept when every committed judge chose the same product and none contradicts it. Of 23,000 judged pairs, 10,511 survive 45.7% . Of the 12,489 discarded, 12,350 are unanimous ties and only 139 are outright conflicts. That discard rate is information, not waste: adjacent-rank pairs really do sit at decision boundaries that surface relevance can’t resolve. Training labels come from a strictly two-family panel Qwen3.5-122B and Gemma-4-31B, both agreeing in both orders on a disjoint slice of traffic, with zero query overlap and zero exact-pair overlap with the benchmark. DeepSeek-V4-Pro is withheld from training entirely, so the benchmark is not adjudicated only by the models that supervised the training. The objective: DPO without the reference Given a query q , a preferred product d+ and a rejected product d− , we minimize the pairwise logistic loss on the reranker’s score f : L = − log σ β · f q, d+ − f q, d− β = 0.1 This is the Bradley–Terry https://doi.org/10.2307/2334029 objective, RankNet https://doi.org/10.1145/1102351.1102363 with a temperature, and the loss used to fit reward models in RLHF https://arxiv.org/abs/2203.02155 . We claim no novelty for it. What is ours is the data it is trained on. Readers coming from preference optimization might ask where the reference policy went. In generative DPO the reference does two jobs: it parameterizes the implicit reward as a log-probability ratio, and it keeps the policy from drifting into degenerate text. Neither applies to a discriminative reranker. The model is the reward the yes/no logit gap , so the reparameterization collapses to the score itself and DPO reduces to the loss above. With no generated text to protect, the anchoring role falls to the low-rank update and to a replay mix of general relevance data: judge-labeled e-commerce pairs are sampled with probability 0.6 in the alignment stage and 0.8 in later stages, and the rest is general retrieval data. On-policy mining runs out of clean labels After the first alignment pass, the obvious next step is active learning: mine the pairs the current model finds hardest, those with the smallest score margin |s+ − s−| , and send them to the judges. One round helped a little development easy-slice accuracy 0.836 → 0.844 . Further rounds stalled, and the reason is worth knowing about if you plan to do the same thing. Alignment polarizes the model’s scores toward the extremes. Once it has, the smallest-margin pairs are no longer genuinely borderline. They are pairs of two good products or two bad ones, which the judges also tie on. The supply of clean, agreed labels dries up: the gold-label yield of a mining batch fell from 26% to 6% , even though those pairs are, by score, maximally confusable. The frontier had become judge-ambiguous rather than merely hard to find. On-policy preference learning, at least with an LLM-panel oracle and an agreement filter, is clean-label-limited. The released models do not use on-policy data. Distillation: volume beats teacher strength The two smaller models are not aligned from their own bases. The aligned 8B scores about 135k query–product examples 134,601 , each student is fit to those soft scores with binary cross-entropy, mixing structured and natural-language renderings of the same pool, and a final stage on judged pairs sharpens its pairwise preferences with the loss above. The teacher supplies coverage; the judges supply precision. We compared this against the direct route, aligning each base on judged pairs plus an on-policy round at 4B . On two-judge development slices: | Size | Stage | Easy | Hard | Hard-NL | |---|---|---|---|---| | 0.6B | Base | 77.7 | 72.6 | 71.2 | | | + judged pairs | 78.3 | 73.9 | 71.6 | | | + distill teacher | 82.5 | 78.3 | 73.8 | | 4B | Base | 82.1 | 76.4 | 75.8 | | | + judged pairs | 83.6 | 78.7 | 75.7 | | | + on-policy | 84.4 | 80.0 | 75.7 | | 8B | Base | 83.5 | 78.0 | 75.1 | | | + judged pairs | 87.2 | 82.3 | 78.5 | Table 1. Stage-by-stage development accuracy for the alignment-from-base route two-judge labels; not comparable to the three-judge benchmark . Judged-pair alignment is the main lever at 8B, distillation at 0.6B. In this ablation the 0.6B was distilled from an earlier 4B teacher; the released models use the 8B teacher. Two observations: - Distillation is the bigger lever at small scale. Judged pairs move the 0.6B by 1.3 points on hard pairs; distillation adds another 4.4. On the benchmark itself, the distilled 4B also beats the directly aligned 4B in both formats 81.2 vs. 80.4 structured, McNemar p = 0.006; 79.5 vs. 78.3 prose, p = 3.7×10⁻⁴ , so distillation, not only the 0.6B result, justifies the 4B’s route. - A more precise teacher lost to a more prolific one. We also built a teacher from DeepSeek and Gemma grades: 95.5% pairwise-consistent in validation, with only 0.5% reversals. At feasible scale it lost badly. The 0.6B distilled from it reached 70.7 on the hard development slice against 78.3 for the same student distilled from a reranker teacher the earlier aligned 4B of Table 1 . An LLM panel grades far fewer examples per unit of compute, and the coverage lost outweighed the precision gained. LoRA, and the stages that didn’t matter Every stage trains only a LoRA adapter: rank 16, α = 32, dropout 0.05, on the attention projections, base weights frozen. This was a choice, not a compute concession. In a controlled comparison at 0.6B, with identical data, recipe and evaluation and differing only in whether the update is low-rank, LoRA matched or slightly edged full fine-tuning: | Method 0.6B | Easy | Hard | Hard-NL | |---|---|---|---| | LoRA, rank 16 | 78.3 | 73.9 | 71.6 | | Full fine-tuning | 77.7 | 73.0 | 71.0 | Table 2. Controlled LoRA vs. full fine-tuning, development slices. That agrees with Thinking Machines’ LoRA Without Regret https://thinkingmachines.ai/blog/lora/ in the small-data regime: our core preference set is about 4.4k judged pairs plus replay, far below the point where adapter capacity should bind. The adapter is a few megabytes. Another stage we kept turned out to be neutral. On 500 hard development pairs, the structured-only 8B lost 4.6 points when the same products were written as prose, against 3.0 for the base 4B 76.6 to 73.6 , which suggested alignment on the structured schema makes a model brittle on prose. We added a final mixed-format pass. On the benchmark, the structured-only and mixed-format 8B are indistinguishable in both formats 83.3 vs. 83.4 structured, p = 0.94; 81.7 vs. 81.8 prose, p = 0.64 . Alignment on structured text alone already carried its gain to prose. We kept the stage, since it costs nothing, but we claim no robustness benefit for it. Other settings, for reproduction: maximum length 1,024 tokens, effective batch 64, learning rate 5×10⁻⁶ for alignment and distillation and 2×10⁻⁶ for the sharpening and mixed-format stages, linear warmup then decay, best checkpoint by pairwise accuracy on a held-out judged split. Results Figure 2. ShopRank-Bench accuracy on 10,511 structured pairs. Bars are overall accuracy ShopRanker in green ; markers show the gold, silver and bronze tiers. The dashed line is chance. The ranking holds on the natural-language view 8B 81.8 vs. 78.4 for the next best . ShopRanker-8B leads at 83.4%, 4.2 points above the strongest open rerankers we tested Jina-m0 and Qwen3-Reranker-8B, both 79.2% , and every ShopRanker significantly beats its own base, at every size and in both formats paired McNemar p ≤ 1.3×10⁻⁴ . The per-tier markers show why we report tiers: almost every system is above 90% on gold, and the separation comes as much from where each model loses on weaker labels as from where it wins. Figure 3. Accuracy against parameter count log scale; 95% query-clustered intervals . On structured text, the distilled 0.6B is statistically indistinguishable from the 4B base: a −0.9-point difference, 95% CI −1.9, +0.2 . That means no detected difference, not demonstrated equality. On prose, the 4B base keeps a significant edge. On one H200, that 0.6B serves 2.8× the throughput of the 4B base 433 vs. 155 documents per second at batch 32 on 40% of the memory 4.5 vs. 11.2 GB . Alignment helps at every scale, but the deployable win is distillation. Alignment also does not trade away general reranking. On three public MTEB https://arxiv.org/abs/2210.07316 reranking tasks we never train on ESCIReranking https://arxiv.org/abs/2206.06588 , SciDocsRR, AskUbuntuDupQuestions , ShopRanker-4B and -8B lead the average at 0.799, slightly above their bases 0.789 and 0.793 . The only slippage against a base is SciDocsRR −0.008 at 4B, −0.003 at 0.6B . Three tasks without uncertainty estimates support “does not degrade”, not “improves uniformly”. What alignment doesn’t install Two diagnostic tracks in ShopRank-Bench ask narrower questions than the preference track, and both returned answers we didn’t expect. A hand-designed attribute hierarchy is the wrong target The attribute-hierarchy AHP track has 1,500 pairs, each differing on exactly two attributes. The winner is whichever satisfies the higher-priority attribute under a hand-designed, intent-conditioned order by default audience product type price style color . Figure 4. Attribute-hierarchy diagnostic 1,500 pairs . Zero-shot, every model is below chance: they systematically invert the designed priority, price worst of all, and scale doesn’t help. Preference alignment lifts this to just above chance at 8B without any hierarchy supervision. The tempting conclusion is “train on the hierarchy”. It is wrong. A model trained to score the seven attributes separately and pool them reaches only 38.8% on judge-labeled preference pairs, and even the best linear re-weighting of those attribute scores reaches 69.5% held out. Judged preference is query-conditional, and a fixed priority order anti-correlates with it. This bounds one learned attribute model under linear pooling, not every possible design, but it was enough for us to keep the hierarchy as a diagnostic and not as a target. Preference alignment shifts the price bias, but doesn’t learn the budget Price was the worst hierarchy attribute, and we suspected missing supervision rather than missing capacity. Soft price preference is genuinely ambiguous, since a higher price can signal quality. A stated budget is not: price ≤ budget can be checked automatically. The budget track 1,302 pairs is adversarial by construction: - in the threshold slice 853 pairs , one product is under the stated budget and one is over, so the cheaper one wins; - in the control slice 449 pairs , the budget is raised above both prices and a secondary attribute makes the pricier product the winner. Always picking the cheaper product scores 100/0 on the two slices and always picking the pricier one scores 0/100, so only real threshold behavior is high on both. Figure 5. Explicit-budget diagnostic 1,302 pairs; markers show the threshold and control slices . General preference alignment overshoots toward cheap; a 0.6B specialist trained on the programmatic rule is strong on both slices. The weaker open rerankers lean toward the pricier product and effectively ignore the cap Jina-m0 35.1 and BGE-v2-m3 32.4 on the threshold slice . General preference alignment flips this, but overshoots: threshold accuracy jumps while control accuracy falls below the bases. That is a soft prefer-cheaper habit, not threshold logic. It is also why budget accuracy is not monotone in size: ShopRanker-8B 85.0% overall trails ShopRanker-4B 88.6% , because its control slice falls further 69.9 vs. 80.8 . Scale sharpens the lean; it does not install the rule. A 0.6B specialist trained with the same pairwise loss on programmatically labeled budget pairs, with replay mixed in, is strong on both slices and highest overall at 94.7% . On this constraint, a little targeted, automatically checkable supervision beats both scale and general alignment. We have not tested whether that transfers to other checkable constraints, and for hard caps in production a price filter is still the simplest guarantee. Is the margin just the training judges? Two of the three benchmark judge families also produced the training labels. If our gains merely reproduced those judges’ idiosyncrasies, they should shrink on pairs the training-disjoint family DeepSeek-V4-Pro also decided. On structured non-gold pairs: | Model | Gold 1,843 | DeepSeek voted 5,100 | Training panel only 3,568 | |---|---|---|---| | ShopRanker-8B | 96.5 | 83.6 | 76.2 | | ShopRanker-4B | 96.3 | 81.6 | 72.8 | | Jina-m0 2.4B | 90.8 | 80.7 | 71.0 | | 8B advantage over Jina-m0 | +5.7 | +2.9 | +5.2 | Table 3. Accuracy split by whether the training-disjoint judge family committed a verdict. The advantage survives on pairs DeepSeek helped decide, but it shrinks, from +5.2 to +2.9 for the 8B. We read this as bounding a training-family effect, not excluding one: the gain is not only an artifact of matching the training judges, but it is not free of that either. As a further spot check, a fourth model family that played no part in training or labeling agreed with 43 of 45 sampled labels. Serving cost Because the LoRA update is merged before serving, each ShopRanker costs exactly what its base costs. At the median, batch-1 latency is 17.6 vs. 18.6 ms at 0.6B, 25.2 vs. 24.5 at 4B and 25.8 vs. 25.4 at 8B one H200, fp16, 1,024-token limit . The accuracy gains are free per query. We also report the comparison that does not favor us. BGE-Reranker-v2-m3 https://huggingface.co/BAAI/bge-reranker-v2-m3 , which the 0.6B beats by 3.3 points on structured pairs, is 4.6× faster at batch 32 and needs a third of the memory. A decoder reranker carries a full language-model stack to emit one scalar; an encoder cross-encoder doesn’t. Where quality binds, the decoder wins; where raw throughput binds, the encoder still does. At the other end, two zero-shot reasoning LLMs are 8.5 to 8.7 points more accurate than ShopRanker-8B, at a median of 3.4 to 11.8 seconds per decision. That is a reference for how much of the benchmark’s preference is recoverable at all, not a deployable alternative. Limitations and open questions - The labels come from an LLM panel, not from shoppers. What we show is alignment to a reasoning panel that transfers to held-out traffic, not that the panel tracks what people actually buy. A human-agreement study is the main open item. - The benchmark is apparel-heavy and English-only. About two-thirds of pairs are clothing, shoes, bags and jewelry, and the queries are US-centric. - On-policy learning needs a better oracle. The clean-label ceiling suggests that once a model polarizes, an agreement-filtered LLM panel cannot keep supplying informative labels. Graded judgments, or oracles that can express “both acceptable, slightly prefer A”, might push past it. - Which other constraints are programmatically checkable? Budgets responded dramatically to a small amount of rule-labeled data. Recipient, exclusion and product-type constraints are partly checkable from catalog attributes. Whether the same trick generalizes is untested. - Is the format gap real? The aligned 8B’s prose gap 1.6 points is larger than its base’s 0.8 , but we have not run the interaction test that would say alignment increases format sensitivity. Try it The models are on Hugging Face under Apache 2.0 0.6B https://huggingface.co/srpone/zoowork-shopranker-0.6b , 4B https://huggingface.co/srpone/zoowork-shopranker-4b , 8B https://huggingface.co/srpone/zoowork-shopranker-8b . Each repository holds the unmodified Qwen3-Reranker base plus the LoRA adapter, so loading the base and then the adapter from the same repo reproduces our scores exactly. ShopRank-Bench https://huggingface.co/datasets/srpone/zoowork-shoprank-bench ships both formats, the per-judge verdicts so you can recompute the tiers and an evaluation script backed by our LookBench https://github.com/SerendipityOneInc/look-bench package. The benchmark data is CC BY-NC 4.0, for research and evaluation. The broader lesson is the one we keep relearning at ZooWork https://zoowork.ai/ : a general model gets you to “relevant”, and getting to “right” for a specific domain takes a clear definition of what better means, a way to label it cleanly, and the discipline to measure where training stops helping. Citation @article{xue2026shopranker, title = {ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker}, author = {Xue, Siqiao and Liu, Shuxuan and Hu, Ning}, journal = {arXiv preprint arXiv:2609.31002}, year = {2026} }