{"slug": "off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code", "title": "Off-Policy Evaluation for Semantic ID Recommenders: Does the Model's Own Code Hierarchy Help?", "summary": "A new arXiv paper (2608.28905v1) finds that off-policy evaluation (OPE) for semantic ID recommenders can be made feasible by using the model's own code hierarchy as an action abstraction, but the benefit comes from coarsening rather than the hierarchy itself. The authors show that per-item OPE is hopeless on production logs due to small effective sample sizes, while marginalizing items to code-prefix clusters restores estimable support and reduces error, with resolution depth as the key control and a conditional bias bound linking coarsening bias to the quantizer's reconstruction residual and target-logging divergence.", "body_md": "arXiv:2608.28905v1 Announce Type: new\nAbstract: Generative recommenders increasingly emit semantic IDs (SIDs): each item is a short sequence of hierarchical discrete codes from a residual quantizer, decoded autoregressively. Before spending scarce A/B-test, a team may decide offline which decoder or reranking variants are worth testing - a job for off-policy evaluation (OPE). We ask a simple question: can the model's own SID tree serve as the action abstraction for that OPE? Our answer has three parts. (i) Under the near-argmax logging real recommenders use, per-item OPE is hopeless - as item-level effective sample size is usually small on production logs - but marginalizing items to code-prefix clusters restores estimable support and cuts error. (ii) This gain is thanks to coarsening, not to the hierarchy specifically; but the SID tree is what makes coarsening feasible in a generative system - each cluster's mass is exactly and cheaply returned by the decoder, whereas flat clustering requires enumerating item/leaf masses that a code-only decoder does not directly expose. (iii) Resolution depth is the operative knob - coarser under scarce support - and a conditional bias bound links the coarsening bias to the quantizer's worst-case reconstruction residual and the target-logging divergence.", "url": "https://wpnews.pro/news/off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code", "canonical_source": "https://arxiv.org/abs/2608.28905", "published_at": "2026-09-01 04:00:00+00:00", "updated_at": "2026-09-01 04:25:30.751060+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code", "markdown": "https://wpnews.pro/news/off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code.md", "text": "https://wpnews.pro/news/off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code.txt", "jsonld": "https://wpnews.pro/news/off-policy-evaluation-for-semantic-id-recommenders-does-the-model-s-own-code.jsonld"}}