Discretizing Reward Models

wpnews.pro

cd /news/machine-learning/discretizing-reward-models · home › topics › machine-learning › article

[ARTICLE · art-45807] src=arxiv.org ↗ pub=2026-07-01T00:53Z topic=machine-learning verified=true sentiment=· neutral

Discretizing Reward Models

Researchers have found that continuous reward models in reinforcement learning are oversensitive, assigning different scores to equally good responses, which can lead to bad policies. They propose evaluating reward models with separate measures of discriminative ability and specificity, and introduce a training-free algorithm using Monte Carlo dropout to discretize rewards, reducing oversensitivity and improving policies.

read2 min views1 publishedJul 1, 2026

Image: source

[Submitted on 19 Jun 2026]


[View PDF](/pdf/2606.21795)

[HTML (experimental)](https://arxiv.org/html/2606.21795v1)

Abstract:Despite their widespread use, the role of reward models in shaping reinforcement learning is poorly understood. Reward models offer a tempting promise: they automatically estimate response quality in the absence of verifiers or human judges. Unlike "verifiable rewards" which typically produce binary scores, reward models typically produce continuous scores, allowing them to be sensitive to fine-grained differences in responses. However, we show this apparent strength is a serious weakness: many popular reward models are oversensitive, assigning different scores to equally good responses. Theoretically, we show that seemingly perfect reward models can be highly oversensitive; empirically, this oversensitivity can lead to bad policies. In place of existing notions of "reward model accuracy," we propose evaluating reward models using distinct measures of "discriminative ability" and "specificity" (the complement of oversensitivity). As a solution, we describe a training-free algorithm that uses Monte Carlo dropout on any neural reward model to produce discrete reward clusters. Theoretically, we prove there exist discretizations that reduce oversensitivity at minimal expense of discriminative ability; empirically we show, in both controlled and natural RL settings, that discretizing rewards leads to less reward hacking and better policies than training on the original rewards.

References & Citations

...

Bibliographic Explorer

(What is the Explorer?) Connected Papers

(What is Connected Papers?) Litmaps

(What is Litmaps?) scite Smart Citations

(What are Smart Citations?)# Code, Data and Media Associated with this Article alphaXiv

(What is alphaXiv?) CatalyzeX Code Finder for Papers

(What is CatalyzeX?) DagsHub

(What is DagsHub?) Gotit.pub

(What is GotitPub?) Hugging Face

(What is Huggingface?) ScienceCast

(What is ScienceCast?)# Demos Influence Flower

(What are Influence Flowers?) CORE Recommender

(What is CORE?) IArxiv Recommender

(What is IArxiv?)# arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

source & further reading

arxiv.org — original article

── more in #machine-learning 4 stories · sorted by recency

smarterarticles.co.uk · 1 Jul · #machine-learning

Nobody to Blame: Who Pays When AI Agents Buy for You

zeit.de · 1 Jul · #machine-learning

Künstliche Intelligenz: US-Regierung hebt Blockade von Anthropics KI-Modellen auf

lesswrong.com · 1 Jul · #machine-learning

'AI allegory steganography' in Claude short stories in the Unslop contest?

9to5mac.com · 1 Jul · #machine-learning

Claude Fable 5 cleared to return as US lifts Anthropic’s export control restriction

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required