cd /news/machine-learning/reward-transport-property-control-in… · home topics machine-learning article
[ARTICLE · art-56865] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Reward Transport: Property Control in Flow Matching via Noise-Space Alignment

Researchers introduce Reward Transport, a method that uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards, enabling controllable generation in flow matching without requiring an oracle, reward model, gradient guidance, or additional computation. On ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control, with opposite structural responses for different targets, ruling out a generic size bias. The approach is complementary to classifier-free guidance and conditional flow matching, while a negative result under epsilon-prediction diffusion clarifies where coupling-level alignment is structurally absent.

read1 min views41 publishedJul 13, 2026

arXiv:2607.08781v1 Announce Type: new Abstract: The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controllable structure directly into the learned flow field. Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards; at inference, varying this coordinate steers the generated distribution without requiring an oracle, reward model, gradient guidance, or additional computation. In the coupling-preserving limit, thresholding this coordinate recovers the Cross-Entropy Method's truncated reward distribution, providing a principled, continuously adjustable distribution-level control knob. Empirically, on ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control over its operating range; most tellingly, the same knob produces opposite structural responses for different targets, growing molecules for logP but shrinking them for QED, which rules out a generic size bias. The interface is complementary to classifier-free guidance and conditional flow matching, while a negative result under epsilon-prediction diffusion clarifies where coupling-level alignment is structurally absent. Code: https://github.com/KehanGuo2/reward-transport

── more in #machine-learning 4 stories · sorted by recency
── more on @zinc-250k 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reward-transport-pro…] indexed:0 read:1min 2026-07-13 ·