cd /news/machine-learning/elite-weighted-supervised-fine-tunin… · home topics machine-learning article
[ARTICLE · art-118534] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization

Researchers introduced Elite-Weighted Supervised Fine-tuning (EW-SFT), a reward-guided method for molecular optimization that updates models using their native pretraining loss on high-scoring molecules, eliminating the need for trajectory-level reinforcement learning. In tests, EW-SFT outperformed native optimizers under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds and achieved comparable performance on a sample-efficiency benchmark, working across autoregressive, masked-diffusion, and discrete-flow generators.

read1 min views2 publishedSep 2, 2026

arXiv:2609.00189v1 Announce Type: new Abstract: Goal-directed optimization is essential for steering molecular generators to propose candidates with desired properties. However, it is often implemented with policy-gradient reinforcement learning, which requires a generation-trajectory log-probability whose form depends on the model architecture and generation procedure. This makes an optimizer difficult to reuse across architectures and conditional generative designs. Supervised fine-tuning needs none of that machinery, but its update is driven by a fixed dataset, so the reward never enters the update. We introduce Elite-Weighted Supervised Fine-tuning (EW-SFT), which uses reward to guide elite selection of high-scoring molecules, and updates the model by its own pretraining loss on that set. Ablations show that reward information is passed primarily through elite selection, rather than through continuous weighting within the selected set. Because the update consumes only scored molecules and the model's native loss, the same rule applies across autoregressive, masked-diffusion, and discrete-flow generators, and across de novo, motif-extension, and linker-design tasks. Under a fixed budget of 3D shape alignment oracle calls on two kinase reference compounds, EW-SFT consistently outperforms the corresponding native optimizers. It further improves goal-directed optimization under a 2D similarity oracle on four held-out references and achieves comparable performance on a sample-efficiency benchmark without a trajectory-level RL formulation. These results demonstrate that EW-SFT is a unified and effective optimizer across molecular generators, design constraints, references, and oracles.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/elite-weighted-super…] indexed:0 read:1min 2026-09-02 ·