cd /news/machine-learning/crystalgrpo-target-aligned-and-cover… · home topics machine-learning article
[ARTICLE · art-89852] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction

Researchers introduced CrystalGRPO, a reinforcement-learning post-training framework for flow-based crystal structure prediction that jointly optimizes coordinates and lattice parameters. In tests on MP-20 and MPTS-52 datasets with PXRDGen and OMatG backbones, both CrystalGRPO-Q and CrystalGRPO-C variants reduced one- and twenty-sample RMSE compared to coordinate-only reinforcement across all four backbone-dataset settings, with CrystalGRPO-Q improving Top-1 recovery and CrystalGRPO-C achieving higher Top-20 recovery.

read1 min views1 publishedAug 10, 2026

arXiv:2608.06582v1 Announce Type: new Abstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. Reinforcement-learning post-training offers a flexible solution, yet existing approaches rely primarily on energy rewards and coordinate-only stochastic policies. Predicted energy does not identify the reference polymorph, while reward-driven concentration can reduce the candidate coverage required for Top-N recovery. We introduce CrystalGRPO, a CSP-aligned post-training framework that extends existing ODE-to-SDE policy constructions to the joint coordinate--lattice state. CrystalGRPO combines MACE-predicted energy with a StructureMatcher-based recovery score and provides two operating modes: CrystalGRPO-Q, which prioritizes single-draw recovery, and CrystalGRPO-C, which combines full-trajectory reference regularization with a coverage-aware group advantage to preserve finite-budget target recovery. Across MP-20 and MPTS-52 with PXRDGen and OMatG backbones, both variants reduce one- and twenty-sample RMSE relative to coordinate-only reinforcement in all four backbone--dataset settings. CrystalGRPO-Q consistently improves Top-1, whereas CrystalGRPO-C achieves a higher Top-20 across all settings.

── more in #machine-learning 4 stories · sorted by recency
── more on @crystalgrpo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/crystalgrpo-target-a…] indexed:0 read:1min 2026-08-10 ·