{"slug": "crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for", "title": "CrystalGRPO: Target-Aligned and Coverage-Preserving Reinforcement Learning for Flow-Based Crystal Structure Prediction", "summary": "Researchers introduced CrystalGRPO, a reinforcement-learning post-training framework for flow-based crystal structure prediction that jointly optimizes coordinates and lattice parameters. In tests on MP-20 and MPTS-52 datasets with PXRDGen and OMatG backbones, both CrystalGRPO-Q and CrystalGRPO-C variants reduced one- and twenty-sample RMSE compared to coordinate-only reinforcement across all four backbone-dataset settings, with CrystalGRPO-Q improving Top-1 recovery and CrystalGRPO-C achieving higher Top-20 recovery.", "body_md": "arXiv:2608.06582v1 Announce Type: new\nAbstract: Flow-based generative models can efficiently produce candidate structures for crystal structure prediction (CSP), but their pretrained objectives do not directly optimize downstream target recovery. Reinforcement-learning post-training offers a flexible solution, yet existing approaches rely primarily on energy rewards and coordinate-only stochastic policies. Predicted energy does not identify the reference polymorph, while reward-driven concentration can reduce the candidate coverage required for Top-N recovery. We introduce CrystalGRPO, a CSP-aligned post-training framework that extends existing ODE-to-SDE policy constructions to the joint coordinate--lattice state. CrystalGRPO combines MACE-predicted energy with a StructureMatcher-based recovery score and provides two operating modes: CrystalGRPO-Q, which prioritizes single-draw recovery, and CrystalGRPO-C, which combines full-trajectory reference regularization with a coverage-aware group advantage to preserve finite-budget target recovery. Across MP-20 and MPTS-52 with PXRDGen and OMatG backbones, both variants reduce one- and twenty-sample RMSE relative to coordinate-only reinforcement in all four backbone--dataset settings. CrystalGRPO-Q consistently improves Top-1, whereas CrystalGRPO-C achieves a higher Top-20 across all settings.", "url": "https://wpnews.pro/news/crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for", "canonical_source": "https://arxiv.org/abs/2608.06582", "published_at": "2026-08-10 04:00:00+00:00", "updated_at": "2026-08-10 04:13:22.977891+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["CrystalGRPO", "MACE", "StructureMatcher", "MP-20", "MPTS-52", "PXRDGen", "OMatG"], "alternates": {"html": "https://wpnews.pro/news/crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for", "markdown": "https://wpnews.pro/news/crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for.md", "text": "https://wpnews.pro/news/crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for.txt", "jsonld": "https://wpnews.pro/news/crystalgrpo-target-aligned-and-coverage-preserving-reinforcement-learning-for.jsonld"}}