cd /news/ai-safety/certified-safety-curation-distributi… · home topics ai-safety article
[ARTICLE · art-128711] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning

A new arXiv paper (2609.12014v1) introduces "certified safety curation," a filter-then-clone pipeline that uses state-only value estimates trained from segment comparisons and Learn-then-Test calibration to certify a selection threshold under a distribution-free (α, δ) bound on the unsafe fraction of an offline reinforcement learning training set. The resulting policies satisfy the cost budget on 11 of 15 DSRL tasks, one short of cloning the ground-truth safe subset, while the uncertified variant reaches 12; retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. The authors report they are not aware of prior work certifying the composition of a training set for offline RL or imitation.

by read1 min views1 publishedSep 14, 2026

arXiv:2609.12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(\alpha, \delta)$ bound on the unsafe fraction of the selection, and behavior cloning follows. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on eleven of fifteen DSRL tasks, one short of cloning the ground-truth safe subset, which needs a label on every trajectory; the uncertified variant reaches twelve. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity the pool attains, which the calibration sample estimates and the scorer enters only through.

── more in #ai-safety 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/certified-safety-cur…] indexed:0 read:1min 2026-09-14 ·