cd /news/machine-learning/segbench-gc-testing-segmentation-inv… · home topics machine-learning article
[ARTICLE · art-116184] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning

Researchers introduced SegBench-GC, a benchmark testing segmentation invariance in offline goal-conditioned reinforcement learning, finding that artificial trajectory cuts degrade performance: in a PointMaze study with 35,000 cuts, success dropped from 50.5% uncut to 39.1% with continuation-valid targets and 19.1% when treated as absorbing, with naive mean success ranging from 4.8% to 31.9% across realizations. The study also showed similar failures on Puzzle-4x5 using an independent baseline, where success fell from 47.2% uncut to 0.27% naive.

read1 min views1 publishedAug 31, 2026

arXiv:2608.27678v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) often uses trajectory structure for future-goal sampling and multi-step targets, yet logged trajectories may be partitioned for administrative reasons that do not correspond to termination. We introduce SegBench-GC, a controlled stress test of segmentation invariance that holds transitions, source trajectories, goal sampling, optimization settings, and evaluation fixed while varying only artificial backup boundaries and whether those boundaries retain continuation value. Continuation-valid targets (CVT) provide the segmentation-consistent control: reward accumulation stops at an artificial cut, but the target bootstraps from its stored successor. In a matched-count PointMaze study with 35,000 artificial cuts, three segmentation realizations, and three optimization seeds, final 50-episode-per-task success is 50.5% uncut, 39.1% with CVT, and 19.1% when the same cuts are treated as absorbing; across segmentation realizations, naive mean success ranges from 4.8% to 31.9%. An independent published n-step baseline (n=25) from the Decoupled Q-Chunking codebase shows the same failure on Puzzle-4x5: 47.2% uncut, 58.5% CVT, and 0.27% naive across three optimization seeds. A target-level diagnostic verifies the analytic target difference to numerical precision, and learned-critic diagnostics show a large optimistic shift under naive handling while CVT remains approximately aligned with the uncut critic. CVT applies standard continuation bootstrapping rather than a new Bellman rule; the contribution is the controlled benchmark, failure isolation, and cross-learner evidence that administrative segmentation can materially change multi-step offline GCRL.

── more in #machine-learning 4 stories · sorted by recency
── more on @segbench-gc 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/segbench-gc-testing-…] indexed:0 read:1min 2026-08-31 ·