{"slug": "preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal", "title": "Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority", "summary": "Researchers at urbanAIthi have developed a preference-conditioned multi-objective reinforcement learning controller for transit signal priority (TSP) that allows runtime tuning of the trade-off between bus delay and overall traffic delay without retraining. The controller, built on IntersectionZoo, outperforms fixed-time and rule-based baselines while maintaining constraint feasibility, though high bus-priority weights can increase non-bus tail delays. The source code is available on GitHub.", "body_md": "arXiv:2607.18286v1 Announce Type: new\nAbstract: Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles. Existing reinforcement-learning (RL) approaches to TSP typically encode transit-aware features (e.g., occupancy and schedule deviation) but optimize a fixed reward or fixed scalarization, which limits operational flexibility when agency priorities change across time-of-day or disruption conditions. We present a preference-conditioned TSP controller, $\\pi(a \\mid s,w)$, that selects the next signal phase under minimum/maximum green and transition-feasibility constraints and can be tuned at runtime via a preference parameter $w$ to trade off bus-priority emphasis against overall traffic delay without retraining. We implement this on top of IntersectionZoo by introducing a constrained signal-control/TSP wrapper, and we extend scenario generation with bus-prevalence augmentation and timetable-based bus insertion to address sparse transit-priority events during training. Experiments against fixed-time control, a rule-based TSP overlay, and fixed-weight PPO specialists show that a single learned conditioned policy spans a smooth empirical trade-off frontier across runtime preferences, outperforms fixed-time and rule-based baselines, and maintains constraint feasibility, while tail-delay diagnostics reveal that non-bus externalities remain limited for moderate preference settings but can increase substantially under high bus-priority weights. The source code of this work is available at https://github.com/urbanAIthi/morl-tsp.", "url": "https://wpnews.pro/news/preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal", "canonical_source": "https://arxiv.org/abs/2607.18286", "published_at": "2026-07-22 04:00:00+00:00", "updated_at": "2026-07-22 04:18:14.626210+00:00", "lang": "en", "topics": ["machine-learning", "autonomous-vehicles"], "entities": ["urbanAIthi", "IntersectionZoo", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal", "markdown": "https://wpnews.pro/news/preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal.md", "text": "https://wpnews.pro/news/preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal.txt", "jsonld": "https://wpnews.pro/news/preference-conditioned-multi-objective-reinforcement-learning-for-runtime-signal.jsonld"}}